مجموعهدادگانِ متونِ فارسیِ نتناک
NETNAK's Dataset of Persian Texts
Persian Texts for Natural Language Processing (NLP)
گردآوری: محمّدِ رجبپور
Compiled by Mohammad Rajabpur
فهرستِ دادگان
| عنوان | تعدادِ فایل | حجمِ کل | دانلود |
|---|---|---|---|
| ۱. اسنادِ سیاسی و حقوقی | ۱۰ | 1.2 MB | |
| ۲. ادبیاتِ کلاسیک | ۵ | 1.4 MB | |
| ۳. آثارِ عمرِ خیّام | ۴ | 299.4 KB | |
| ۴. آثارِ صادقِ هدایت | ۱۴ | 2.2 MB | |
| ۵. آثارِ جلالِ آلِاحمد | 16 | 4.5 MB | |
| ۶. اشعارِ سهرابِ سپهری | ۸ | 206.7 KB | |
| ۷. پیامهایِ متنیِ آیتاللّه سیّدمجتبی خامنهای | ۲۶ | 166.4 KB | |
| ۸. اخبارِ علم و فناوری | ۱۷ | 147.6 KB | |
| کلِ دادگان | ۱۰۰ | 10.1 MB |
منظور از «حجمِ کل» حجمِ کل فایلهایِ متنی در حالتِ غیرفشرده و زیپنشده است.
برایِ مشاهدهیِ جزئیاتِ دادگان در هر موضوعِ خاص بر رویِ عنوانِ دلخواهِ خود کلیک کنید.
Table of Contents
| Title | Data Items | Total Size | Download |
|---|---|---|---|
| 1. Legal & Political Documents | 10 | 1.2 MB | |
| 2. Classical Literature | 5 | 1.4 MB | |
| 3. Works of Omar Khayyam | 4 | 299.4 KB | |
| 4. Works of Sadegh Hedayat | 14 | 2.2 MB | |
| 5. Works of Jalal Al-e-Ahmad | 16 | 4.5 MB | |
| 6. Poems of Sohrab Sepehri | 8 | 206.7 KB | |
| 7. Textual Messages of Ayatollah Seyyed Mojtaba Khamenei | 26 | 166.4 KB | |
| 8. Science & Technology News | 17 | 147.6 KB | |
| The Whole Dataset | 100 | 10.1 MB |
The term "total size" refers to the cumulative file volume of the textual data prior to any compression or archival extraction.
To access comprehensive data pertaining to a particular subject, select the corresponding title of interest.
معرفیِ مجموعهدادگانِ متونِ فارسیِ نتناک
Introduction to NETNAK's Dataset of Persian Texts
مجموعهدادگانِ زیر شاملِ متونِ ارزشمندِ فارسی در حوزهها و موضوعاتِ مختلف است. این دادگان در قالبِ فایلهایِ متنی با فرمتِ کاربردی و مناسبِ Unicode utf-8 در اختیارِ شما پژوهشگران و مخاطبانِ گرامی جهتِ دانلود و استفاده قرار گرفته است. یکی از منابعِ مهم در گردآوریِ این دادگان ویکینبشته است و بایستی قدردانِ مشارکتکنندگانِ مختلف آن باشیم. متونی از ویکینبشته یا منابعِ دیگر انتخاب شدهاند که کامل یا نسبتاً کامل بودهاند و بهجز تنظیمِ فرمتِ فایل دخل و تصرفی در متنها صورت نگرفته است. در انجامِ پژوهش خود بکوشید تا این متون را با نسخههایِ مختلفِ الکترونیک و کاغذیِ موجود با دیدِ انتقادی مقایسه کنید و در صورتِ لزوم بههنجارسازی (normalization) متونِ مدنظرتان را در مرحلهیِ پیشپردازش فراموش نکنید. بهویژه رسمالخطِ متونِ کهن و قدیمیتر ممکن است تفاوتهایی با رسمالخطِ کنونی زبانِ فارسی داشته باشند و همچنین ممکن است کاتبانِ متون در سدههایِ پیش دچارِ لغزشهایی شده باشند و بایستی نسخهای را برگزینید که بیشترین اعتبار و اصالت را داشته باشد. اینجانب میکوشم تا به مرور متنهایِ بیشتری را به این مجموعه بیافزایم و گامهایی را برایِ تبدیلِ آن به یک پیکرهیِ بزرگ بردارم. در انتخابِ متون نهایتِ تلاش صورت گرفته است تا از متونِ فاقدِ کپیرایت استفاده شود. امیدوارم این مجموعهدادگان برایِ دانشجویان، اساتید و پژوهشگران در حوزههایِ «پردازشِ زبانِ طبیعی»، «زبانشناسیِ محاسباتی» و سایر رشتههایِ مرتبط سودمند باشد و انجامِ پروژهها و تحقیقات را برایشان سادهتر کند.
The present dataset comprises a rich collection of Persian texts spanning a diverse range of disciplines and subject areas. The dataset is made available to researchers and interested users in the form of text files encoded in Unicode UTF-8, a practical and appropriate format for download and subsequent utilization. A significant source contributing to the compilation of this corpus is Wikisource, to whose numerous contributors due acknowledgment is extended. The texts included have been selected from Wikisource or other sources on the basis of their completeness or relative completeness, and aside from necessary adjustments to file formatting, no substantive modifications have been made to the original content. Researchers are encouraged to approach these materials with a critical perspective, comparing them against existing electronic and print editions. Where necessary, attention should be given to normalization during the preprocessing stage, particularly in light of potential orthographic variations between older manuscripts and contemporary Persian usage. It is also advisable to remain mindful of possible scribal errors introduced in earlier centuries, and to prioritize versions that demonstrate the greatest scholarly credibility and authenticity. Efforts will continue to augment this collection progressively, with the long-term objective of developing it into a comprehensive corpus. In the selection of materials, every precaution has been taken to ensure the use of texts that are free of copyright restrictions. It is hoped that this dataset will prove beneficial to students, academics, and researchers in fields such as natural language processing, computational linguistics, and related domains, facilitating their project work and scholarly investigations.
اهمیتِ مجموعهدادگان
الف) صرفهجویی در زمان
دادگانِ متنی در برخی وبسایتها در چندین صفحهیِ مختلف برایِ یک اثر ارائه شدهاند و رونوشتِ مطالبِ تمامِ صفحات و ذخیرهیِ آنها بهصورتِ فایلِ متنیِ ساده زمانبر است، بهویژه اگر این کار را بخواهید برایِ تمامِ آثارِ یک مولف انجام دهید. بنابراین ضروری است که برایِ شتاببخشی به فرایندِ جمعآوریِ داده، متونِ موردِ نظر بهصورتِ یکجا و بهفرمتِ مناسب وجود داشته باشند.
ب) شناساندنِ منابعِ جمعآوریِ داده
منبعِ هر دادهیِ متنیِ موجود در این مجموعه ذکر شده و هایپرلینکِ آن موجود است. اگر شما نیز در حوزهیِ جمعآوریِ دادگان فعالیت دارید، خود نیز میتوانید با مراجعه به این منابع، نسبت به استخراجِ دادگانِ متنی دیگر یا بیشتری از آنها مبادرت ورزید.
ج) فراهمسازیِ دادهیِ متنی در زمانِ قطعی یا اختلالِ اینترنتِ بینالملل
در سالهایِ ۱۴۰۴ و ۱۴۰۵ چند بار به مدت زمانِ طولانی دسترسی به اینترنتِ جهانی قطع شد و امکانِ دسترسی به وبسایتهایی مانندِ ویکینبشته که دادگانِ متنیِ رایگان و بدونِ کپیرایت در اختیارِ کاربران قرار میدهند فراهم نبود. در آن زمان این نیاز احساس شد که باید برایِ اینگونه مواقع وبسایتهایی که در داخلِ ایران میزبانی میشوند نیازهایِ پژوهشگران و توسعهکنندگان را دستِکم برایِ مدتی کوتاه برطرف کنند. وبسایتِ نتناک توسطِ خدماتدهندگانِ ایرانی میزبانی میشود و در زمانِ قطعیِ اینترنتِ بینالملل در دسترس میباشد.
Significance of the Dataset
The significance of the present dataset may be articulated along three principal dimensions:
a) Temporal Efficiency in Data Acquisition
In numerous online repositories, a single textual work is often fragmented across multiple web pages, necessitating the laborious process of manually copying content from each page and consolidating it into a plain text file. This procedure becomes markedly more time-consuming when undertaken for the complete oeuvre of a given author. Consequently, the aggregation of desired texts in a unified location and in a readily processable format constitutes an essential measure for streamlining the data collection workflow.
b) Transparency and Traceability of Data Sources
For every textual item included in this collection, the original source is explicitly cited, and the corresponding hyperlink is provided. This attribution not only ensures scholarly accountability but also offers a valuable resource for researchers engaged in their own data-gathering endeavours, as they may refer to these primary sources for the extraction of supplementary or alternative textual materials.
c) Resilience during International Internet Disruptions
During the year 2026, Iran experienced multiple protracted interruptions in access to the global Internet, during which platforms such as Wikisource—renowned for providing freely accessible, copyright-free textual data—became temporarily unavailable. These occurrences underscored the pressing need for domestically hosted websites capable of addressing the requirements of researchers and developers, at least over limited durations, under such constrained conditions. The NetNak website, being hosted by Iranian service providers, remains accessible even during periods of international connectivity failure, thereby ensuring continuity of access to scholarly resources.
سلبِ مسئولیت
الف) احتمالِ نیاز به بررسیِ اصالت، ویرایش یا بههنجارسازیِ متون
دادگانِ موجود در این مجموعه فقط به منظورِ پردازشِ زبانِ طبیعی گردآوری شدهاند و هیچ کدام را نباید بهمانندِ کتابِ الکترونیکِ موردِ مطالعه قرار بدهید یا با دیگران به اشتراک بگذارید زیرا اصالتِ متون و میزانِ وفاداریِ آنها به نسخههایِ کاغذی آنها باید توسط کارشناسان کلمهبهکلمه و جملهبهجمله موردِ بررسی و موشکافی قرارگیرد و پس از پالایش و ویرایشِ متون آنگاه میتوان نسبتِ به انتشارِ آنها در قالبِ کتابِ الکترونیک اقدام کرد و آنها را در اختیارِ مخاطبان قرار داد. اینجانب در بررسی متون ویکینبشته (حتی متونِ برتر) و متونِ سایرِ منابع خطاها و لغزشهایی از جمله غلطهایِ املایی و تایپی و تفاوتهایِ جزئی با اصلِ متن را یافتهام. چون درصد این خطاها کم است در بررسیِ آماری و پردازشِ کامپیوتریِ متون احتمالاً میتوان از آنها چشمپوشی کرد اما برایِ ارائهیِ آنها به مخاطبِ انسانی باید با وسواسِ هرچهتمامتر تعدادِ لغزشها را به صفر میل داد. به فایلهایِ این مجموعه بهمثابهیِ متونِ بازی بنگرید که شما خود میتوانید در صورتِ لزوم اصلاحاتِ لازم را در آنها با رویکردی وفادارانه به اصل و خواستگاهشان انجام دهید.
ب) عدمِ تاییدِ محتوایِ متون و نظراتِ مولفانِ آنها
گنجاندنِ یک متن در این مجموعه لزوماً به معنایِ تایید یا ردِ محتوایِ آن توسطِ اینجانب نیست. تلاشِ من این بوده است که آثارِ مولفانی با دیدگاهها و سبکهایِ گوناگون را در این مجموعه گردآوری کنم تا دادگان دچارِ سوگیری به سمتِ افراد و متون خاصی نشوند. طبیعی است که بدین منظور حتی ممکن است متنهایی را برگزینم که به مذاقِ اینجانب خوش نیایند و ممکن است تمامی یا برخی ایدههایِ مطرح شده در آنها را نادرست پندارم. اما رعایتِ اصلِ بیطرفیِ علمی به اینجانب اجازه نمیدهد که بر اساسِ معیارهایِ شخصی متون را انتخاب کنم و میکوشم متنهایی را برگزینم که بازتابدهندهیِ اندیشهها و باورهایِ گروهها و طیفهایِ مختلفِ جوامعِ فارسیزبان باشند، چه اکثریت و چه اقلیت. متونی که در اینجا قرار دارند تنها برایِ پردازشِ کامپیوتری آنهاست و بههیچوجه قصد و نیتی برایِ اشاعهیِ افکار و عقایدِ پدیدآورندگانِ آنها، چه درست و چه نادرست، وجود ندارد.
ج) قوانینِ کپیرایت
تلاش شده است تا تنها از متونی استفاده شود که در مالکیتِ عموم قرار دارند و طبقِ قوانینِ جمهوریِ اسلامی ایران فاقدِ کپیرایت هستند. اما در استفادهیِ پژوهشی یا تجاری از این متون خود نیز بکوشید قوانینِ کپیرایت را دربارهیِ هر متن بسنجید بهویژه اگر در کشورِ دیگری ساکن هستید و قوانینِ کپیرایت آنجا متفاوت است.
Disclaimers
The following disclaimers are hereby articulated with respect to the dataset:
a) On the Necessity of Authenticity Verification, Editorial Intervention, and Normalization
The textual materials compiled in this dataset are intended exclusively for applications in natural language processing and computational analysis. Under no circumstances should they be treated as authoritative editions or disseminated as electronic books, as their fidelity to the original print sources and their overall authenticity require meticulous, line-by-line scrutiny by subject-matter experts. Only subsequent to rigorous refinement and editorial revision may such texts be considered suitable for publication in e-book format and for wider public distribution. In the course of reviewing materials drawn from Wikisource—including its higher-quality editions—as well as from other repositories, occasional orthographic and typographical errors, along with minor divergences from the source texts, have been observed. While the incidence of such inaccuracies is relatively low and may, for statistical and computational purposes, be deemed negligible, any presentation intended for human readership demands the utmost care to reduce such anomalies to zero. Users are therefore encouraged to regard the files in this collection as open-source textual resources, which they may themselves amend as necessary, provided that such corrections remain faithful to the original source material.
b) On the Non-Endorsement of Content and Authorial Views
The inclusion of any given text within this dataset does not imply, either explicitly or implicitly, the compiler's endorsement of or opposition to its content. The overarching objective has been to assemble works representing a broad spectrum of perspectives and stylistic variations, thereby mitigating any potential bias toward particular individuals or ideological positions. To this end, it has at times been necessary to incorporate texts with which the compiler may personally disagree or whose propositions he may consider partially or wholly untenable. Nevertheless, adherence to the principle of scholarly impartiality precludes the application of subjective criteria in the selection process; rather, every effort has been made to reflect the intellectual and doctrinal diversity characteristic of Persian-speaking communities, encompassing both majority and minority viewpoints. It must be emphasized that the texts are provided solely for computational processing, and no intention exists to propagate the ideas or beliefs of their authors, irrespective of their veracity.
c) On Copyright Compliance
Every reasonable effort has been made to include only those texts that fall within the public domain and are free of copyright restrictions under the applicable laws of the Islamic Republic of Iran. Nonetheless, users are strongly advised to independently ascertain the copyright status of each text prior to employing it for research or commercial purposes, particularly if they reside in jurisdictions where copyright regulations may differ substantially from those in force in Iran.
نحوهیِ دانلود
برایِ دانلودِ هر فایل بر رویِ نامِ آن یا آیکونِ آن کلیک کنید. در صورتی که میخواهید فایل را در دیرکتوریِ مطلوبِ خود ذخیره سازید بر رویِ آن رایتکلیک کنید و از گزینهیِ save as استفاده کنید.
How to Download the Data
To obtain any individual file, simply click on its corresponding name or icon. Should you wish to save the file to a specific directory of your choosing, right-click on the file and select the "save as" option from the contextual menu.
اسنادِ سیاسی و حقوقی
Legal & Political Documents
ادبیاتِ کلاسیک
Classical Literature
آثارِ عمرِ خیّام
Works of Omar Khayyam
آثارِ صادقِ هدایت
Works of Sadegh Hedayat
آثارِ جلالِ آلِاحمد
Works of Jalal Al-e-Ahmad
اشعارِ سهرابِ سپهری
Poems of Sohrab Sepehri
پیامهایِ متنیِ آیتاللّه سیّدمجتبی خامنهای
Textual Messages of Ayatollah Seyyed Mojtaba Khamenei
اخبارِ علم و فناوری
Science and Technology News
استفاده از دادگانِ متنیِ نتناک در پردازشِ زبانِ طبیعی با ذکرِ منبع مجاز است.
The use of NETNAK's textual dataset for natural language processing purposes is hereby permitted, provided that proper attribution is given.