Release Date: 5/16/2026
Format: TXT
Size: 18.93 MB
This dataset consists of a collection of Gujarati news articles and blog posts gathered from various online sources. It covers a wide range of topics, including current affairs, politics, lifestyle, culture, and general interest content, providing diverse linguistic patterns and writing styles. The dataset has been compiled to support research and development in natural language processing (NLP) for low-resource languages, particularly Gujarati. It can be used for tasks such as text classification, sentiment analysis, summarization, and language modeling. The data may include variations in tone, formality, and structure, reflecting both journalistic and informal writing. Basic preprocessing may have been applied, such as text cleaning and normalization, though users are encouraged to perform task-specific preprocessing as needed.
Licensing
Creative Commons Attribution Non Commercial 4.0 International (CC-BY-NC-4.0)
https://spdx.org/licenses/CC-BY-NC-4.0.htmlRestrictions/Special Constraints
This dataset is intended for research and educational purposes only. Users must ensure compliance with applicable laws and respect the rights of original content creators. Redistribution or commercial use may be subject to additional permissions.
Forbidden Usage
Users must not attempt to identify individuals or original authors from the dataset. The dataset must not be used to generate harmful, misleading, or unlawful content, including misinformation or abusive material. Any use that violates privacy, copyright, or applicable laws is strictly prohibited.
Intended Use
This dataset is intended for use in natural language processing tasks such as text classification, sentiment analysis, and language modeling for Gujarati.
Gujarati (ગુજરાતી) is an Indo-Aryan language of the Indo-European language family, belonging to the Indo-Aryan branch. It is the official language of the Indian state of Gujarat and is widely spoken in Dadra and Nagar Haveli, Daman and Diu, and among large diaspora communities in the United Kingdom, United States, East Africa, and Canada. According to Glottolog, it belongs to the Western Indo-Aryan group alongside Rajasthani and Sindhi. Gujarati holds a rich literary tradition spanning several centuries, with notable contributions in poetry, prose, and philosophical writing. Most speakers are bilingual in Hindi or English depending on their region and level of education.
Gujarati Script
અ, આ, ઇ, ઈ, ઉ, ઊ, ઋ, એ, ઐ, ઓ, ઔ, અં, અઃ, ક, ખ, ગ, ઘ, ઙ, ચ, છ, જ, ઝ, ઞ, ટ, ઠ, ડ, ઢ, ણ, ત, થ, દ, ધ, ન, પ, ફ, બ, ભ, મ, ય, ર, લ, વ, શ, ષ, સ, હ, ળ, ક્ષ, જ્ઞ, ં, ઃ, ઁ
The dataset is organized by author and source, each containing domain-specific sub-collections:
Gujarati News & Blogs Corpus/
│
├── Bakul Shah/
│ └── Literature Article Blog/
│
├── Capt Narendra/
│ ├── Article/
│ ├── Travel & Article Blog/
│ └── Travel & Literature Blog/
│
├── Saryu Parikh/
│ ├── Literary and Reflective Article Blog/
│ └── Literature and Personal Blog/
│
├── Suresh Jani/
│ └── Reflective Literature & Article Blog/
│
├── Vikas Nayak/
│ ├── Article and Literature Blog/
│ └── Inspirational Article Blog/
│
└── Webgujari/
└── Literature and Culture Magazine/
| Field | Details |
|---|---|
| Dataset Name | Gujarati News & Blogs Corpus |
| Language | Gujarati (ગુજરાતી) |
| Language Family | Indo-European — Indo-Aryan Branch |
| Script | Gujarati Script (Unicode) |
| Number of Authors | 5 Authors + 1 Website Source |
| Number of Domains | 9 |
| File Format | Plain Text (.txt) |
| Annotation | Unannotated — raw natural text |
[##]-Gujarati [Domain] Collection.txt