Release Date: 8/10/2026
Format: WAV, TSV
Size: 2.87 GB
Share
The Noakhali Text-to-Speech (TTS) Dataset is a high-quality speech corpus designed to support the development of speech technologies for the Noakhali dialect of Bangla. It contains professionally recorded speech paired with accurate text transcriptions, making it suitable for supervised training and evaluation of text-to-speech systems. The recordings are captured in a controlled environment to ensure consistent audio quality and natural speech. The dataset consists of literary content that reflects the phonetic and linguistic characteristics of the Noakhali dialect. It is intended for research in speech synthesis, pronunciation modeling, language modeling, and other speech and language processing tasks. By providing a structured and well-aligned speech corpus, the dataset also contributes to the preservation and digital representation of this regional language variety and supports the advancement of AI technologies for low-resource languages.
Pricing details
Your purchase is the license to the raw data. Once purchased, you're responsible for storing and using this dataset.
Your purchase includes a license to the data, paid directly to the dataset vendor, and a 5% ($50.00) platform fee paid to MDC.
Restrictions/Special Constraints
This dataset is intended exclusively for lawful research, educational, and speech technology development in compliance with applicable licenses and regulations. Any use that infringes intellectual property rights, privacy, or applicable laws is strictly prohibited.
Forbidden Usage
This dataset must not be used to generate deceptive, fraudulent, or harmful synthetic speech, including impersonation or misinformation.
Intended Use
This dataset is intended for research, education, and the development of speech technologies, including text-to-speech (TTS), speech synthesis, and language modeling for the Noakhali dialect. It also supports linguistic research and the preservation of regional language resources through responsible and ethical use.
Noakhali is a regional variety of Bangla (Bengali) primarily spoken in the Noakhali region of Bangladesh and by diaspora communities. It possesses distinctive phonetic, lexical, and grammatical characteristics that differentiate it from Standard Bangla, making it valuable for speech technology, linguistic research, and the preservation of regional language varieties.
Bengali (Bangla) Script: অ, আ, ই, ঈ, উ, ঊ, ঋ, এ, ঐ, ও, ঔ, ক, খ, গ, ঘ, ঙ, চ, ছ, জ, ঝ, ঞ, ট, ঠ, ড, ঢ, ণ, ত, থ, দ, ধ, ন, প, ফ, ব, ভ, ম, য, র, ল, শ, ষ, স, হ, ড়, ঢ়, য়, ং, ঃ, ঁ।
| Field | Value |
|---|---|
| Speaker ID | Spk-01 |
| Speaker Age | 21 |
| Speaker Region | Noakhali, Bangladesh |
| Speaker Gender | Male |
| Field | Value |
|---|---|
| Total Recordings | 6740 |
| Total Duration | 10 Hours 59 Minutes 21 Seconds |
| Field | Value |
|---|---|
| File Type | WAV (Studio Quality) |
| Bit Depth | 16-bit PCM |
| Sampling Rate | 48 kHz |
| Channels | Mono |
| Background Noise | Low (Studio Quality) |
| Signal-to-Noise Ratio (SNR) | High |
| Reverberation Time (RT60) | Estimated ~0.18–0.25 seconds |
| # | Domain | Clips |
|---|---|---|
| 1 | Literature | 6740 |
| Total Clips | 6740 |
| Field | Value |
|---|---|
| Total Clips | 6740 |
| Total Files (Including TCV Files) | 6740 |