411 Downloads view our datasets

MirasAI

  • LinkedIn

About MirasAI

MirasAI LLC

About Us

MirasAI LLC is an NLP research company dedicated to advancing language technology for underserved South Asian languages: Urdu, Punjabi, Saraiki, Sindhi, Pashto, Balochi, Burushaski, and others. Based in Indiana, we develop high-quality datasets and AI tools that serve researchers, developers, and communities across South Asia.

Our Mission

To democratize language technology for low-resource South Asian languages, enabling equitable AI development and ensuring these communities benefit from the language resources they create.

What We Do

Research & Data: High-quality NLP datasets across speech, text, and multimodal domains (ASR, TTS, speech emotion recognition, linguistic annotation).

Technology Services: Custom AI solutions for organizations working in South Asia.

Recent Work

  • Burushaski–English Speech Translation Corpus (15 hours, 14,970 utterances)

  • Karachi Colloquial Urdu Speech Dataset (50-hour collection)

  • Punjabi Multimodal Dataset (113 prompts, 10+ hours)

  • Hazargi Multimodal Dataset (181 video segments)

Why We Share on Mozilla Data Collective

We publish on MDC because it aligns with our values: open access, community accountability, and ethical data governance. Publishing here ensures researchers worldwide can access language resources equitably, elevates South Asian language visibility globally, and builds critical infrastructure for underserved research communities.

All our datasets follow rigorous ethical standards including community consent, fair contributor compensation, and transparent documentation.

Contact: mirasai.net | Meesum Alam, Director

Datasets

Associated Press Pakistan 21.4 Million Pashto CorpusCC-BY-NC-SA-4.0pusNLPTXT38.81 MB
Associated Press Pakistan 74.4 Million Urdu CorpusCC-BY-NC-SA-4.0urdNLPTXT143.39 MB
Associated Press Pakistan 1.8 Million Chinese CorpusCC-BY-NC-SA-4.0zhoNLPTXT2.76 MB
Associated Press Pakistan 10.4 Million Saraiki CorpusCC-BY-NC-SA-4.0skrNLPTXT19.51 MB
Associated Press Pakistan 16.4 Million Sindhi CorpusCC-BY-NC-SA-4.0sndNLPTXT23.23 MB
Associated Press Pakistan 16.9 Million Balochi CorpusCC-BY-NC-SA-4.0balNLPTXT22.85 MB
Associated Press Pakistan 5.1 Million Arabic CorpusCC-BY-NC-SA-4.0arbNLPTXT9.60 MB
Associated Press Pakistan 78.1 Million English CorpusCC-BY-NC-SA-4.0engNLPTXT178.01 MB
Balti Literature corpusCC-BY-NC-SA-4.0bftNLPTXT, DOCX292.67 KB
Bangladesh Traffic Signs DatasetCC-BY-NC-4.0ben, engCVJSON, JPEG1.35 GB
Chittagonian (চাটগাঁইয়া, saṭgãia) Text CorpusCC-BY-NC-SA-4.0ctgNLPTXT6.59 MB
Dameli Literature CorpusCC-BY-NC-4.0dmlNLPTXT, DOCX9.83 MB
Gawar-Bati Literature CorpusCC-BY-NC-SA-4.0gwtNLPTXT, DOCX876.83 KB
Gujarati News and Blogs CorpusCC-BY-NC-4.0gujNLPTXT18.93 MB
Hindi Literature & News CorpusCC-BY-NC-SA-4.0hinNLPTXT19.52 MB
India Traffic Signs DatasetCC-BY-ND-4.0NACVJPEG, JSON190.44 MB
Kannada Text CorpusCC-BY-NC-SA-4.0kanNLPTXT2.61 MB
Kannada Time Aligned Speech CorpusCC-BY-NC-SA-4.0kanASROGG, SRT355.77 MB
Marathi Blog & Literature CorpusCC-BY-NC-4.0marNLPTXT4.20 MB
Multispeaker Hindi ASR DatasetCC-BY-NC-SA-4.0hinASROGG , SRT63.03 MB
Noakhalian (নোয়াখাইল্লা) Text CorpusCC-BY-NC-4.0oakNLPTXT, DOCX3.61 MB
Ormuri Bilingual Dictionary by Rozi Khan BurkiCC-BY-ND-4.0oruTTSJPG, PDF, DOCX15.68 MB
Ormuri Literature corpus by Rozi BurkiCC-BY-NC-SA-4.0oruNLPPDF, DOCX18.92 MB
Pakistan Traffic Signs DatasetCC-BY-NC-4.0urd, engCVJPEG, JSON2.64 GB
Palula Literature CorpusCC-BY-NC-4.0phlTTSDOCX, TXT, PDF34.05 MB
Punjabi 10 Hours TTS CC-BY-NC-SA-4.0pnbTTSWEBM, TSV481.96 MB
Rangpuri (অংপুরি Ôṅgpuri) Text CorpusCC-BY-NC-4.0rktNLPTXT3.68 MB
Rohingya Literature CorpusCC-BY-NC-4.0rhgNLPTXT, DOCX7.01 MB
Saraiki 10 Hours TTS DatasetCC-BY-NC-SA-4.0srkTTSWEBM, TSV584.44 MB
Sylheti Text Corpus by Haque PublishersCC-BY-NC-4.0sylNLPTXT3.53 MB
Tamil Literature CorpusCC-BY-NC-SA-4.0tamNLPTXT41.85 MB
Tamil Time Aligned Speech DatasetCC-BY-NC-SA-4.0tamASROGG, SRT37.11 MB
Telugu Text CorpusCC-BY-NC-SA-4.0telNLPTXT19.87 MB
Wakhi Literature CorpusCC-BY-ND-4.0wblNLPDOCX, TXT74.83 KB
🚧 Noakhali 10 Hours TTS (Male Speaker) 🚧MirasAI Data LicenseoakTTSWAV, TSV2.87 GB