Release Date: 9/29/2026
Format: JSONL, CSV, MD
Size: 333.79 KB
Created by RohingyaLanguage.org (https://rohingyalanguage.org/), this multilingual lexical dataset connects English dictionary headwords and phrases with Rohingyalish (Latin-script Rohingya) and Hanifi Rohingya script. The full release contains 15,926 rows in four UTF-8 JSONL shards, derived from 6,510 source dictionary entries. Each row includes id, english, rohingyalish, hanifi, part_of_speech, source_url, and validation. Hanifi forms were generated by the project’s rule-based converter. Published rows passed an exact normalized Rohingyalish → Hanifi → Rohingyalish round-trip check; duplicates, converter warnings, and non-exact round trips were excluded. This validation establishes converter consistency, not independent human linguistic verification. Explore the source dictionary at https://rohingyalanguage.org/tools/dictionary/ and the script converter at https://rohingyalanguage.org/tools/script-converter/. The full dataset is also hosted on Hugging Face; the related Zenodo record contains a separate 500-entry demonstration sample.
Licensing
Creative Commons Attribution 4.0 International (CC-BY-4.0)
https://spdx.org/licenses/CC-BY-4.0.htmlRestrictions/Special Constraints
No additional restrictions beyond Creative Commons Attribution 4.0 International (CC BY 4.0). Credit RohingyaLanguage.org, provide a link to the source dataset and license, and indicate if changes were made.
Forbidden Usage
No additional forbidden usages are imposed beyond the terms of CC BY 4.0 and applicable law. Commercial use, adaptation, and redistribution are permitted subject to the license.
Ethical Review
The source documentation does not report a formal ethics review or participant recruitment process. This release consists of dictionary headwords, translations, part-of-speech/origin annotations, and source URLs. Hanifi forms are generated by a rule-based converter. Automated round-trip validation was performed, but this is not an independent linguistic or ethical review. No contributor-consent or compensation claim is made in this submission.
Intended Use
Rohingya NLP and low-resource language research; Hanifi/Rohingyalish transliteration; dictionary and educational applications; search and text normalization; OCR post-processing experiments; and language-model evaluation or data preparation. Human linguistic review is recommended before relying on generated Hanifi forms in deployed applications.
Created and maintained by RohingyaLanguage.org, a source of Rohingya language tools and learning resources. Full source dataset: https://huggingface.co/datasets/rohingyalanguage/rohingya-hanifi-rohingyalish-english
Source export: https://github.com/abahziz0/rohingya-language/tree/hf-dataset-export/huggingface/rohingya-hanifi-rohingyalish-english
15,926 rows across four UTF-8 JSONL shards in a single train split. Fields: id, english, rohingyalish, hanifi, part_of_speech, source_url, validation. The source dictionary was flattened into one English–Rohingyalish translation per row. Hanifi text was generated using the project’s rule-based converter. Rows with unknown/unmapped characters, converter warnings, duplicate records, or non-exact normalized round trips were excluded. The validation field is exact_roundtrip. Source files are preserved without changes in this mirror; the archive includes provenance documentation.
Converter consistency does not establish linguistic accuracy or independent human review. Hanifi tone-marking conventions may vary among writers. This is a lexical resource, not a representative corpus of natural discourse, a word-frequency list, or a complete machine-translation benchmark.
The optional preview is the 500-entry CSV sample published at https://zenodo.org/records/22886278 (DOI: 10.5281/zenodo.22886278, version 1.0.0). That DOI identifies only the sample, not the full dataset. Its README documents deterministic sampling across 6,408 distinct retained English headwords, with one translation per headword.
Credit RohingyaLanguage.org, link to the source dataset and https://creativecommons.org/licenses/by/4.0/, and indicate changes as required by CC BY 4.0.
Compensated · similar languages
Versioned software-engineering corpus with code, configuration, workflows, and AI-tooling artifacts across frontend, backend, cloud, and infrastructure.
50 hours of premium, rights-cleared English-language scripted film content, with English subtitles and accompanying metadata, available for AI training