License:
CC-BY-NC-4.0
Steward:
Information and Language Processing Research LabDataset ID:
cmssxpm5q02ylmf07ro4vzrls
Release Date: 8/14/2026
Format: WAV, TSV
Size: 477.54 MB
Share
Nwāchā Munā is a 5.39-hour manually transcribed Devanagari speech corpus for Nepal Bhasha (Newar), an endangered and digitally under-resourced language primarily spoken in the Kathmandu Valley. The corpus was developed to address the scarcity of publicly available annotated speech resources for Nepal Bhasha and to support research in automatic speech recognition (ASR) and low-resource speech technology. The corpus contains 5,727 utterances collected through a combination of original field recordings and existing web-based audio resources. The field recordings were contributed by 18 native Nepal Bhasha speakers from different regions and were recorded using smartphone microphones in open environments. An additional portion of the corpus was derived from web-sourced audio and transcribed/transliterated into Devanagari. The audio was standardized to 16 kHz, mono-channel WAV format. The dataset is accompanied by a proximal-transfer ASR benchmark that investigates whether transfer from the geographically and linguistically related Nepali language can provide competitive performance in an ultra-low-resource Nepal Bhasha setting. The benchmark compares Nepali Conformer-based transfer with multilingual pretraining approaches and evaluates the effect of data augmentation. Nwāchā Munā is intended to support research in automatic speech recognition, speech technology, low-resource language processing, linguistic research, and language technology development for Nepal Bhasha. The corpus is being made openly available to improve the discoverability and accessibility of Nepal Bhasha speech resources and to encourage further research and development for the language.
Licensing
Creative Commons Attribution Non Commercial 4.0 International (CC-BY-NC-4.0)
https://spdx.org/licenses/CC-BY-NC-4.0.htmlRestrictions/Special Constraints
This dataset is intended for research, educational, and scientific purposes. Users must comply with the terms of the CC BY-NC-SA 4.0 license, including attribution, non-commercial use, and sharing adaptations under the same license.
Forbidden Usage
Attempting to identify, re-identify, or infer the identity of individual speakers from the dataset. Using the dataset for speaker identification, profiling, surveillance, or other purposes intended to identify or track individual speakers. Attempting to clone, impersonate, or reproduce the voice of an individual speaker for deceptive or harmful purposes. Using the dataset to create or deploy systems that impersonate individual speakers without their consent. Attempting to infer sensitive personal information about individual speakers from their recordings.
Ethical Review
Participants were informed about the purpose of the data collection and the intended use of their speech recordings for Nepal Bhasha speech and language technology research. Consent was obtained from participants for the recording and use of their speech data. The data is intended to be used responsibly for research and language technology development, with measures in place to respect participant privacy and confidentiality.
Intended Use
This dataset is intended for research and development in Nepal Bhasha speech and language technology, including automatic speech recognition, low-resource speech processing, linguistic research, language preservation, and the development and evaluation of speech-based applications and models.
Nwāchā Munā is a 5.39-hour manually transcribed Devanagari speech corpus for Nepal Bhasha (Newar), an endangered and digitally under-resourced language primarily spoken in the Kathmandu Valley.
The corpus was developed to address the limited availability of publicly accessible, annotated speech resources for Nepal Bhasha and to support research and development in automatic speech recognition (ASR) and low-resource speech technology.
The corpus contains:
Total duration: 5.39 hours
Utterances: 5,727
Word tokens: 27,644
Unique words: 8,599
Speakers: 18 native Nepal Bhasha speakers
Script: Devanagari
Audio format: WAV
Sampling rate: 16 kHz
Channels: Mono
Transcription: Manual transcription
The corpus includes speech collected from native Nepal Bhasha speakers across different regions, together with speech obtained from publicly available web-based sources.
The field-recorded portion includes speakers from Banepa, Dhulikhel, Panauti, and Patan, with recordings covering different speaking contexts and community backgrounds.
The speech recordings were standardized to 16 kHz, mono-channel WAV format to provide a consistent format for speech recognition research.
The speech was manually transcribed in Devanagari script, providing textual resources suitable for training and evaluating Nepal Bhasha ASR systems.
Nwāchā Munā is accompanied by a proximal-transfer ASR benchmark designed to investigate the use of related-language transfer for low-resource Nepal Bhasha speech recognition.
The benchmark explores whether knowledge transferred from the geographically and linguistically related Nepali language can improve ASR performance in an ultra-low-resource Nepal Bhasha setting.
The benchmark includes experiments involving Nepali Conformer-based transfer, multilingual pretraining, data augmentation, and other ASR configurations.
Nwāchā Munā is intended to support:
Automatic speech recognition
Low-resource speech processing
Speech and language technology research
Linguistic research
Language documentation and preservation
Development and evaluation of Nepal Bhasha speech-based applications
By making the corpus publicly accessible, Nwāchā Munā aims to improve the availability and discoverability of Nepal Bhasha speech resources and encourage further research and development in Nepal Bhasha language technology.
For detailed information about the corpus, data collection, transcription methodology, preprocessing, experimental setup, and ASR benchmark, see the published paper:
Nwāchā Munā: A Devanagari Speech Corpus and Proximal Transfer Benchmark for Nepal Bhasha ASR
CHiPSAL 2026 / LREC 2026 Publication
The accompanying GitHub repository contains the ASR training scripts, benchmarking code, notebooks, and related resources used for the experiments:
If you use the Nwāchā Munā corpus or its associated benchmarks, please cite the accompanying publication.