License:
CC-BY-NC-SA-4.0
Steward:
CommunityDataset ID:
cmtr89ry305csny07rrpxqe9d
Release Date: 9/7/2026
Format: WEBM, TSV
Size: 554.21 MB
Share
This dataset contains speech data in Javanese language, especially Purworejo dialect, collected from native speakers. It consists of audio recordings and transcriptions covering various everyday topics, including family, education, food, work, community, culture, and daily activities.
Licensing
Creative Commons Attribution Non Commercial Share Alike 4.0 International (CC-BY-NC-SA-4.0)
https://spdx.org/licenses/CC-BY-NC-SA-4.0.htmlRestrictions/Special Constraints
The dataset is intended to be used in accordance with its license terms. Users must provide attribution to the dataset creator and must not remove source information. Use of the data should respect the rights and privacy of the speakers.
Forbidden Usage
This dataset must not be used to identify, expose, or disclose the personal identities of speakers. It must not be used to create systems that misleadingly imitate or impersonate the speaker's voices or identities. This dataset must not be used for commercial purposes, including selling, renting, licensing, or otherwise exploiting the data or voices in the dataset for commercial gain. Any use that violates the rights, privacy, or consent of the speakers is also prohibited.
Ethical Review
This dataset was created by writing texts in the Purworejo dialect of Javanese with code-mixing of Indonesian and English.The files were read and recorded by native speakers through the hosting platform https://sabre-2.onrender.com/ and https://mdc-dataset-toolbox-ifuhj.ondigitalocean.app/app/sabre . The collection of audio recordings was compiled into a comprehensive dataset.
Intended Use
This dataset is intended for research, education, and language technology development, with a focus on preserving and supporting regional languages.
The dataset contains Javanese used in everyday communication by native speakers, with a primary geographical focus on Purworejo, Central Java Province, Indonesia.
Created by the owner of the dataset, considered as a linguist and native speaker.
The dataset covers general topics such as everyday life, family, education, community, food, restaurants, markets, business, culture, environment, and personal experiences and opinions.
Approximately 10 hours
Audio file name, text
“Ing wekdal esuk punika wekdal ingkang sae sanget kangge ngapalaken hafalan Al-Qur’an saha nindakaken murojaah.”
“Amargi ing wayah esuk, hawa taksih seger lan pikiran isih resik, saengga saged langkung gampil kangge ngapalaken ayat-ayat suci.”
“Kajaba punika, ing wekdal punika durung kathah swara gangguan saking tangga teparo utawi kagiyatan sanesipun, dados saged langkung fokus lan tenang.”
“Nalika ngapalaken Al-Qur’an ing esuk, ati rumasa langkung ayem, pikiran langkung cetha, lan kalodhangan punika saged nuwuhaken semangat tumrap nglajengaken sinau lan ngibadah ing dinten punika.”
“Wontenipun kabersihan hawa lan ketenangan suwasana ndadosaken wekdal ésuk punika wekdal ingkang paling trep tumrap para penghafal Al-Qur’an.”
The dataset uses the latin alphabet with spelling that follows everyday language usage. Latin alphabet (A–Z), Arabic numerals (0–9).