License:
CC-BY-NC-SA-4.0
Steward:
CommunityDataset ID:
cmu1ityrc0024o107b4n7rqo6
Release Date: 9/14/2026
Format: WEBM, TSV
Size: 414.61 MB
This dataset contains multilingual data reflecting the use of Javanese in the Klaten region, including instances of code-mixing with Indonesian, the speakers’ national language, and English, an international language. Generally, it employs ‘ngoko’ variants of Javanese speech levels. Language use in Klaten Regency is interesting to document due to its geographical location, situated between Yogyakarta and Surakarta, two major cultural centers on the island of Java.
Licensing
Creative Commons Attribution Non Commercial Share Alike 4.0 International (CC-BY-NC-SA-4.0)
Restrictions/Special Constraints
This dataset is not intended for commercial use. It is provided solely for research and educational purposes. Any use of the data must include proper attribution and citation of the source. All use of this dataset requires written permission through an access request that clearly states the intended purpose of use. The data may not be used for any purpose that could result in negative or harmful impacts. The owner shall not be held responsible for any improper or unlawful use of the data, including any use that violates applicable laws or regulations.
Forbidden Usage
Modification of this dataset without the owner's permission is prohibited. The dataset may not be used for criminal, unlawful, or unethical purposes. The dataset may not be used to clone specific voices, impersonate specific individuals, or train models designed to imitate specific individuals. The dataset may not be used to train chatbots or large language models (LLMs). Re-uploading or redistributing the original dataset is prohibited. Any attempt to identify the speakers in this dataset is prohibited.
Ethical Review
This dataset was created by writing texts in the Javanese dialect spoken in Klaten Regency, incorporating code-mixing with Indonesian and English. The files were read and recorded by a native speaker through the hosting platform https://mdc-dataset-toolbox-ifuhj.ondigitalocean.app/app/sabre. The collection of audio recordings was compiled into a comprehensive dataset.
Intended Use
This dataset is designed to document Javanese speech in Klaten Regency, Central Java Province, Indonesia. It is intended to support the preservation and dissemination of the Javanese language spoken in Klaten Regency for research, educational, and cultural purposes. Specifically, this dataset is intended for linguistic studies, including dialectology, comparative-historical linguistics, ethnolinguistics, sociolinguistics, and etymology particularly research concerning regional dialects and languages, with a specific focus on Javanese as spoken in Klaten Regency.
Javanese language of ‘ngoko’ variant speech level from Klaten dialect.
Native speakers of Javanese, aged between 20 and 30 years old,
The data consists of a collection of Javanese sentences and their corresponding spoken utterances in the form of recordings. The author is a Javanese speaker from the Central Klaten District, Klaten Regency. The dataset was compiled between May and August 2026.
This dataset consists of general domains such as culture, daily life activity, social interactions, personal experiences, social media usage, family, education, technology, etc.
5 hours
Public open access is permitted with proper attribution and citation of the dataset source.
Deshinta Amalia Putri. (2026). Speech Corpus of Javanese-Klaten Regency [Data set]. Mozilla Data Collective. URL [dataset].
Audio file name, text
“”Salah siji cerita legenda nang daerahku seng sering dicritakne naliko nang sekolah utawa nang lingkungan sosial yaiku bab bulus Jimbung.
“Covid-19 dadi pageblug paling nganyeli.”
“Saiki media sosial wes akeh ngangkat konten-konten gawean AI koyo video lan foto seng mirip karo tokoh-tokoh terkenal.“
“Na, tradisi Yaqowiyu iki wujud e yaiku festival seng kegiatan utamane yaiku nyebar puluan ewu apem.”
“Siji, ngandakne nak Klaten kasal soko tembung "kelathi" seng artine buah bibir utawa bahan rasan-rasan.”
Latin alphabet (A–Z), Arabic numerals (0–9).