License:
CC-BY-NC-SA-4.0
Steward:
CommunityDataset ID:
cmsja4wtp01t3mi07ug56oto9
Release Date: 8/7/2026
Format: WAV, TXT
Size: 4.99 GB
Share
The Ladino Database Project, set up in 2008 by the Sephardic Center of Istanbul (SKAD) to document the spoken Judeo-Spanish (Ladino) of the last native speakers in Istanbul, produced 68 recorded interviews conducted with a structured five-section questionnaire by trained Ladino-speaking interviewers. This release contains a subset of that database: 13 of the 68 interviews, each with a full-length recording and its plain-text transcription (one pair per speaker). For 9 of the 13 speakers, the audio is also provided as short clips individually aligned to their transcripts, so the release supports both long-form speech processing and segment-level training and evaluation of speech recognition models. The original corpus was created by SKAD; digitization, segmentation, packaging and publication of this subset were carried out by Col·lectivaT in collaboration with Karen Gerson Sarhon of SKAD.
Licensing
Creative Commons Attribution Non Commercial Share Alike 4.0 International (CC-BY-NC-SA-4.0)
Restrictions/Special Constraints
Research and non-commercial use only. Attribution must be given to the original creators: the Sephardic Center of Istanbul and Karen Gerson Sarhon.
Forbidden Usage
Training text-to-speech models which replicates voices of original participants is strictly forbidden. Commercial use without the permission of the Sephardic Center of Istanbul is not permitted.
Ethical Review
Interviews were conducted by trained SKAD staff with adult native speakers who consented to participate in a language documentation project conducted by the Sephardic Center of Istanbul. The interviews were recorded and transcribed by SKAD in the course of its own documentation project; Col·lectivaT did not carry out the interviews or collect the data.
Intended Use
Speech corpus for automatic speech recognition training and evaluation in Ladino; linguistic research and documentation of an endangered language; sociolinguistic research on language shift and native-speaker fluency patterns.
The original corpus is described in:
Sarhon, Karen Gerson. 2011. "Ladino in Turkey: The Situation Today as Reflected by the Ladino Database Project". European Judaism 44(1): 62-71.
Interviews were conducted by a team of trained Ladino-speaking interviewers: Karen Sarhon (coordinator), Feride Petilon, Dora Niyego, Seli Gaon, Coya Delevi, Meri Schild and Anet Pase. Interviewees were native speakers of Ladino selected without restrictions on age or sex. The five-section questionnaire included questions such as "In what language do you count?" and "In what language do you dream?", questions about the interviewees' past and present, hypothetical/future questions, and free-association prompts eliciting proverbs, songs, traditions, superstitions and stories.
The original database comprises 68 recorded interviews of 30 minutes to 1.5 hours each, transcribed into Word files; the project was intended for 100 interviews but was completed at 68 due to funding constraints. The segmented version in this release contains 5,741 clips, 3,963 of them with aligned transcripts.
The project remains one of the only systematic spoken-language documentation efforts for Judeo-Spanish native speakers.
The digitization, segmentation and packaging of this release, carried out by Col·lectivaT in collaboration with SKAD, are described in:
Öktem, Alp, Rodolfo Zevallos, Yasmin Moslem, Özgür Güneş Öztürk, and Karen Gerson Şarhon. 2022. "Preparing an endangered language for the digital age: The Case of Judeo-Spanish". In Proceedings of the Workshop on Resources and Technologies for Indigenous, Endangered and Lesser-resourced Languages in Eurasia within the 13th Language Resources and Evaluation Conference, pages 105–110, Marseille, France. European Language Resources Association.
This subset contains 13 full-length recordings summing to about 12.9 hours of speech, each paired with a full transcription, plus 5,741 segmented clips (3,963 of them aligned to transcripts) totalling about 3.7 hours.
For other Ladino datasets from related work, see the Ladino Data Hub.
The original Ladino Speech Dataset was created by the Sephardic Center of Istanbul (SKAD). Col·lectivaT's role here was limited to digitization, packaging, and publication, carried out with the permission of Karen Sarhon. For access to original dataset, please contact Sephardic Centre of Istanbul.
Packaging and publication of this dataset was done as part of project "Judeo-Spanish: Connecting the two ends of the Mediterranean" carried out by Col·lectivaT and Sephardic Center of Istanbul within the framework of the "Grant Scheme for Common Cultural Heritage: Preservation and Dialogue between Turkey and the EU–II (CCH-II)" implemented by the Ministry of Culture and Tourism of the Republic of Turkey with the financial support of the European Union. The content of this repository is the sole responsibility of Col·lectivaT and does not necessarily reflect the views of the European Union.