Task: N/A
Release Date: 9/15/2026
Format: WAV, MP4, EAF, MD, CSV
Size: 4.46 GB
This collection contains eighteen recordings from January 2019, constituting the largest portion of the Amahuaca corpus. It brings together discourse genres that fall outside biography and technical description, including traditional narratives such as stories, legends, riddles, and the narrative of oma; three accounts of the history of the Amahuaca people, including one with commentary in Spanish; free conversations and unscripted recordings of spontaneous speech; and metadiscursive materials addressing the Amahuaca language itself, encounters with missionaries, an anecdote, and a commented reading session of a book. The spontaneous speech materials are particularly valuable for acoustic modeling. The recordings were made in Atalaya and the Native Communities of San Juan and San Martín, Ucayali, Peru. Documentation was carried out by Pilar Valenzuela, Candy Angulo, and Roberto Zariquiey Biondi.
Licensing
Licencia Chana 2.0 — Licence for Peruvian Indigenous Language Documentation Collections
https://github.com/rzariquiey/licencia-chanaRestrictions/Special Constraints
ACCESS IS GRANTED ON REQUEST. Requesters must identify themselves and state their institutional affiliation. The donor reviews and approves or declines each request individually. This dataset is released under the Licencia Chana 2.0, not under a Creative Commons licence. Download is direct and unrestricted, but use is subject to the following terms. PERMITTED WITHOUT FURTHER AUTHORISATION - Academic research and publication, with attribution. - Teaching and educational use, with attribution. - Language revitalisation work by the source communities and by organisations working with them. - Evaluating and benchmarking existing speech or language models, with attribution. - Non-production academic machine-learning experimentation, provided that (a) the resulting model is not deployed in production, (b) model weights are not released, and (c) any results published cite this collection. REQUIRES PRIOR WRITTEN AUTHORISATION FROM THE DONOR - Training models intended for production or public release. - Release or distribution of model weights derived from this material. - Any commercial use. - Redistribution of the dataset, in whole or in part, on any other platform. ATTRIBUTION Cite as: Zariquiey, Roberto (donor). Amahuaca: Narratives and Conversations. Mozilla Data Collective. RIGHT OF WITHDRAWAL Speakers retain the right to withdraw their recordings from this collection at any time. This is the principal reason a Creative Commons licence was not used: CC licences are irrevocable and cannot accommodate this commitment. Contact for authorisations: [email protected]
Forbidden Usage
- You agree not to attempt to determine the identity of the pseudonymised speakers in this dataset, nor to link the speaker codes (S01, S02...) to real individuals. - Any attempt to clone, synthesise or imitate the voices of the speakers in this dataset is forbidden. - Training models intended for production deployment or public release is forbidden without prior written authorisation from the donor. This includes releasing model weights trained wholly or partly on this material. - Commercial use of any kind is forbidden without prior written authorisation. - Redistribution of this dataset, in whole or in part, on any other platform or in any other repository is forbidden without prior written authorisation. - Use of this material in ways that misrepresent, decontextualise or commercially exploit the cultural knowledge it contains is forbidden. Note on machine learning: this is not a blanket prohibition on ML research. Evaluating and benchmarking existing models on this data is permitted, and so is non-production academic experimentation, provided the model is not deployed, the weights are not released, and the results cite the collection. What requires authorisation is production training and weight release. The reason is that the consent forms signed before approximately 2020 do not mention AI training and cannot reasonably be read as covering it; the communities have not yet been consulted on this point. Contact for authorisations: [email protected]
Intended Use
Amahuaca traditional narrative, three independent versions of the history of the Amahuaca people, free unscripted conversation, and metalinguistic reflection by speakers on their own language. Intended for discourse research, for acoustic modelling from unscripted multi-party speech, and for the public record of language shift.