Release Date: 9/2/2026
Format: TSV
Size: 2.93 MB
Share
This dataset contains 49,448 sentences of colloquial US English translated into Hausa and thoroughly reviewed by native speakers. Sourced from informal user communication, product reviews, and social interactions, the corpus captures authentic, everyday phrasing and localized tone. It is specifically designed to train machine translation engines, chatbots, and generative models to handle casual, non-formal voices effectively.
Pricing details
Your purchase is the license to the raw data. Once purchased, you're responsible for storing and using this dataset.
Your purchase includes a license to the data, paid directly to the dataset vendor, and a 5% ($79.75) platform fee paid to MDC.
Licensing
TAUS Data Use Agreement
https://www.taus.net/legal-policies/data-use-agreement-taus-on-mozilla-data-collectiveRestrictions/Special Constraints
No Data Extraction: Design, export, or operate any service or AI Model built on the Dataset in a manner intended to allow end-users or third parties to extract, reverse-engineer, or reconstruct the raw underlying translation pairs of the Dataset.
Forbidden Usage
No Resale or Sublicensing of Dataset: Sell, lease, rent, sublicense, assign, host, or redistribute the raw or modified Dataset (in whole or in part, including source-target text pairs) to any third party as a standalone data asset or database product. No Circumvention: Circumvent, remove, or alter any digital rights management or data structure associated with the Dataset.
Structure: Tab-separated parallel text pairs containing the source segment and target translation.
The source text was sampled from the TAUS Data Cloud—a repository accumulated over two decades from industry contributors, translation workflows, and curated corpora—specifically focusing on our colloquial Matching Data corpus. It consists of curated user-generated content (UGC) and informal business communications across multiple domains, including:
Product user reviews and blog comments
Social media interactions and chat logs
Everyday informal business small talk
While the English source sentences often capture colloquial, real-world usage that appears across various web sources and public crawls, the human target translations were created for and are owned by TAUS. These parallel paired translations do not exist elsewhere online.
Regarding third-party web matches or overlapping short phrases (such as subtitle snippets or dictionary examples), TAUS operates under a long-standing legal framework developed in consultation with tech and data specialist lawyer Wouter Seinen (Who Owns My Language Data White Paper, Section 2.2):
Originality Thresholds: Under both EU law (European Court of Justice precedent requiring the "author's own intellectual creation") and US copyright law (requiring more than a de minimis quantum of creativity), short, commonplace, or functional phrases do not meet the bar for copyright protection.
Segment Isolation: When documents are ingested into translation databases, they are broken into isolated sentence segments. Unless an individual short segment retains the distinct "signature of the author" (e.g., highly distinctive artistic prose or song lyrics), the isolated string itself falls outside copyright protection.
Functional & Colloquial Phrases: Sentences comprising everyday speech, common idioms, or standard dialogue lines generally do not qualify as copyrightable works in isolation.
Because our repository operates at a segment level rather than a document level, URL-to-sentence MANIFEST tracking is not maintained. The legal and operational clearance of these segments relies on these established legal thresholds for short, non-original text fragments.
To capture natural, spoken-sounding language for conversational AI and chatbots, native-speaker translators were given the following criteria:
Tone & Style: Prioritize a friendly, casual, and natural spoken tone over rigid literal accuracy. Local idioms and informal phrasing should be adapted naturally into the target language.
Handling Difficult Colloquialisms: When translating highly creative or figurative phrasing (e.g., idiom-dense phrases or slang), translators were instructed to find an equivalent natural colloquial expression in the target language rather than translating word-for-word.
Review & Verification: All translations underwent a mandatory two-tier workflow where a second native speaker reviewed and validated the segments to correct over-formalization or misinterpretations of informal source text.