Release Date: 9/2/2026
Format: TXT
Size: 9.60 MB
Share
The Associated Press Pakistan 5.1 Million Arabic Corpus is an Arabic-language news dataset comprising 41,963 articles (approximately 5.1 million words) published by the Associated Press of Pakistan (APP) on its Arabic-language section at app.com.pk, covering the period 2016–2026. The corpus offers a distinctive South Asian perspective in Arabic-language news text and was assembled by MirasAI LLC using an automated web crawler that parsed each article's HTML to extract the headline, body text, publication date, and source URL, followed by deduplication, boilerplate removal, and normalization to plain UTF-8 text. APP retains full and exclusive ownership of the data, with MirasAI serving solely as documentation and hosting facilitator; access is restricted and approval-based via the Mozilla Data Collective, for non-commercial research use only, with no sale, sublicensing, or commercial redistribution permitted, and APP may request takedown within a 48-hour compliance window. The dataset is intended for non-commercial NLP research, including Arabic language modeling and cross-lingual studies, and all users are required to credit the Associated Press of Pakistan as the sole and primary source.
Licensing
Creative Commons Attribution Non Commercial Share Alike 4.0 International (CC-BY-NC-SA-4.0)
https://spdx.org/licenses/CC-BY-NC-SA-4.0.htmlRestrictions/Special Constraints
Access is restricted and approval-based through the Mozilla Data Collective for non-commercial research only. The Associated Press of Pakistan retains full ownership, while MirasAI LLC acts solely as documentation and hosting facilitator. Sale, sublicensing, and commercial redistribution are prohibited. Users must credit the Associated Press of Pakistan as the primary source.
Forbidden Usage
Commercial use, sale, sublicensing, redistribution, and unauthorized derivative works are prohibited. Use is limited to approved non-commercial NLP research through the Mozilla Data Collective. Users must credit the Associated Press of Pakistan as the primary source and must not misrepresent MirasAI LLC as the data owner.
Ethical Review
Hosted with the permission of the Associated Press of Pakistan, which retains full ownership. MirasAI acts solely as a documentation and hosting facilitator. Access is restricted to approved, non-commercial research; sale, sublicensing, and redistribution are prohibited.
Intended Use
Intended for non-commercial NLP research, including Arabic language modeling and cross-lingual studies. All uses must credit the Associated Press of Pakistan as the sole and primary source.
The Associated Press Pakistan Arabic News Corpus is a collection of Arabic-language news articles published by the Associated Press of Pakistan, Pakistan’s national news agency. Established in 1947, the Associated Press of Pakistan is the country’s premier news service and publishes news across multiple languages, including English, Urdu, Chinese, Arabic, Sindhi, Saraiki, Pashto, Balochi, and Brahvi.
This dataset specifically contains Arabic-language text from the published Arabic-language archive of the Associated Press of Pakistan. It provides a distinctive South Asian perspective in Arabic news reporting, covering topics related to Pakistan, international affairs, diplomacy, politics, culture, and regional developments.
Arabic is a Central Semitic language and one of the world's most widely spoken languages, serving as the official language of over 20 countries across the Middle East and North Africa and as one of the six official languages of the United Nations. It is widely used in education, media, literature, religion, and everyday communication across the Arabic-speaking world. Arabic has a rich linguistic and literary heritage spanning over a thousand years and plays an important role in NLP research, including language modeling and cross-lingual studies.
| Field | Value |
|---|---|
| Articles | 41,963 |
| Word count | ~5.1 million |
| Coverage | 2016–2026 |
| Language | Arabic (arb) |
| Source | app.com.pk (Arabic section) |
Arabic.tar
└── Arabic-20260828T091135Z-1-001/
└── Arabic/
├── 2016.txt
├── 2017.txt
├── 2018.txt
├── 2019.txt
├── 2020.txt
├── 2021.txt
├── 2022.txt
├── 2023.txt
├── 2024.txt
├── 2025.txt
└── 2026.txt
إسلام آباد: 01 – يناير 2018م (وكالة الأنباء الباكستانية الرسمية).
أصيب ثمانية أشخاص بينهم رجال الأمن إثر الانفجارين المزدوجين بمدينة "تشامن" في إقليم بلوشستان بجنوب غرب البلاد اليوم، وأوضحت الشرطة الباكستانية بأن الانفجارين وقعا قرب نقطة التفتيش التابعة للشرطة ما أسفرا عن إصابة 8 أشخاص بينهم رجال الأمن تم نقلهم إلى مستشفى قريبة، من جانبها هرعت القوات إلى مكان الحادث وطوقتها لجمع الأدلة.
Articles were collected from the Associated Press of Pakistan's publicly published Arabic-language section on app.com.pk using an automated web crawler developed by MirasAI. Each article's HTML was parsed to extract the headline, body text, publication date, and source URL; articles were deduplicated, cleaned of boilerplate, and normalized to plain text in UTF-8. Crawling was rate-limited and conducted with the knowledge and permission of the Associated Press of Pakistan.
This corpus is uploaded and hosted with the express permission of the Associated Press of Pakistan. The Associated Press of Pakistan retains full and exclusive ownership of the data; MirasAI acts solely as documentation and hosting facilitator. Access on the Mozilla Data Collective is restricted and approval-based, for non-commercial research use only. No sale, sublicensing, or commercial redistribution is permitted.
Intended for non-commercial NLP research, including Arabic language modeling and cross-lingual studies. All uses must credit the Associated Press of Pakistan as the sole and primary source.