Release Date: 9/14/2026
Format: JS, TS, HTML, CSS, JSON, PY, PS1, SH, YAML, ENV
Size: 419.55 KB
Software Engineering & AI Tooling Corpus is a curated, versioned collection of software-development artifacts spanning frontend engineering, backend engineering, full-stack workflows, API foundations, authentication and security, cloud deployment, storage and file services, AI model integration, reliability and infrastructure, and application bootstrap. The dataset contains implementation files, configuration artifacts, code iterations, and development-stage outputs across JavaScript, TypeScript, HTML, CSS, JSON, Python, PowerShell, shell, YAML, ENV, and related software-engineering formats. The corpus preserves iterative development states and version progression so that changes in implementation, debugging, refactoring, configuration, feature development, and system integration can be studied over time. Files are organized by engineering domain and technology type to support structured analysis of software evolution, code transformation, implementation decisions, and multi-file development workflows. The dataset is intended for software-engineering research and commercial AI/ML applications including code generation, code completion, debugging, refactoring, code understanding, developer tooling, software-engineering agents, configuration reasoning, repository-level reasoning, and evaluation of AI-assisted programming systems. Its versioned structure also supports research into development trajectories, code-change modeling, automated software maintenance, and longitudinal analysis of software-engineering processes.
Pricing details
Your purchase is the license to the raw data. Once purchased, you're responsible for storing and using this dataset.
Your purchase includes a license to the data, paid directly to the dataset vendor, and a 5% ($625.00) platform fee paid to MDC.
Licensing
Mr. Zay LLC Software Engineering & AI Tooling Non-Exclusive License
https://drive.google.com/file/d/1KjKnGyUjVw5u5w4-oTqo56GthIaSXtUA/view?usp=sharingRestrictions/Special Constraints
Use of this dataset is permitted only under the applicable Mozilla Data Collective agreement and the dataset license. Commercial AI/ML training, evaluation, code understanding, developer-tooling research, software-engineering research, and related applications are permitted. The raw dataset, substantially equivalent copies, or extracted corpus may not be resold, republished, sublicensed, or redistributed as a standalone dataset without written permission from the rights holder. Users must respect applicable intellectual-property rights, notices, privacy requirements, and any third-party restrictions identified within the dataset.
Forbidden Usage
Users may not redistribute or resell the raw dataset or a substantially equivalent derivative corpus without authorization. The dataset may not be used to identify, profile, or contact individuals; extract or exploit credentials, secrets, access tokens, or private information; gain unauthorized access to systems; develop malware, credential-theft systems, destructive exploits, or other unlawful capabilities; misrepresent authorship or ownership; or use the dataset in violation of applicable law or third-party rights.
Ethical Review
No formal institutional human-subjects ethics review was conducted because this corpus primarily consists of software source code and configuration artifacts rather than data collected through a human-subject research study. Release review focuses on intellectual-property rights, privacy, confidentiality, and security. Credentials, private keys, access tokens, unnecessary personal information, and materials not authorized for distribution must be excluded from the released dataset. Contributor-created materials should only be distributed where the submitter has sufficient rights or authorization to do so.
Intended Use
This dataset is intended for training, fine-tuning, benchmarking, and evaluating systems for code generation, code completion, debugging, refactoring, code understanding, software-engineering agents, developer tooling, configuration reasoning, multi-file workflow understanding, repository-level reasoning, and longitudinal analysis of software-development processes. It may also support research into AI-assisted programming and automated software engineering.
The dataset was assembled from a broader software-development corpus and organized into ten primary engineering domains: Frontend Engineering, Backend Engineering, Full Stack Workflows, API Foundations, Authentication & Security, Cloud Deployment, Storage & File Services, AI Model Integration, Reliability & Infrastructure, and Application Bootstrap. Verified format groupings include JavaScript, TypeScript, HTML, CSS, JSON, Python, PowerShell, shell, YAML, and ENV/configuration artifacts. The corpus contains iterative implementation states that may be useful for longitudinal development analysis, change modeling, debugging research, code transformation, and software-engineering trajectory modeling. Version identifiers should generally be interpreted as development lineage rather than unrelated standalone samples.