AI Governance Documentation for Pashto, Dari and 22 More Afghan Languages
EU AI Act, California SB 53 and ISO/IEC 42001 evidence
Ariana Nexus prepares the Afghan-language evidence that AI governance files are usually missing: what your model or system was tested on in Pashto and Dari, how it performed, where it fails, who reviewed it and how it is monitored. Each document is written to the EU AI Act, California SB 53 or ISO/IEC 42001 requirement it supports.
See what the evidence file containsWhat is AI governance documentation for language coverage?
AI governance documentation for language coverage is the written evidence of how an AI model or system was built, tested, limited and monitored in a specific language. For Afghan languages it records the Pashto and Dari training and test data, evaluation results by language and dialect, known failure modes, human oversight arrangements and incident history, in the structure that the EU AI Act, California SB 53 and ISO/IEC 42001 ask for.
How AI governance documentation for Pashto and Dari works
From the records you already hold to one evidence file that three frameworks can read.
- Model cards and system cards
- Data sheets
- Evaluation logs
- Risk registers
- Public language claims
- Native-written Pashto and Dari tests
- Data review by variety and script
- Safety evaluation against English
- Drafting to the requirement
- Independent second review
- Ten controlled documents
- Requirement-to-evidence crosswalk
- Evaluator record
- JSON and CSV index
- Version history
- EU AI Act: AI Office, notified body, deployer
- California SB 53: the public, the Attorney General
- ISO/IEC 42001: certification body
- Your own counsel and auditors
Most AI governance files are silent on Pashto and Dari
Conformance is not coverage.
A model card says the model supports more than a hundred languages. The technical documentation reports accuracy on English benchmarks. The risk assessment lists multilingual performance as a single line. None of that tells a regulator, an auditor or a deployer what happens when a person writes to the system in Pashto or Dari.
That gap is now a documentation problem, not only a quality problem. California SB 53 requires frontier developers to publish the languages a model supports. The European Commission’s training-content template asks general-purpose AI providers to describe the linguistic characteristics of their data. The EU AI Act requires high-risk systems to be trained and tested on data that represents the people and settings they are used in, and to declare accuracy and known limitations. ISO/IEC 42001 asks for an assessment of the system’s impact on individuals and groups.
A management system can conform to a standard and still hold no evidence about an entire language community. Conformance is not coverage. This service supplies the coverage.
What we deliver: the Afghan-language evidence file
Ten controlled documents. Each one can be lifted into your technical documentation, framework, transparency report or management system without rewriting.
M1: Regional requirements
Where each module is filed
M2: Summaries and public statements
Language coverage statement
Safety evaluation summary
Instructions for use and transparency text
M3: Data
Data governance record
M4: Testing
Evaluation and test report
M5: People and operations
Language-specific risk register
Impact assessment input
Human oversight and escalation procedure
Monitoring and incident plan
Delivered as controlled Word and PDF documents with version history, plus a machine-readable index (JSON and CSV) that imports into your document control or GRC system.
How the evidence maps to the EU AI Act, SB 53 and ISO/IEC 42001
One body of evidence, three readers. The same Pashto and Dari test results answer a different question in each framework.
References are a navigation aid prepared on 19 September 2026 against the current legal texts. They are not legal advice. Your counsel decides which obligations apply to you.
When the obligations apply
Checked against the Official Journal and the California code on 19 September 2026.
- ISO/IEC 42001:2023 published. Certification is available now.In effect
- EU AI Act enters into force.In effect
- EU obligations for general-purpose AI model providers apply (Articles 53 and 55). Code of Practice and training-content template published in July 2025.In effect
- California SB 53 takes effect: frontier AI frameworks, transparency reports listing languages supported, incident reporting.In effect
- Regulation (EU) 2026/1744, the Digital Omnibus on AI, enters into force and moves the high-risk dates.In effect
- European Commission enforcement powers over general-purpose AI providers begin. Article 50 transparency obligations apply.In effect now
- General-purpose AI models placed on the market before 2 August 2025 must comply.Upcoming
- High-risk obligations apply to stand-alone Annex III systems: asylum and migration, public benefits, education, employment, justice.Upcoming
- High-risk obligations apply to AI in Annex I regulated products, including medical devices.Upcoming
Why Pashto and Dari are hard to evidence
Both are low-resource languages. That changes what can be assumed from English results, and what has to be shown directly.
Very little web text
Dari is not counted separately
Safety training does not transfer evenly
Machine-translated benchmarks mislead
Several varieties, more than one way of writing
Few reviewers can read the output
Where Afghan-language coverage decides the outcome
More than forty years of war and displacement have placed Afghan communities in front of AI systems in asylum offices, hospitals, schools, job centres and social platforms across Europe, North America and Australia.
Asylum, migration and border systems
Healthcare and patient communication
War, displacement and trauma content
Public services, education and employment
Online platforms and content moderation
Frontier model safety
Who needs Afghan-language governance evidence
General-purpose AI model providers and frontier developers
Providers of high-risk AI systems
Deployers: public authorities, hospitals, agencies and NGOs
Online platforms and trust and safety teams
Organizations certifying to ISO/IEC 42001
Auditors, notified bodies, certification bodies and counsel
How we produce the evidence
Six steps, each with a written output you can inspect before the next begins.
- Step 1
Scope and regulatory map
Confirm your role (provider, deployer, frontier developer, certification candidate), the systems in scope, the languages you claim, and the exact articles, sections and clauses the evidence must serve.Output: Signed scope and crosswalk skeleton · Client and Ariana Nexus - Step 2
Evidence inventory and gap analysis
Read what you already hold (model cards, data sheets, evaluation logs, risk registers) and mark what is missing, untested or unreadable for Afghan languages.Output: Gap register · Ariana Nexus - Step 3
Native evaluation and data review
Build or select test sets written by native speakers, run the evaluation, and review data samples for variety, script, quality and provenance. Every item is scored twice; disagreements are adjudicated.Output: Results with agreement statistics · Ariana Nexus - Step 4
Drafting to the requirement
Write each document in the structure its reader expects: Annex IV sections, framework and transparency report entries, ISO documented information. Numbered claims, every figure traceable to its source.Output: Draft evidence file · Ariana Nexus - Step 5
Independent second review
A reviewer who did not run the work checks each claim against its source, each legal reference against the current text, and the language content against the firm’s Cultural Compliance Bureau standards.Output: Review record and sign-off · Second reviewer - Step 6
Handover and update cycle
Deliver into your document control or GRC system with the machine-readable index. Re-test and re-issue on each model release, material change or annual review.Output: Issued file, version 1.0 · Client and Ariana Nexus
A first evidence file for Pashto and Dari on one system typically takes six to ten weeks from a signed scope. Updates for a new model version are faster, because the test sets, procedures and crosswalk already exist.
Ways to engage
Every engagement is scoped and priced as a whole, with a named partner accountable for it. We do not take per-item review work or place unnamed reviewers in another firm’s pipeline.
Why Ariana Nexus
Governance consultancies know the frameworks and cannot read the languages. Language vendors can read the languages and do not know what an auditor will sample. This firm was built to do both.
What we bring
Afghan languages are the entire practice
Scholars, not a crowd
The standard where no certification exists
How we deliver
Written for the person who audits it
Produced in-house
A clean data path
How we are different
Honest coverage
Alongside your certifier and counsel
We do not build or sell AI models, governance software or certification. We do sell evaluation, red teaming and data services; where an evidence file reports on work we performed, the file says so.
What this service does not do
- It does not give legal advice or an opinion on whether you comply.
- It does not certify anything. ISO/IEC 42001 certificates come from accredited certification bodies; EU conformity assessment belongs to the provider or a notified body.
- It does not issue a badge, seal or “compliant” label. No accredited scheme exists for language coverage, so we do not invent one.
- It does not write your English-language governance program. It supplies the Afghan-language evidence that belongs inside it.
- It does not adjust findings. If the model fails in Pashto, the file says so, and says where.
Afghan languages covered
All 24 appear in the coverage statement. We publish readiness tiers rather than a flat claim, because speaker availability differs by two orders of magnitude across this list.
Dari deliverables are Afghanistan Dari, not Iranian Persian. Turkmeni deliverables are Afghan Turkmeni, not the standard Turkmen of Turkmenistan. Hazaragi is a variety of Dari and is staffed separately because the lexical distance is large enough to change a result.
The team behind the evidence
Ariana Nexus is led and operated by alumni and scholars of leading universities who are native speakers of the languages they assess. The people who read the Pashto and Dari output are the people who write the findings.

Hassan Ukasha
Managing Partner, Ariana Nexus, Washington, D.C.Hassan Ukasha oversees the firm’s operations and its AI governance documentation program. He is the accountable partner on every evidence file Ariana Nexus issues: he approves the scope, signs the independence and conflicts declaration, and is the point of escalation for your general counsel, notified body or certification auditor.

Zeba Haqbani

Hussain Ahmad

Wasil Peroz

Maryam Safi
Who is accountable for each part of your file
How this team works, and why it is different
Read, not relayed
Tested in the variety
Two disciplines on every file
Named and accountable

Questions buyers ask
Does the EU AI Act require testing in every language a system supports?
Not in those words. The Act does not list languages. It requires high-risk systems to use training, validation and test data that is relevant and sufficiently representative for the intended purpose and the setting of use (Article 10), to declare accuracy levels and known limitations (Articles 13 and 15), and to document testing in the Annex IV technical file. If Afghan-language speakers are within the intended purpose, the file needs evidence that covers them, or a stated limitation.
What does California SB 53 require about languages?
Before deploying a new or substantially modified frontier model, a frontier developer must publish a transparency report that includes the languages the model supports, its intended uses and its restrictions. Large frontier developers must also summarise their catastrophic-risk assessments and the role of third-party evaluators. SB 53 prohibits materially false or misleading statements about catastrophic risk and its management, which is one reason safety claims should be evidenced in every language the report lists. The law has applied since 1 January 2026.
How does Afghan-language evidence fit into ISO/IEC 42001 certification?
ISO/IEC 42001 certifies the management system, not the model. An auditor samples documented information: the AI system impact assessment (clause 6.1.4 and Annex A.5), data quality and provenance records (A.7), verification and validation (A.6.2.4) and information for users (A.8.2). Where a system in scope is used in Pashto or Dari, those records need language-specific content. We supply it; your certification body audits it.
Does ISO/IEC 42001 certification mean EU AI Act compliance?
No. ISO/IEC 42001 is a voluntary management system standard and is not a harmonised standard under the EU AI Act. It gives a governance structure that much of the Act’s documentation can sit in, but the Act’s specific obligations still have to be met and evidenced separately.
Are Pashto and Dari low-resource languages, and why does that matter for compliance?
Yes. In Common Crawl’s August 2026 crawl, 0.0043% of pages were identified as Pashto against 40.45% English, and Dari is not identified separately from Persian at all. Models see little Afghan-language data, safety training transfers unevenly and public benchmarks are scarce, so performance and safety claims cannot be assumed from English results. They have to be evidenced directly.
Can machine-translated benchmarks be used as evidence?
They can be disclosed, but they are weak evidence. A machine-translated test set carries the translation system’s errors and contains no Afghan names, places, institutions or idiom, so it measures translation quality as much as model quality. We use test items written by native speakers and state the method in the report so an assessor can judge its weight.
Do you provide legal advice or certify compliance?
No. Ariana Nexus is not a law firm, a notified body or a certification body. We produce Afghan-language evidence and documentation and map it to the requirements it supports. Your counsel decides what the law requires of you; your certification body or notified body decides conformity. We work alongside them, never in place of them.
Who does the work?
University-educated native speakers tested in the specific variety they review, led by partners and principals who are alumni and scholars of Cornell University, the University of Chicago, the University of British Columbia and Otto-von-Guericke University. No U.S. certification exists for Pashto or Dari, so every reviewer is assessed against the firm’s written standard and the result is recorded in the evaluator record. All work is done in-house. We do not use subcontractors or crowd platforms.
Which Afghan languages do you cover?
All 24 appear in the coverage statement, in three published readiness tiers. Pashto and Dari, including the Hazaragi variety, receive the full ten-document evidence file today. Uzbeki, Turkmeni, Balochi, Pashayi and the Nuristani group are evidenced on an agreed scope. The smallest languages, such as Ormuri, Parachi, Tirahi and the Pamir group, are scoped to verified speaker availability, and the file says “not evidenced” where that is the truth.
We deploy AI but do not build it. Do we need this?
Often, yes. Under the EU AI Act, deployers of high-risk systems must use them according to the instructions, assign competent human oversight and, for public bodies and certain services, complete a fundamental rights impact assessment before first use. If your clients, patients or applicants include Afghan-language speakers, those steps need language-specific content. We also review a vendor’s language claims before you buy.
How long does it take?
A gap assessment takes two to three weeks. A first full evidence file for Pashto and Dari on one system typically takes six to ten weeks from a signed scope, depending on how much evaluation already exists. Updates for a new model version are faster.
How are our model, data and documents protected?
Work begins under a mutual non-disclosure agreement. Access is limited to the named engagement team, data stays in the environment agreed in the scope, and contributors are pseudonymised in deliverables. Ariana Nexus routes no data, documents or inquiries through channels controlled by the de facto authorities in Afghanistan. European engagements run under GDPR and UK GDPR terms.
How is this different from your LLM evaluation and red teaming services?
Those services produce the measurements. This service turns measurements, ours or your own, into the governance record: coverage statements, data and risk files, oversight procedures and the crosswalk that auditors and regulators read. Many clients start here with a gap assessment, then commission the evaluation the gaps call for.
When do the EU AI Act obligations apply?
Obligations for general-purpose AI model providers have applied since 2 August 2025, and Commission enforcement powers over them began on 2 August 2026. Regulation (EU) 2026/1744 moved the high-risk dates: stand-alone Annex III systems must comply from 2 December 2027, and AI in Annex I regulated products from 2 August 2028.
Related services
Multilingual AI red teaming and safety testing
Generates the adversarial test evidence summarized in D-06.LLM evaluation and benchmark development
The native-written test sets and human evaluation behind D-02.AI interpreting and translation product quality assurance
Accuracy testing for systems that translate or interpret Pashto and Dari.Deepfake and synthetic media analysis in Pashto and Dari
Detection and labeling evidence for Article 50 obligations.Online hate speech and harassment monitoring
Systemic-risk evidence for platforms under Digital Services Act Articles 34 and 35.Find out what your file can and cannot show
Tell us which systems and language claims are in scope and which deadline you are working to. A partner will reply with the questions we need answered to scope a gap assessment.
Request a scoping conversationEngagements begin under a mutual non-disclosure agreement. Washington, D.C. Serving the United States, the European Union and clients worldwide.
Sources
- Regulation (EU) 2024/1689 of the European Parliament and of the Council (Artificial Intelligence Act). Official Journal of the European Union, 12 July 2024. https://eur-lex.europa.eu/eli/reg/2024/1689/oj
- Regulation (EU) 2026/1744 (Digital Omnibus on AI). Official Journal, 24 July 2026; in force 27 July 2026. Council of the EU press release, 29 June 2026. https://www.consilium.europa.eu/en/press/press-releases/2026/06/29/artificial-intelligence-council-gives-final-green-light-to-simplify-and-streamline-rules/
- California Senate Bill 53 (2025), Transparency in Frontier Artificial Intelligence Act. Business and Professions Code § 22757.10 et seq. https://leginfo.legislature.ca.gov/faces/billNavClient.xhtml?bill_id=202520260SB53
- ISO/IEC 42001:2023, Information technology — Artificial intelligence — Management system. https://www.iso.org/standard/81230.html
- European Commission, Template for the public summary of training content for general-purpose AI models, 24 July 2025. https://digital-strategy.ec.europa.eu/en/faqs/template-general-purpose-ai-model-providers-summarise-their-training-content
- Regulation (EU) 2022/2065 (Digital Services Act), Articles 34 and 35. https://eur-lex.europa.eu/eli/reg/2022/2065/oj
- Yong, Z.-X., Menghini, C. and Bach, S. H. (2023). Low-Resource Languages Jailbreak GPT-4. arXiv:2310.02446. https://arxiv.org/abs/2310.02446
- Common Crawl Foundation. Statistics of Common Crawl monthly archives: distribution of languages, CC-MAIN-2026-34. https://commoncrawl.github.io/cc-crawl-statistics/plots/languages
Reviewed by the Ariana Nexus Cultural Compliance Bureau. Last updated 19 September 2026.