Skip to main content
Publications
/
Language and data

Afghan Language Governance in Healthcare, AI, Law Enforcement and Cross-Border Data

The 24 languages of Afghanistan, and the identification, handling and transfer standards each sector is held to.

Edition
2026
Length
For
Compliance, data and language-access leads
Download the abstract
PDF ·
80 KB
The full document is in preparation. The abstract above is available now.
Register to receive it on release
The download above is the published abstract. The complete document is distributed to Ariana Nexus clients and partners, and to institutions in an active evaluation.
Request the full document
Afghan Language Governance in Healthcare, AI, Law Enforcement and Cross-Border Data

What this document covers

Afghanistan is not a one-language country, and “Afghan” is not a language. This document sets out the 24 languages of Afghanistan across five language families, gives each one’s name in Arabic script alongside its English name, and treats correct identification as the first control rather than a preliminary step. Misidentification is where most failures begin — a Pashto interpreter booked for a Dari speaker, or Iranian Persian supplied where Afghanistan Dari is required. Afghanistan Dari is not Iranian Persian, and Afghan Turkmeni is not the standard Turkmen of Turkmenistan; a system that flattens either will mis-serve the speaker while recording a match.

The healthcare section deals with what correct identification requires at intake. A language-access obligation is not discharged by recording a language name. It depends on getting the language, the dialect and the region right rather than the language alone, and on carrying that record forward so that a second appointment is staffed as accurately as the first. The section sets out what an intake process has to capture for that to be possible.

The AI and data section deals with what a laboratory needs to establish before it trains or evaluates on Afghan-language material. The governing question is which varieties are represented in a corpus and which are not. A dataset labelled only “Pashto” or only “Dari” cannot answer that question, and a model evaluated against such a set will report a competence it does not have. The section describes what a laboratory should be able to state about provenance and variety coverage before it publishes a result.

The final section covers law enforcement and cross-border transmission: the handling and transfer standards that apply to Afghan-language material moving between jurisdictions, and the questions an institution should be able to answer about where that material has been, who has read it, and under what authority it moved.

It also states the firm’s own compliance boundary plainly. Ariana Nexus does not route documents, data or inquiries through channels controlled by the de facto authorities in Afghanistan. That boundary is a condition of the work rather than a preference, and it is published here so that a counterparty can hold the firm to it.

The document is written for the compliance, data and language-access leads who have to defend a decision after it has been made.

Document details

Edition
2026
Pages
Format
Booklet, 210 × 280 mm
Languages
English
Published
Audience
Language and data
Series

Inside the document

Afghan Language Governance in Healthcare, AI, Law Enforcement and Cross-Border Data
Afghan Language Governance in Healthcare, AI, Law Enforcement and Cross-Border Data

Work with Ariana Nexus

Start an engagement