Skip to main content
Technology, AI, and Digital Platforms — Service 04

AI Interpreting and Translation Product Quality Assurance — Pashto, Dari, and 22 More Afghan Languages

Independent accuracy testing, adversarial evaluation, and deployment assurance for the vendors who build AI interpreting and machine translation products — and for the hospitals, federal agencies, courts, and school districts that deploy them.

What an evaluation reports
Speech recognition error, by variety
Translation error, ISO 5060 typology
Silent failure rate
Directional asymmetry
Dialect spread
Failure and uncertainty behavior
Illustrative structure. Every published figure carries its specification, sample design, and confidence interval.
Definition

What is AI interpreting and translation quality assurance?

AI interpreting and translation quality assurance is independent testing that measures how accurately an AI interpreting, speech translation, or machine translation product performs in a specific language before and after deployment. Ariana Nexus evaluates these products in Pashto, Dari, and 22 more Afghan languages using the ISO 5060 error typology and MQM scoring, then reports where the product is safe to use, where it is not, and why.

Two buyers. Vendors — AI interpreting platforms, speech translation companies, machine translation providers, and frontier labs shipping translation features — buy evidence they can put in front of customers and regulators. Deploying institutions — health systems, federal agencies, courts, school districts, and resettlement organizations — buy verification before they sign, and monitoring after they go live.

Reviewed

Last reviewed. Next scheduled review March 2027. Reviewed by the Technology, AI, and Digital Platforms practice and the Cultural Compliance Bureau.

24

Afghan languages evaluated, with dialect-level scoring for Pashto and Dari rather than a single language-level average.

7

error categories in the ISO 5060 typology, harmonized with MQM, applied to every segment we score.

Aug 2, 2026

EU AI Act Article 50 transparency obligations became enforceable. The Digital Omnibus deferred the high-risk regime; it did not defer Article 50.

59.3%

of healthcare leaders surveyed in 2026 said they were not confident AI interpreting works correctly in real interactions.

Into English

is where accuracy falls hardest. Published reviews find AI translation degrades substantially translating into English from non-European languages.

0

AI interpreting or translation products sold by Ariana Nexus. We evaluate; we do not compete with the systems we test.

Why AI interpreting and machine translation fail in Pashto and Dari

The dangerous errors are not the obvious ones. A garbled sentence gets caught. A fluent, confident, grammatical sentence that says the wrong thing does not.

Failure 01

Afghan languages are low-resource. Pashto, Dari, and the 22 other languages of Afghanistan have a fraction of the training data behind English, Spanish, or Mandarin, which is why Pashto machine translation accuracy and Dari translation accuracy testing produce numbers that look nothing like the benchmarks a vendor quotes. Published systematic reviews of machine translation in healthcare report substantial variability in accuracy and call for evaluation before clinical use, not after. Low resource language MT evaluation is a separate discipline, not a smaller version of the same job.

Failure 02

Accuracy collapses in the direction that matters. Research on AI translation tools finds performance is reasonable out of English and degrades substantially translating into English, particularly for non-European languages. In a clinical or legal encounter, the patient-to-provider direction is the one carrying the symptom, the consent, the testimony.

Failure 03

Speech adds a second failure layer. AI interpreting is not one model. It is speech recognition, then machine translation, then speech synthesis. Recognition accuracy already drops for accents and dialects; every recognition error is inherited by the translation and spoken aloud with full confidence.

Failure 04

A language-level score hides a dialect-level failure. A product can post a strong average on Dari and fail on Hazaragi, or hold on Kabuli Pashto and break on Kandahari. An average is not parity, and the patient in front of the clinician is not an average.

Failure 05

The error changes the decision without announcing itself. A published example: a patient describes chest tightness, the system renders it as discomfort rather than pressure, and the clinician defers further evaluation. Nothing in the output looks wrong. That is the failure mode this service exists to find.

Who this service is for

For vendors

You build the product

AI interpreting platforms · speech translation companies · machine translation providers · language technology vendors · frontier AI labs shipping translation features · healthcare communication platforms

You need accuracy evidence a customer's procurement team and a regulator will both accept, in languages your internal evaluators cannot read.

A pre-release accuracy evaluation in the language pairs and varieties you claim to support

A benchmark set for Pashto, Dari and the other languages you ship, which you keep and re-run against every model version

Red teaming translation models in Afghan languages: code-switching, transliteration, dialect, noise, overlapping speech

Human parity claim testing before a marketing or procurement claim goes out

Evidence structured for EU AI Act Article 50 transparency files and NIST AI Risk Management Framework translation documentation

A regression harness so the next model release does not quietly lose a language

For deploying institutions

You buy and deploy the product

health systems and hospitals · federal agencies and primes · courts and legal aid · school districts · payers · resettlement and refugee-serving organizations

You need to know whether the tool in front of you is safe for Afghan patients, applicants, litigants, and families — before it is in the room, and continuously after.

An AI interpreting procurement checklist and AI translation vendor due diligence on every shortlisted product, scored on one rubric

Verification of the vendor's own published accuracy claims against your real content

A machine translation risk assessment: where the tool may be used, where it may be used only with a qualified human interpreter present, and where it must not be used at all

An AI interpreter pilot a health system can actually defend to its board, with acceptance thresholds and remedies written into the contract

Post-deployment monitoring with escalation triggers and a human-interpreter fallback protocol

Section 1557 machine translation and Title VI language access AI documentation that survives an audit or accreditation review

What we test: the eight layers of an AI interpreting and translation evaluation

L01

Speech recognition accuracy

Pashto ASR word error rate and character error rate, Dari speech recognition accuracy, and Hazaragi speech recognition scored separately, by speaker gender, age band, and audio condition. Accented speech, background noise, overlapping talk, disfluency, and code-switching between Pashto, Dari, Urdu, and English.

L02

Machine translation quality

Segment-level scoring under the ISO 5060 translation evaluation standard and the MQM error typology, plus chrF++ and neural quality estimation as secondary signals. Human scoring is primary; automatic metrics never stand alone. Where post-editing is in scope, ISO 18587 post-editing requirements set the bar.

L03

End-to-end speech translation

The full pipeline scored as one product: semantic fidelity from source speech to target speech, latency, turn-taking behavior, interruption handling, and what the system does when it cannot hear.

L04

Critical-content integrity

Numbers, doses, units, dates, durations, frequencies, negation, named entities, and place names. A single negation flip or dose error is scored critical regardless of the rest of the segment.

L05

Cultural and register fidelity

Kinship terms, honorifics, gendered address, religious register, indirect description of symptoms, and the vocabulary Afghan patients and families actually use for pain, grief, trauma, and shame.

L06

Dialect and variety parity

Kandahari, Kabuli, and eastern Pashto; Kabuli and Herati Dari; Hazaragi as a variety of Dari staffed and scored separately. Scored as a spread, reported as a range.

L07

Uncertainty and failure behavior

What the product does when it is unsure. Does it signal low confidence, ask for repetition, hand off to a human — or does it produce a fluent guess? This layer is where silent failure lives.

L08

Script, rendering, and safety

Perso-Arabic rendering, right-to-left handling, diacritics, Pashto letters absent from Persian keyboards, and orthography drift toward Iranian Persian. Red teaming translation models in Afghan languages sits here: refusal behavior, toxicity, and jailbreak resistance under code-switching and transliteration attack.

Headline metric

The number we built this practice around: silent failure rate

Most evaluation reports tell you how often a system is wrong. We also tell you how often it is wrong in a way nobody in the room could detect.

Silent failure rate is the share of scored segments that contain a critical or major error while remaining fluent, grammatical, and confident in the target language — output that a monolingual clinician, officer, judge, or teacher has no way of flagging. We report it separately from overall error rate because it is the only number that predicts harm in a setting where no second language is present in the room. A product can hold a respectable overall score and still carry an unacceptable silent failure rate, and that is exactly the product that gets deployed.

Silent failure rate

Critical and major errors that are fluent and undetectable without a second language present. Reported separately, by language and by setting.

Overall error rate

Every error, weighted by severity, per hundred scored segments. The number most reports stop at.

Hallucination rate, machine translation

Content present in the output that has no basis in the source — invented symptoms, invented instructions, invented entities. Counted as unsupported additions and reported as its own rate.

Omission rate, interpreting

Source content that never reaches the output at all. In speech, the commonest and least visible failure: nothing looks wrong because nothing is there.

Directional asymmetry

The gap between English-into-Pashto and Pashto-into-English performance, scored as its own figure rather than averaged away.

Dialect spread

The distance between the best-performing and worst-performing variety inside one language — the Pashto dialect parity model in practice. Reported as a range, never a mean.

How we measure AI interpreting accuracy: ISO 5060 translation evaluation, the MQM error typology, and dual independent scoring

ISO 5060:2024 is the first international standard for evaluating translation output — human, post-edited, and raw machine translation. Its error typology is harmonized with MQM. We run it as written, and we hand you the scoring sheets.

Phase 1

Specification

Before anything is scored, we write the specification: language pairs, varieties, domain, audience, register, intended use, and what counts as a critical error in your setting. ISO 5060 calls this the pre-evaluation phase. Quality has no meaning without it.

Phase 2

Sampling

We draw a statistically defensible sample from real content — your transcripts, your documents, your recorded encounters under agreement — not from a demo script. Sample size is stated, and so is the confidence interval on every reported figure.

Phase 3

Dual independent scoring

Two qualified evaluators score every segment blind to each other. We report inter-rater agreement. Where they disagree beyond threshold, a third adjudicates and the adjudication is logged.

Phase 4

Severity and weighting

Each error is typed and assigned a severity — neutral, minor, major, critical. Penalty points convert to a quality rating against the threshold set in the specification, not against a number we invented afterward.

Phase 5

Back-translation review and targeted probes

Back-translation review is run on selected segments by an evaluator who has not seen the source — the standard method for surfacing meaning loss that reads fluently. Adversarial probes then run against the failure classes the specification identified as highest-risk.

Phase 6

Report, dispute, re-test

You receive the report, the scoring sheets, and the benchmark set. ISO 5060's post-evaluation phase includes a dispute route, and we honor it: you may challenge any score, and the adjudication is documented.

What you receive

Evaluation report

Findings by language, variety, direction, and setting, with the specification, sample design, and confidence intervals stated on the face of the report.

Scoring sheets

Every scored segment, every error type, every severity, every adjudication. Not a summary. The underlying record.

Language-pair scorecard

A one-page scorecard per language pair, built to be handed to a procurement committee, a clinical safety group, or a customer.

Failure taxonomy

The named, ranked list of how this specific product fails in these specific languages — the document an engineering team can actually act on.

Benchmark set, Pashto and Dari

The evaluation corpus, delivered to you with rights to keep and re-run, covering every language pair in scope. You own it. It does not expire with the engagement.

Regression harness

A re-runnable test suite so the next model version can be checked against the same bar without commissioning a new study.

Deployment decision memo

For institutions: a setting-by-setting statement of where this product may be used, where it may be used with a human interpreter present, and where it must not be used at all.

Compliance annex

Findings mapped to the instruments that apply to you — EU AI Act Article 50, NIST AI RMF, ISO/IEC 42001, Section 1557, Title VI, and accreditation expectations.

Vendor demo, automatic metric, or independent evaluation

Vendor demo or pilot
Automatic metric only
Ariana Nexus evaluation
 Content tested
Vendor demo or pilotScripted or curated
Automatic metric onlyPublic test set
Ariana Nexus evaluationYour real content, sampled to a stated design
 Who scores
Vendor demo or pilotThe vendor
Automatic metric onlyA model
Ariana Nexus evaluationTwo independent qualified evaluators, blind, plus adjudication
 Dialect resolution
Vendor demo or pilotNone
Automatic metric onlyNone
Ariana Nexus evaluationScored and reported per variety
 Direction tested
Vendor demo or pilotUsually into the language
Automatic metric onlyBoth, averaged
Ariana Nexus evaluationBoth, reported separately
 Silent failure
Vendor demo or pilotNot measured
Automatic metric onlyNot measured
Ariana Nexus evaluationMeasured and reported as its own figure
 Evidence handed over
Vendor demo or pilotSlide deck
Automatic metric onlyA score
Ariana Nexus evaluationReport, scoring sheets, benchmark set, harness
 Usable in a compliance file
Vendor demo or pilotNo
Automatic metric onlyPartly
Ariana Nexus evaluationStructured for Article 50, NIST AI RMF, ISO/IEC 42001

The standards and regulations this evaluation is built against

Instrument
Title
Requirement
Status
InstrumentISO 5060:2024
TitleTranslation services — Evaluation of translation output
RequirementError typology harmonized with MQM, severity and weighting, penalty points, sampling, evaluator competence, and a defined dispute route.
StatusISO 5060 translation evaluation — applied as the scoring method
InstrumentMQM
TitleMultidimensional Quality Metrics
RequirementThe error framework ISO 5060's typology is harmonized with. Also the basis of ASTM WK46396, the draft US practice for analytic evaluation of translation quality.
StatusMQM error typology — applied as the error framework
InstrumentISO 18587:2017
TitlePost-editing of machine translation output
RequirementThe requirements standard when raw machine translation is the starting point and a qualified linguist must bring it to publishable quality.
StatusISO 18587 post-editing — applied where post-editing is in scope
InstrumentISO 17100:2015
TitleTranslation services — Requirements
RequirementThe process and competence standard for human translation, used as the reference point when a product claims parity with human output.
StatusApplied as the comparison baseline
InstrumentEU AI Act, Article 50
TitleTransparency obligations, Regulation (EU) 2024/1689
RequirementEnforceable since August 2, 2026. Not deferred by the Digital Omnibus. Deployer and provider disclosure duties apply regardless of risk classification; the Article 50(2) marking deadline for systems already on the market is December 2, 2026.
StatusIn force — evidence prepared
InstrumentEU AI Act, high-risk regime
TitleAnnex III and Annex I, as amended by Regulation (EU) 2026/1744
RequirementThe Digital Omnibus on AI entered into force July 27, 2026, moving Annex III standalone high-risk compliance to December 2, 2027 and Annex I product-embedded systems to August 2, 2028.
StatusDated — build to the new calendar
InstrumentISO/IEC 42001:2023
TitleAI management systems
RequirementThe management-system standard an AI vendor is most often asked to evidence. Our evaluation output is structured to drop into its performance-evaluation clauses.
StatusApplied to report structure
InstrumentNIST AI RMF 1.0 and AI 600-1
TitleAI Risk Management Framework and Generative AI Profile
RequirementThe US framework federal buyers and primes ask about by name. Measure and Manage functions are where third-party language evaluation belongs.
StatusNIST AI Risk Management Framework translation mapping
Instrument45 CFR 92.201(c)(3)
TitleSection 1557, Affordable Care Act
RequirementMachine translation of vital health content requires review by a qualified human translator where the text is critical to rights, benefits, or informed consent, or where accuracy is essential. The AI boundary is already written into federal law.
StatusSection 1557 machine translation boundary
InstrumentTitle VI, Civil Rights Act
TitleMeaningful access for limited-English-proficient persons
RequirementApplies to every recipient of federal financial assistance. An inaccurate AI tool does not discharge the obligation; it creates the record that you failed it.
StatusTitle VI language access AI obligation

The 24 Afghan languages we evaluate, and the varieties inside them

Coverage is stated by family, with the varieties that actually change a score named. Pashto and Dari carry full dialect-level evaluation today; the remaining languages are evaluated on a scoped basis with the panel named in the specification.

Iranian

13 languages

PashtoDariHazaragiAimaqBalochiOrmuriParachiWakhiShughniSanglechiIshkashimiMunjiYidgha
Turkic

3 languages

UzbekiTurkmeniKyrgyz
Indo-Aryan

3 languages

PashayiGawarbatiTirahi
Nuristani

4 languages

Nuristani (Ashkun group)KatiPrasunWaigali
Dravidian

1 language

Brahui
Language
Variety note
LanguagePashto
Variety noteKandahari (southern), Kabuli (central), and eastern varieties scored separately. Retroflex consonants and Pashto-only letters are a recurring recognition failure.
LanguageDari
Variety noteAfghanistan Dari, not Iranian Persian. Kabuli and Herati varieties scored separately. Orthography and lexical drift toward Iranian Persian is scored as an error, not a stylistic choice.
LanguageHazaragi
Variety noteA variety of Dari, evaluated and staffed separately because recognition and translation performance on it diverges sharply from standard Dari.
LanguageUzbeki and Turkmeni
Variety noteAfghan Turkmeni is distinct from the standard Turkmen of Turkmenistan. Products trained on the latter are scored against the former.

Independence, and the conflict we declare in writing

Ariana Nexus does not build, license, resell, or white-label an AI interpreting or machine translation product. We have no engine in the market and no revenue that depends on one system scoring above another. That is the first condition of an evaluation being worth anything.

We do sell human interpreting and translation in Afghan languages. That is a conflict, and we disclose it on the face of every report rather than leaving a buyer to discover it. The controls are structural: evaluators assigned to a product evaluation are not drawn from the interpreting bench, scoring is done against a rubric fixed in the specification before any output is seen, and you receive the raw scoring sheets — so our conclusion can be checked against our own evidence.

No evaluation is sold with a result attached. We do not offer a pass, a badge, a seal, or a certification, because no accredited certification scheme exists for AI interpreting accuracy in Afghan languages. Any firm offering you one is selling you something that does not exist.

Findings belong to the client. We do not publish a client's results, name a client's product, or use a result in our own marketing without written consent. Where a vendor asks us to review a public accuracy claim, we will say plainly whether the evidence supports it.

What we will not do

We do not tune the system we are evaluating in the same engagement.

Remediation advice is delivered as a findings document. If you want us to help fix what we found, that is a separate engagement with a separate team, disclosed in both reports.

We do not issue certifications or conformity marks.

We issue evidence. Certification against the EU AI Act runs through notified bodies and harmonized standards, and the harmonized standards are not ready. We say so rather than filling the gap with a logo.

We do not publish exploit tooling.

Adversarial findings, including jailbreak and harmful-content probes in Afghan languages, go to the system owner to remediate. They are not released, demonstrated publicly, or reused against another client's system.

We do not route content through channels controlled by the de facto authorities in Afghanistan.

No data, documents, or inquiries. This is a firm compliance and security line, and it is one reason our evaluation work is deliverable at all.

We do not tell you an AI product is unsafe as a route to selling you interpreters.

Several products we have assessed are appropriate for defined, lower-risk uses, and we say so in writing. The recommendation follows the evidence, in both directions.

Why Ariana Nexus, how we deliver, and how we are different

Anyone can run an automatic metric over a test set. The scarce thing is a qualified human who can read Kandahari Pashto, recognize a clinical register error, and defend the score in front of an engineering team.

There is no US certification for Afghan-language interpreting. We set the standard.

No accrediting body certifies interpreters or evaluators in Pashto, Dari, or any of the other 22 Afghan languages. Where no certification exists, someone has to own the bar. Ariana Nexus trains its own linguists and Afghan diaspora interpreters, publishes the rubric it scores against, and hands over the evidence so the bar can be inspected rather than trusted.

We are a consulting and professional services firm, not a language vendor with an evaluation add-on.

The engagement is scoped, specified, sampled, and reported the way an assurance engagement is. The people writing the specification are the people who will sit in front of your safety committee and defend it.

Every part is produced by our own people.

No subcontractors, no brokered specialists, no referral to an outside bench. One engagement, one point of accountability, one signature on the report.

Our evaluators are scholars and alumni of leading universities, not bilinguals.

Cornell, Brown, the University of Chicago, the University of British Columbia, Otto-von-Guericke. Native command of the variety, plus the research training to design a sample, run inter-rater statistics, and write a finding that holds up.

We score dialects, not languages.

Pashto is not one thing and neither is Dari. We report Kandahari separately from Kabuli, Herati separately from Kabuli Dari, and Hazaragi as a variety of Dari with its own score. A single language number is a marketing figure, not an evaluation.

Afghan context is the work, not a label on it.

The people scoring your output have lived the context the output describes: Afghan patients, the Afghan diaspora, the Afghan war and displacement, and the indirect vocabulary families use for loss, moral injury, moral trauma, and shame. Those are the segments where fluent machine output is most confidently wrong, and they are the reason Afghan languages AI evaluation cannot be outsourced to a crowd platform.

One bench, every adjacent question.

Afghan patients language access technology, LLM translation evaluation in Pashto, AI medical interpreter evaluation, speech translation accuracy testing, and Afghan diaspora AI services sit in one practice with one standard. You are not re-briefing a new vendor every time the question changes shape.

How engagements are structured

Baseline Evaluation

A single product, one to three language pairs, one domain.


Specification and sample design

Dual independent scoring under ISO 5060

Silent failure and dialect spread reporting

Evaluation report, scoring sheets, benchmark set


Typical use: a vendor preparing a claim, or an institution evaluating one shortlisted product.

Deployment Assurance

Multi-product or multi-setting, scoped to a deployment decision.


Everything in Baseline, across the products or settings in scope

Comparative scorecard on a single rubric

Setting-by-setting deployment decision memo

Acceptance criteria and contract thresholds

Compliance annex and fallback protocol


Typical use: a health system, agency, or prime choosing between vendors and documenting the choice.

Continuous Assurance Program

An annual program, not a project.


Regression testing on every model release

Quarterly re-scoring against the maintained benchmark set

Live monitoring thresholds and escalation triggers

Named partner and standing evaluation panel

Annual report to your board, safety committee, or regulator


Typical use: a vendor shipping continuously, or an institution with the product in live clinical, legal, or educational use.

The team behind this service

Ariana Nexus is operated and led by alumni and scholars of leading universities. The people who specify, score, and defend an evaluation are named here, with their degrees and institutions, because in a field with no certification the credential of the evaluator is the credential of the evaluation.

This is the part most firms cannot staff. An AI interpreting evaluation in Pashto needs three things in the same person or the same room: native command of the specific variety, subject-matter depth in the domain being translated, and the research training to design a sample and run inter-rater statistics. Bilingual is not one of the three. Every evaluator on this service holds a university degree, is assessed in the specific variety claimed, and works to a published rubric.

Program oversight

Hassan Ukasha

Managing Partner

B.S. Cornell University

M.P.H. Cornell University

Every evaluation this practice publishes is released under the Managing Partner's signature. Scope, independence, and the conflict declared on the face of each report are his to approve before any finding leaves the firm.

Hassan Ukasha oversees the firm's operations and this program. He approves the evaluation specification before any system output is seen, confirms that the evaluator panel is drawn from outside the interpreting bench, and signs the independence statement that accompanies every report. Where a finding is contested, the dispute route under ISO 5060 runs to him.

Senior Partner

Zeba Haqbani

B.A.
The American University of Afghanistan
B.A.
University of British Columbia

Leads evaluation engineering for this service. Builds and runs the firm's institutional systems, technology, and AI platforms — including the benchmark infrastructure, scoring pipeline, and regression harness behind every evaluation. Lived in Kabul.

Principal

Hussain Ahmad

M.Eng.
Cornell University
Ph.D.
University of Chicago

AI and data engineering. Designs the measurement layer: speech recognition error analysis, quality-estimation modeling, sampling design, and the statistics that turn a set of scores into a defensible finding.

Principal

Wasil Peroz

B.A.
Milli University
M.Sc.
Otto-von-Guericke University Magdeburg

Institutional law and AI regulation. Maps findings to the instruments that govern the deployment — EU AI Act Article 50 and the amended high-risk calendar, GDPR and UK GDPR, Section 1557, and Title VI — and writes the compliance annex.

Principal

Maryam Safi

B.A.
Cornell University

Cultural Compliance Bureau review. Owns the cultural and register layer: kinship terms, honorifics, gendered address, religious register, and the indirect vocabulary Afghan families use for pain, grief, and shame.

Built in Washington, D.C., for institutions that have to answer for the result

Marble and glass institutional lobby

Ariana Nexus, 1717 Pennsylvania Avenue NW, Washington, D.C. Evaluations are specified, scored, and adjudicated in-house — no subcontracted panels, no brokered reviewers.

Questions buyers ask

Is AI interpreting safe for hospitals?

Not as a single yes or no. Safety is a property of a specific product, in a specific language and direction, in a specific setting — not of the category. A product may be acceptable for English-into-Dari wayfinding and unsafe for a Pashto consent discussion in the same hospital on the same day. What a health system needs is a machine translation risk assessment that states which is which, in writing, with the evidence attached. That is what this service produces.

How accurate is Pashto machine translation?

There is no single answer, and any vendor that gives you one is quoting a benchmark rather than measuring your use. Pashto machine translation accuracy varies by direction, by dialect, by domain, and by whether the input is clean text or recognized speech. Published research consistently finds that accuracy falls hardest translating into English from non-European languages — which is the direction carrying the symptom, the consent, and the testimony. We measure your product on your content and report the figure with its confidence interval.

How do you evaluate an AI interpreting vendor?

We score every shortlisted product on one rubric, using your real content rather than the vendor's demo, with two independent qualified evaluators scoring blind. The output is an AI interpreting procurement checklist you can put in a committee pack: comparative scorecards, a setting-by-setting decision memo, acceptance thresholds for the contract, and the raw scoring sheets so the conclusion can be checked against the evidence. AI translation vendor due diligence that stops at the vendor's own slide deck is not due diligence.

What does AI medical interpreter evaluation cover?

Speech recognition accuracy on accented and noisy clinical audio; translation quality under the ISO 5060 error typology; end-to-end speech translation fidelity and latency; integrity of numbers, doses, dates and negation; clinical and cultural register; dialect parity; and what the system does when it is unsure. Each layer is scored separately, because collapsing them into one number is what hides the failure a clinician would never catch.

Do you run speech translation accuracy testing, or only text?

Both, and they are different jobs. Text-only machine translation quality evaluation scores a written segment against a specification. Speech translation accuracy testing scores the full pipeline — recognition, translation, synthesis — as the user meets it, including latency, turn-taking, interruption handling, and behavior when the system cannot hear. A product can score well on text and fail in the room.

What does low resource language MT evaluation involve that standard evaluation does not?

Reference scarcity, dialect fragmentation, and unreliable automatic metrics. There is no large trusted reference corpus to score against, so references are built as part of the engagement. Automatic metrics such as chrF++ correlate poorly with human judgment at low resource, so they run only as secondary signals. And a language-level score is meaningless where varieties diverge as sharply as Kandahari and Kabuli Pashto, so scoring resolution has to go one level deeper than it does in Spanish or French.

Is AI interpreting safe to use with Afghan patients?

It depends on the product, the language, the direction, and the setting — which is why the question has to be answered with evidence rather than a policy. Some products perform acceptably on English-into-Dari for low-risk, scripted exchanges such as wayfinding or appointment confirmation. The same product may be unsafe for consent, discharge instructions, or a behavioral health assessment in Pashto. Our deployment decision memo states which is which for your specific product and setting.

Can machine translation be used for vital documents under Section 1557?

Not on its own. 45 CFR 92.201(c)(3) requires review by a qualified human translator when machine translation is used for text that is critical to rights, benefits, or informed consent, or where accuracy is essential. The federal rule already draws the AI boundary; the practical question is whether your human review layer is actually catching what the machine gets wrong, which is what we measure.

How do you measure AI interpreting accuracy?

Human scoring under the ISO 5060:2024 error typology, harmonized with MQM, is primary. Each segment is typed by error category, assigned a severity, and converted to a quality rating through penalty points against a threshold fixed in advance. Automatic metrics such as chrF++ and neural quality estimation run alongside as secondary signals — they are useful for regression detection and unreliable as a verdict, particularly in low-resource languages.

What is a silent failure rate?

The share of scored segments containing a critical or major error while remaining fluent, grammatical, and confident in the target language. It is the error that a monolingual clinician, officer, judge, or teacher cannot detect because nothing about the output looks wrong. We report it separately from overall error rate because it is the figure that predicts harm.

Why do you score Pashto dialects separately instead of reporting one Pashto number?

Because a single number hides the failure. A product can hold a strong average on Pashto and break on Kandahari, and the speaker in front of your clinician is not an average. We report Kandahari, Kabuli, and eastern varieties separately, and we report the spread between best and worst as its own figure. The same applies to Kabuli and Herati Dari, and to Hazaragi as a variety of Dari.

Does the EU AI Act apply to our AI translation product?

Article 50 transparency obligations have been enforceable since August 2, 2026 and apply to providers and deployers regardless of whether the system is classified as high-risk — the Digital Omnibus deferred the high-risk regime, not Article 50. The Annex III high-risk deadline moved to December 2, 2027 and Annex I to August 2, 2028 under Regulation (EU) 2026/1744. Our compliance annex maps findings to the specific obligation and date.

Do you certify AI interpreting products?

No. No accredited certification scheme exists for AI interpreting accuracy in Afghan languages, and the harmonized standards that would underpin EU conformity assessment are not ready. We issue evidence — report, scoring sheets, benchmark set — not a badge. Any firm offering you a certification for this is selling something that does not exist.

You also sell human interpreting. Isn't that a conflict?

Yes, and we declare it on the face of every report rather than leaving you to find it. The controls are structural: evaluators on a product evaluation are not drawn from the interpreting bench, the rubric is fixed in the specification before any output is seen, and you receive the raw scoring sheets so our conclusion can be checked against our own evidence. We do not build or resell an AI interpreting product, so no result moves revenue toward a system we own.

How large a sample do you need, and can you use our real data?

Sample size follows from the confidence interval you need on the reported figures, the number of language pairs and varieties in scope, and the error rate we observe in a pilot draw. We state the design before scoring begins. Real content produces a valid evaluation and demo scripts do not, so we work from your transcripts, documents, or recorded encounters under a data agreement — with a business associate agreement in place before any protected health information is touched.

What is the difference between this and your LLM evaluation service?

LLM evaluation and benchmarking assesses a model's general capability in Afghan languages. This service assesses a shipped product — the interpreting app, the speech translation feature, the machine translation engine in your workflow — as the user actually meets it, including speech recognition, latency, failure behavior, and the human review layer around it.

Can you evaluate a product we have already deployed?

Yes, and it is the more common engagement. Post-deployment work draws from live content under agreement, establishes the baseline you should have had at procurement, sets monitoring thresholds and escalation triggers, and defines the fallback protocol for handing an encounter to a qualified human interpreter.

Which languages can you evaluate beyond Pashto and Dari?

All 24 Afghan languages, with the depth stated honestly. Pashto and Dari carry full dialect-level evaluation today. Uzbeki, Turkmeni, Hazaragi, Balochi, and Pashayi are evaluated on a scoped basis. For the smallest languages — Ormuri, Parachi, Tirahi, Prasun, and the Pamir group — the evaluator panel is named in the specification before the engagement begins, because a fill guarantee in a language with a few thousand speakers would not be honest.

Related services

Pashto and Dari Training Data and Annotation for AI Models

Corpus collection, expert annotation, and preference data in Afghan languages.

View service

LLM Evaluation and Benchmarking for Pashto, Dari and Afghan Languages

Model-level benchmarking, dialect parity reporting, and rubric-based human evaluation.

View service

Multilingual Red Teaming and AI Safety Testing in Afghan Languages

Adversarial testing for jailbreaks, harmful content, and code-switching attacks.

View service

AI and Machine Translation Quality Review for Patient Communications

For health systems reviewing their own patient-facing translated content, rather than a vendor's product.

View service

Engage

Commission an evaluation

Tell us the product, the language pairs, and the setting. We will come back with a specification and a sample design before any commercial conversation. Engagements begin under a mutual non-disclosure agreement.

Start an engagement

Ariana Nexus · 1717 Pennsylvania Avenue NW, 10th Floor, Washington, D.C. 20006 · +1 202-771-0224