Skip to main content

Technology, AI and Digital Platforms

Afghan Name Matching, Transliteration and Identity Data for Screening and KYC

One Afghan name can be written twenty or more ways in the Latin alphabet, and almost none of those spellings are wrong. Screening systems, KYC onboarding checks and vetting files built on English-language name logic break on Pashto and Dari records — flagging thousands of ordinary Afghans and missing the records that matter. Ariana Nexus supplies the name-variant data, transliteration rules, tuning evidence and human review that make those systems work.

Service status · 20 September 2026
Scripts read in sourcePashto and Dari, plus 22 more Afghan languages
Standards mappedBGN/PCGN Afghanistan 2007 · BGN/PCGN Pashto · ICAO 9303 · ALA-LC · ISO 233
Lists tested againstUN 1988 (Taliban) List · OFAC SDN · EU and UK consolidated lists · PEP and adverse media
Delivered byAfghan scholars and analysts, in-house — no subcontractors
Engagement baseWashington, D.C.

Exhibit · one name, twenty-two renderings

Twenty-two Latin renderings of one illustrative name, each built from a spelling pattern common in Afghan records. An exact-match system sees twenty-two people. A fuzzy matcher at 0.90 Jaro-Winkler still splits six of them off, mostly where a title, an abbreviation or a missing name changes the shape of the string — and it groups “Mohammad Ajmal”, a fragment many unrelated people share. A caseworker sees one person.

محمد اجمل زدران

Illustrative example · Mohammad Ajmal Zadran · given name, given name, qawm name

Mohammad Ajmal ZadranMohammed Ajmal ZadranMuhammad Ajmal ZadranMohammad Ajmal JadranMohamad Ajmal ZadranMohd Ajmal ZadranMohammad AjmalAjmal ZadranAjmal Mohammad ZadranMohammad A. ZadranM. Ajmal ZadranMohammadajmal ZadranMohammad Ajmal ZdranMuhammed Ajmal JadranMohammad Ajmal ZadraanAjmal Khan ZadranHaji Mohammad Ajmal ZadranMohammad Ajmal ZardanMohamed Ajmal ZadranMohammad Ajmal ZadrānMuhammad Agmal ZadranMohammad Ajmal Zadran Khost

Solid marker: Jaro-Winkler similarity of 0.90 or higher against the first record (lower-cased, accents folded). The matcher groups it.Hollow marker: below 0.90. The matcher treats it as a different person.

01

What is Afghan name matching?

Afghan name matching is the problem of deciding whether two records refer to the same Afghan person after a name written in Pashto or Dari script has been romanized into the Latin alphabet inconsistently, without a surname, and often without a reliable date of birth. Ariana Nexus builds the name-variant corpora, transliteration rule sets, screening test data and written analyst review that sanctions screening, KYC and vetting programs need to make that decision defensibly.

02

Nine reasons Afghan names break a screening system

Every failure below is a data problem with a documented cause. None of them are solved by raising or lowering a match threshold.

Spelling

One name, twenty spellings, none of them wrong

محمد is attested in Latin script as Mohammad, Mohammed, Muhammad, Mohamad, Mohamed, Muhammed, Mohd and more. عبدالرحمن appears as Abdul Rahman, Abdulrahman, Abdurrahman, Abd al-Rahman and Abdul Rehman. The spelling on a customer record is usually the one a clerk, an airline, a resettlement agency or the person themselves chose years ago — not the output of any standard.

Structure

There is no surname to match on

Afghanistan has no general tradition of family names, and the national identity document does not normally carry one. What looks like a surname on a Western form is usually a father's name, a grandfather's name, a tribal or qawm name, a place name, or a religious title. Mapping those into first, middle and last name fields silently destroys the link between records.

Order

Name fields carry a lineage, not a first and a last name

The United Nations 1988 Sanctions List records Afghan names across four numbered name fields, alongside the name in original script and separately graded good-quality and low-quality aliases. A system that reads field 1 as a forename and field 4 as a surname is reading a patronymic chain as a Western name and will mismatch on both ends.

Titles

Honorifics enter and leave the record

Mullah, Maulavi, Haji, Qari, Hafiz, Sayed, Shah, Khan, Jan, Gul, Agha and Engineer attach to names in some records and not others. Some are religious titles, some are inherited, some are affectionate. Stripping all of them loses real identifiers; keeping all of them produces matches on the title instead of the person.

Dates

Date of birth is a weak discriminator, not a strong one

Afghan civil records run on the Solar Hijri calendar, whose year begins at Nowruz in late March. Conversions are frequently defaulted to 1 January or 1 Hamal, and sanctions entries themselves carry approximations and ranges rather than dates. Most screening configurations are tuned on the assumption that a matching date of birth confirms identity and a differing one clears it. On Afghan records, both assumptions are wrong.

Places

Place of birth and address do not normalise either

Herat, Hirat and Herāt; Kandahar and Qandahar; Kunduz, Qunduz and Kondoz; Nangarhar and Nangrahar; Pul-e-Khumri and Pol-e Khomri. District names repeat across provinces, and administrative boundaries have been redrawn. Address matching adds noise instead of confidence unless the gazetteer is built for Afghanistan.

Algorithms

Latin-alphabet phonetics encode the wrong language

Soundex, Metaphone and edit-distance measures such as Levenshtein and Jaro–Winkler operate on Latin strings and encode English sound patterns. Pashto carries retroflex and affricate consonants — څ, ځ, ښ, ږ, ړ, ڼ, ټ, ډ — with no settled Latin equivalent, and Dari leaves short vowels unwritten. Matching on the romanization instead of the source script compounds the error rather than correcting it.

Encoding

The same name is not the same string

Perso-Arabic text arrives with Arabic yeh against Farsi yeh (ي / ی), Arabic kaf against keheh (ك / ک), heh against teh marbuta, zero-width non-joiners, two families of Arabic-Indic digits, bidirectional control marks and unnormalised presentation forms. Two byte-different strings render identically on screen and compare as different names.

Collision

Common Afghan names collide with listed entries

Several of the most frequently held given names in Afghanistan appear inside entries on the UN 1988 List and the OFAC SDN List. A configuration tuned for European name distributions will treat an ordinary customer's name as a partial hit, generate an alert, and do it again at every rescreen.

03

What Ariana Nexus delivers

Six lines of work. Most engagements begin with one and add the others as the institution's Afghan-population data is brought under control.

Reference data built for your population

Afghan Name Variant Corpus

Script-anchored name records with their attested Latin renderings, patronymic chains, honorific handling, qawm and tribal flags, frequency weighting and provenance for every variant. Built for your Afghan-population data and delivered as a dated, versioned file your matching engine or data platform ingests directly.

  • Given, father and grandfather name chains held separately
  • Frequency weighting so common names are scored as common
  • Provenance recorded per variant, not per record
  • Each delivery dated, versioned and documented
Rules and normalisation

Transliteration standards mapping

A single normalisation layer that reconciles the romanization standards your records were built on with the spellings your customers actually use — so Pashto and Dari names resolve to one canonical form before any matching logic runs.

  • BGN/PCGN National Romanization System for Afghanistan (2007)
  • BGN/PCGN Romanization of Pashto (1968 system, 2017 revision)
  • ICAO Doc 9303 Part 3 machine-readable travel document transliteration
  • ALA-LC and ISO 233 for library, legal and archival records
  • Unicode normalisation, digit folding and bidi handling as a written rule set
Configuration evidence

Screening tuning and false-positive reduction

We take your existing screening configuration and test it against an Afghan-specific variant corpus, by variant class, and report what it caught and what it missed. The output is written in the format your model validation function and your examiner will ask for.

  • Control testing by variant class, not by aggregate hit rate
  • Documented false-negative exposure as well as false-positive volume
  • Threshold and algorithm recommendations with the evidence behind each
  • Written in model-documentation format for independent review
Human review

Alert adjudication and escalation review

Native Pashto and Dari analysts review level-two and level-three alerts on Afghan names and write the reasoning into the case record — in English, citing the script evidence, so the decision survives an audit years later.

  • Script-level reasoning recorded in the case file
  • Consistent disposition language across analysts
  • Escalation criteria agreed in writing before the first alert
  • No decision made on the institution's behalf — evidence and rationale only
Document review

Afghan identity document examination

Paper tazkira and e-Tazkira, Afghan passports including the machine-readable zone, driving licences, nikah nama, Se-Parcha and education documents examined for internal consistency, name-chain agreement and authenticity indicators — with a written opinion where a matter is contested.

Examination of the underlying document — Tazkira, e-Tazkira or passport — is a separate engagement.

  • Name chain reconciled across every document in a file
  • Machine-readable and visual inspection zones compared
  • Calendar conversion checked against the issuing period
  • Written expert opinion available where evidence will be tested
Existing customer files

Data remediation and entity resolution

Your existing Afghan-population records cleaned, de-duplicated and resolved to canonical entities — the remediation that makes every later screening run cheaper and every later audit shorter.

  • Duplicate resolution across name, date and place fields
  • Canonical entity keys written back to your system of record
  • Remediation decisions logged with the evidence for each
  • Run once as a programme, or standing as a monthly service
04

How an engagement runs

Six weeks is the baseline shape for a first engagement on a single system. A reference-data delivery alone moves faster; multi-system remediation takes longer.

Stage 1 · Week 1

Data read

Under a mutual non-disclosure agreement, we take a representative sample of your Afghan-population records in their current state — names, dates, places, document references — and report what is actually in the fields rather than what the schema says should be.

Stage 2 · Weeks 1–3

Variant expansion

Names are anchored to source script where it exists and reconstructed where it does not. Patronymic chains, honorifics, qawm names and place elements are parsed into separate fields and expanded into their attested variants.

Stage 3 · Weeks 3–5

Control testing

Your live screening configuration is run against the expanded corpus and against a held-back test set. Results are reported by variant class — what was caught, what was missed and at which threshold the behaviour changes.

Stage 4 · Weeks 5–6

Rule and data delivery

Normalisation rules, the name-variant corpus, threshold recommendations and the supporting documentation are delivered together, in the format your model validation function and your examiner expect.

Stage 5 · Ongoing

Standing review

List changes, corpus updates and configuration drift are reviewed on a named schedule with a named owner on both sides. Nothing about this data set stays correct on its own.

05

What the record shows

over 90%

of sanctions-screening alerts are false positives, and each one takes human review to clear.

Kim and Yang, Frontiers in Artificial Intelligence, 2024

4 name fields

carry an Afghan entry on the UN 1988 Sanctions List, alongside the original script and separately graded good-quality and low-quality aliases.

UN Security Council 1988 Committee, list structure

no surname

is recorded on the Afghan national identity document in the ordinary case. There is no general tradition of family names to match on.

Landinfo country of origin report on tazkera and Afghan ID documents

06

Standards and instruments this work is built on

Every row below is a published source. None of them is a claim about Ariana Nexus — they are the documents an examiner will hold your configuration against.

InstrumentReferenceWhat it governs
InstrumentBGN/PCGN National Romanization System for AfghanistanReference2007 systemWhat it governsThe joint United States and United Kingdom standard covering both Dari and Pashto names, designed to apply to Afghan names regardless of language origin.
InstrumentBGN/PCGN Romanization of PashtoReference1968 system, 2017 revisionWhat it governsPashto-specific romanization, including the additional consonants absent from Arabic and the handling of the definite article in names of Arabic origin.
InstrumentICAO Doc 9303, Part 3ReferenceMachine-readable travel documentsWhat it governsThe transliteration rules that govern the machine-readable zone of an Afghan passport, and the reason the name in the MRZ often differs from the name on the same page.
InstrumentUN Security Council 1988 Sanctions ListReferenceResolution 1988 (2011) regimeWhat it governsThe Afghanistan and Taliban list. Four numbered name fields, the name in original script, and aliases graded as good quality or low quality — a structure most matching engines discard on import.
InstrumentOFAC SDN List and Afghanistan general licencesReferenceE.O. 13224 · GL 14–20What it governsThe Taliban and the Haqqani Network are designated as Specially Designated Global Terrorists. General licences 14 through 20 authorise defined humanitarian, remittance, agricultural and NGO activity.
InstrumentALA-LC and ISO 233ReferenceLibrary and archival romanizationWhat it governsThe standards behind name spellings in academic, court and archival records, which frequently differ from both the passport spelling and the bank record.
InstrumentAfghan civil registrationReferenceTazkira and e-TazkiraWhat it governsIssued by the National Statistics and Information Authority. The paper tazkira carries no security elements and normally no surname; the e-Tazkira records a Latin-script name the applicant supplies themselves.
07

Where generic matching stops and this work starts

Failure modeGeneric multilingual screeningWith Afghan name data in place
Failure modeRomanization varianceGeneric multilingual screeningFuzzy distance applied to the Latin stringWith Afghan name data in placeVariants resolved to a script-anchored canonical record
Failure modeMissing surnameGeneric multilingual screeningBlank or duplicated last-name fieldWith Afghan name data in placeGiven, father and grandfather names held in separate fields
Failure modeHonorificsGeneric multilingual screeningStripped globally or kept globallyWith Afghan name data in placeClassified by type, with rules per class
Failure modeDate of birthGeneric multilingual screeningTreated as a strong confirming or clearing signalWith Afghan name data in placeWeighted as weak, with calendar provenance recorded
Failure modePlace of birthGeneric multilingual screeningFree-text comparisonWith Afghan name data in placeResolved against an Afghan administrative gazetteer
Failure modeTuning evidenceGeneric multilingual screeningAggregate hit rateWith Afghan name data in placeCatch and miss reported by variant class
Failure modeAlert reviewGeneric multilingual screeningEnglish-language analyst, no script accessWith Afghan name data in placeFirst-language review with written script-level reasoning
An Afghan flag among many United States flags on a lawn
08

Where the market says multilingual, we say Pashto and Dari.

Every enterprise screening platform advertises multilingual matching. The phrase usually means Arabic, Cyrillic and Chinese, with Persian handled as a variant of Arabic and Pashto handled not at all. Afghan names are the hardest sub-case in the set and the least represented in the training data, the gazetteers and the test suites those platforms ship.

Ariana Nexus does not compete with those platforms. We supply the layer they do not have. Our people read the source scripts, know how a Kandahari name is built differently from a Herati one, can tell a qawm name from a patronymic on sight, and can put the reasoning in writing for a regulator, an examiner or a court.

The firm is run by alumni and scholars of Cornell, the University of Chicago, the University of British Columbia and Otto-von-Guericke University Magdeburg, working alongside Afghan analysts with first-language command of the scripts. Every engagement is produced in-house: one team, one statement of work, one point of accountability. Nothing is brokered out.

Source script first

Matching decisions are anchored to the Perso-Arabic original, not to whichever romanization happened to reach your database.

Evidence you can hand over

Every recommendation arrives with the test that produced it, written for an audience that will challenge it.

Scholars, not bilinguals

Graduate degrees from Cornell, the University of Chicago and Otto-von-Guericke University Magdeburg, applied to a problem that is linguistic, legal and statistical at once.

Delivered in-house

No subcontractors, no outsourced review pool, no third-party data brokers in the chain of custody.

09

The people who do this work

Afghan name matching sits between linguistics, computation, document law and financial-crime compliance. It cannot be delivered by a bilingual speaker, and it is not delivered here by one. The people below read the source scripts, trained at the institutions named, and sign the written work that leaves the firm.

Program oversight

Hassan Ukasha

Managing Partner
B.S. Cornell University
M.P.H. Cornell University

Oversees the firm's operations and this programme. Sets the boundary conditions every identity engagement runs under, and owns the client relationship from the first data read to standing review.

Zeba Haqbani

Senior Partner
B.Sc.
University of British Columbia

Builds and runs the firm's institutional systems, technology and AI platforms, and the data engineering behind each name-variant delivery.

Hussain Ahmad

Principal
M.Eng.
Cornell University
Ph.D.
University of Chicago

Leads matching methodology — variant expansion, scoring, control test design and the model documentation that accompanies every tuning recommendation the firm makes.

Wasil Peroz

Principal
B.A.
Milli University
M.Sc.
Otto-von-Guericke University Magdeburg

Leads document and civil-registration analysis: tazkira and e-Tazkira, passports and the machine-readable zone, and the reconciliation of a name chain across every document in a file.

Maryam Safi

Principal
B.A.
Cornell University

Leads analyst review and the written record — disposition language, escalation criteria, and the reasoning that has to still make sense to an auditor several years after the alert was cleared.

10

Who this is built for

Banks, money services and fintechs

Sanctions and PEP screening, KYC onboarding, periodic rescreening and remittance corridors serving Afghan customers and the Afghan diaspora.

Screening and RegTech platforms

Afghan name data, test corpora and language expertise supplied to an existing matching engine, so the platform can claim Afghan coverage it can actually evidence.

Federal agencies and prime contractors

Biographic screening, applicant and partner vetting, and document examination where the underlying records are Afghan and the standard of proof is high.

Humanitarian and resettlement organisations

Partner and staff vetting under sanctions constraints, and beneficiary deduplication across agencies working with the same displaced population.

Law firms and courts

Identity and civil-registration evidence in immigration, asylum, criminal and family matters, with written opinion and testimony where required.

AI and data platforms

Entity resolution, knowledge-graph construction and training data where Afghan names, places and documents have to resolve correctly at scale.

11

What this service does not do

This page describes work on identity data. The limits below are part of the engagement, written into the statement of work, and they do not move.

We do not build watchlists, targeting packages or surveillance products, and we do not sell list data.
We do not route documents, data or enquiries through channels controlled by the de facto authorities in Afghanistan.
We do not teach or produce document forgery or spoofing technique. Document work runs in the direction of detection only.
We do not decide a customer's outcome, file a suspicious activity report or act as your compliance function. We supply the data, the test evidence and the written reasoning; the institution decides.
We do not give legal advice and we do not replace your sanctions counsel or your independent model validation function.

The purpose of this work is narrower than the tools it improves. Most of the people caught by a poorly tuned Afghan name rule are not the people the rule was written for. Getting the data right takes them out of the alert queue and leaves the queue to the records that belong in it.

12

Terms used on this page

Tazkira

The Afghan national identity document, issued by the National Statistics and Information Authority. The older paper form carries no security elements and normally no surname.

e-Tazkira

The electronic, biometric form of the tazkira. The Latin-script spelling of the holder's name is supplied by the applicant on the application form, which is why two siblings can hold different official spellings.

Patronymic chain

A name built from the person's given name followed by the father's and often the grandfather's given name. It is a lineage, not a surname, and it changes shape depending on who is filling in the form.

Qawm

A community, tribal or solidarity-group name that sometimes appears in the surname position on Western forms and sometimes does not appear at all.

Romanization

The conversion of a name from Perso-Arabic script into the Latin alphabet. Several competing standards exist, and most real records follow none of them.

Good-quality alias

On United Nations sanctions lists, an alias assessed as a reliable alternative name for the listed person, as distinct from a low-quality alias.

Solar Hijri calendar

The Afghan civil calendar, whose year begins at Nowruz in late March. Conversion errors are one of the most common sources of date mismatch in Afghan records.

Machine-readable zone

The two encoded lines at the foot of a passport page. Its transliteration follows ICAO Doc 9303 and frequently differs from the printed name above it.

13

Questions institutions ask before they start

Why do Afghan names have so many different English spellings?

Because no single romanization standard governs how Pashto and Dari names are written in the Latin alphabet, and because the spelling on most records was chosen by a person rather than produced by a system. Applicants supply their own Latin spelling on Afghan identity documents; airlines, banks, resettlement agencies and immigration systems each add their own. Twenty or more spellings of one name is normal, and none of them is a mistake.

Do Afghan names have surnames?

Usually not in the Western sense. Afghanistan has no general tradition of family names, and the national identity document does not normally record one. What appears in a surname field is typically a father's name, a grandfather's name, a tribal or qawm name, a place name or a religious title — and the same person may supply a different one on different forms.

Which transliteration standard should we use for Pashto and Dari names?

For a United States or United Kingdom institution the reference point is the BGN/PCGN National Romanization System for Afghanistan (2007), with the BGN/PCGN Pashto system for Pashto-specific names and ICAO Doc 9303 for anything derived from a travel document. In practice no single standard is enough, because your customer records were never built on one. The workable answer is a normalisation layer that maps the standards and the real-world spellings to one canonical record.

Why does our screening system miss Afghan names written in script?

Most matching engines compare Latin strings. When a name arrives in Perso-Arabic script it is either transliterated on import — introducing a new variant that may match nothing — or held as a string that is byte-different from a visually identical name because of yeh and kaf variants, zero-width joiners or unnormalised Unicode. Both paths produce silent false negatives.

Why do so many Afghan records show 1 January as the date of birth?

Because Afghan civil records run on the Solar Hijri calendar, whose year begins at Nowruz in late March, and because many people hold no recorded birth date at all. Conversions and intake forms default to the first day of the year. Sanctions entries themselves often carry approximate years or ranges rather than dates. Any configuration that treats date of birth as a strong confirming or clearing signal is mis-tuned for this population.

What is a tazkira, and how is it different from an e-Tazkira?

The tazkira is the Afghan national identity document, issued by the National Statistics and Information Authority. The older paper form carries no security elements and normally no surname. The e-Tazkira is the electronic biometric card; its Latin-script name is supplied by the applicant on the application form, which is why official spellings differ between members of the same family.

Can you tell whether two Afghan records are the same person?

We can tell you what the evidence supports and what it does not, in writing, with the reasoning shown. On a well-documented file that is often a clear answer. On a file with two paper tazkiras, no shared date and a name chain that changes across documents, the honest answer is a stated confidence with the specific evidence that would resolve it. We do not manufacture certainty that the records cannot carry.

How do you reduce false positives without weakening sanctions detection?

By testing rather than loosening. We build a variant corpus for your Afghan population, run your live configuration against it, and report catch and miss rates by variant class. That shows which alerts are structurally unavoidable, which come from a fixable data defect, and where the configuration is already missing true matches. Recommendations move specific rules with evidence attached; nothing moves on a global threshold alone.

Do you screen against the UN 1988 List and the OFAC SDN List?

We do not sell list data and we do not replace your screening vendor. We make your screening work on Afghan names, and we test your configuration against the way Afghan entries are actually structured on those lists — four numbered name fields, original script, and aliases graded good quality or low quality.

Is this available for Afghan languages other than Pashto and Dari?

Yes. Pashto and Dari carry most Afghan identity records and have the deepest coverage. Uzbeki, Turkmeni, Balochi and the Hazaragi variety of Dari are covered for name and place handling, and the remaining Afghan languages are available on request, scoped case by case rather than promised as standing coverage.

Who reviews the work, and what are their qualifications?

Named members of the firm, not an anonymous pool. Matching methodology sits with Hussain Ahmad (M.Eng., Cornell; Ph.D., University of Chicago). Document and civil-registration analysis sits with Wasil Peroz (M.Sc., Otto-von-Guericke University Magdeburg). Analyst review and the written record sit with Maryam Safi (B.A., Cornell). Hassan Ukasha, Managing Partner, oversees every engagement. The work is produced in-house, with no subcontractors.

How does a first engagement start?

With a scoped data read under a mutual non-disclosure agreement. You send a representative sample of your Afghan-population records; we report what is actually in the fields and what it will cost you in alerts and in missed matches. That report stands on its own whether or not the engagement continues.

14

Sources

  • 1BGN/PCGN National Romanization System for Afghanistan (2007)United States Board on Geographic Names and the Permanent Committee on Geographical Names
  • 2BGN/PCGN Romanization of Pashto, 1968 system, 2017 revisionUnited States Board on Geographic Names and the Permanent Committee on Geographical Names
  • 3ICAO Doc 9303, Machine Readable Travel Documents, Part 3International Civil Aviation Organization
  • 41988 Sanctions List and committee amendment noticesUnited Nations Security Council Committee established pursuant to resolution 1988 (2011)
  • 5Afghanistan-related general licences 14 to 20 and counter terrorism designationsOffice of Foreign Assets Control, United States Department of the Treasury
  • 6Afghanistan: Tazkera, passports and other ID documentsLandinfo, Norwegian Country of Origin Information Centre
  • 7Legal identity and civil documentation in AfghanistanUnited Nations High Commissioner for Refugees and partners
  • 8Accuracy improvement in financial sanction screening: is natural language processing the solution? (2024), doi:10.3389/frai.2024.1374323Kim, S. and Yang, S., Frontiers in Artificial Intelligence

Start with the data read

Send a representative sample of your Afghan-population records under a mutual non-disclosure agreement. We report what is in the fields, what your current configuration does with it, and what that is costing you in both directions.

Initiate an engagement