The workLabels for training and validation, made in the language.
One principle governs everything below: annotation happens in the language — never the industry shortcut of translating to English and labeling the translation, which destroys exactly the dialect, register, and cultural signal the label exists to capture.
Classification & labeling.
Intent, topic, sentiment, and safety or policy labels on Afghan-language text — judged in the register and dialect band the text was written in.
Translation-quality annotation.
Segment-level quality labels and MQM-style error typing — accuracy, fluency, terminology, register, and cultural-validity errors, each categorized, not just flagged.
RLHF & preference data.
Rankings, ratings, and rubric-based preference judgments over Afghan-language model outputs; instruction-following evaluation by people who can tell compliant from fluent.
Span & structure annotation.
Named entities, relations, and linking — with Afghan naming conventions and transliteration variants handled as the first-class problem they are.
Speech annotation.
Transcription, dialect-band tagging, and speaker attributes — the labeled substrate behind the firm's ASR/TTS reference work.
Safety & cultural-validity labeling.
The labels behind the firm's hallucination and trust-and-safety audit work — harm, appropriateness, and cultural-validity judgments that require exactly the qualification a crowd cannot supply.