TextPashto and Dari text data
Written data for supervised fine-tuning, chat, translation and classification, written natively in Pashto and Dari rather than machine-translated from English.
- Instruction and chat data (SFT)
- Prompts and ideal responses written by native speakers across the tasks your model must perform.
- Multi-turn conversations
- Realistic dialogues in everyday, health, legal, public-service, education and commercial settings.
- Domain text corpora
- Licensed, consented writing in the registers your model must handle: formal, colloquial, broadcast and online.
- Parallel corpora
- Pashto–English, Dari–English and Pashto–Dari sentence pairs for machine translation, aligned and reviewed.
- Text annotation
- Named entities, intent and slots, sentiment, topic, toxicity and safety labels, under written guidelines.
- OCR ground truth
- Printed and handwritten Pashto and Dari in Naskh and Nastaliq styles, transcribed line by line.
- Romanized and code-switched text
- Pashto and Dari as people actually type them, in Latin script and mixed with English, normalized and labeled.
SpeechPashto and Dari speech data
Spoken data for voice assistants, speech-to-speech models and spoken-language understanding, recorded with informed consent.
- Conversational speech
- Two-speaker dialogues and spontaneous monologues, with speaker panels balanced by dialect, region, gender and age.
- Spoken prompts
- Questions and commands for voice assistants and speech-to-speech models, spoken the way users actually ask.
- Phone, wideband and far-field capture
- 8 kHz telephone, 16 kHz wideband and room-distance recordings for the channel your product runs on.
- Transcription
- Verbatim or clean transcripts with timestamps, speaker turns and non-speech events, in normalized Pashto or Dari script.
- Speech annotation
- Speaker diarization, dialect and accent labels, emotion and intent tags on audio.
- Speech translation pairs
- Pashto or Dari audio aligned to English translations for speech translation models.
PreferencePashto and Dari preference data for RLHF and DPO
Human judgments that teach a model which response is better in Pashto and Dari: the data behind reward models, RLHF and DPO.
- Pairwise preferences
- Native raters choose between two model responses and write the reason, in English or the source language.
- Rankings and rubric ratings
- Responses scored on helpfulness, accuracy, safety, dialect fit and cultural fit.
- Rewrites
- Native speakers correct or rewrite a response to produce the preferred answer for DPO and SFT.
- Safety preference data
- Refusals and safe completions judged by raters who understand the local context of harm.
- Multi-turn conversation ratings
- Whole conversations rated turn by turn, including tone and register across the exchange.
- Spoken-response preference
- Voice model replies judged for pronunciation, dialect and naturalness by native listeners.
- Rubrics for reinforcement learning
- Expert-written grading rubrics that score Pashto and Dari responses for rubric-based rewards.