Services
AI & LLM Language Services
Language models learn from language and are judged by it. We bring two decades of professional translation to Turkish and English data for AI teams: written and reviewed by native speakers, with the care literary work demands.
At a glance
- Turkish and English, by native speakers
- Training data, evaluation, annotation and localization
- Paid pilot before production
- Two-stage review on every batch
- Your guidelines, your tools or simple file formats
- NDA and data-handling terms on request
What we do
Language work for models, done by linguists.
Training data writing
Prompts, responses, instructions and multi-turn dialogues written by native Turkish and English speakers, including long-form and literary styles.
Dataset and benchmark localization
English→Turkish translation of instruction sets, evaluation benchmarks and product text, adapted to Turkish culture rather than carried over word for word.
Human evaluation and preference data
Rating model outputs for accuracy, fluency, tone and cultural fit, and side-by-side comparisons for preference and RLHF data.
MT evaluation and post-editing
Post-editing of machine translation, error annotation with frameworks such as MQM, and quality reviews of translated datasets.
Annotation and classification
Entities, sentiment, intent, topic and toxicity labels for Turkish and English text, following your guidelines.
Red-teaming and safety review
Adversarial prompts and harm review in Turkish, with attention to cultural, social and legal context in Türkiye.
Transcription
Accurate Turkish and English transcription of audio and video for speech and multimodal datasets.
Why it matters
Where Turkish trips models up.
Turkish is spoken by tens of millions of people, yet it is still a lower-resource language for most models. These are the places where careful human work makes the biggest difference.
Rich morphology
One Turkish word can carry what English says in a sentence. Tokenizers and metrics built for English often miss this.
Register
The difference between sen and siz, and between plain and formal Turkish, changes how a reply lands with a reader.
İ and ı
Turkish has dotted and dotless i in both cases. Careless lower- and upper-casing quietly corrupts text and evaluation.
Culture
Idioms, humour, names, holidays and politeness rules that a literal translation flattens or gets wrong.
How we work
Pilot first, then production.
- 1
Scoping
A short call about the task, your guidelines, volumes, formats and deadlines.
- 2
Paid pilot
A small batch to calibrate against your guidelines and agree what “good” looks like.
- 3
Production
Each item is done by one linguist and checked by a second; we track agreement and flag guideline gaps.
- 4
Delivery
JSONL, CSV or directly in your annotation platform, with a short quality report.
Confidentiality
Your data stays your data.
- Non-disclosure and data-processing agreements on request
- Access limited to the linguists on your project
- Your data is used only for your project, never to train anything of our own
- Files deleted at the end of the project on request
- We can work entirely inside your own tools and accounts
FAQ
Questions from AI teams
What do you do for AI teams?
We write, translate, evaluate and annotate Turkish and English data for language models, such as instruction and dialogue data, localized benchmarks, preference rankings and safety reviews. Native speakers do the work and a second linguist reviews it.
Can you work in our annotation platform?
Yes. We can work inside your own tools and accounts, or deliver JSONL or CSV files, whichever suits your pipeline.
How do you keep data quality consistent?
We start with a paid pilot to calibrate against your guidelines. In production one linguist does each item and a second reviews it; we track agreement and raise unclear guideline cases with you instead of guessing.
Will you sign our NDA or data-processing agreement?
Yes. We review and sign non-disclosure and data-processing agreements, use your data only for your project and delete files at the end on request.
Tell us about your model.
Share the task, language direction and volume. We’ll suggest a pilot and send a quote.
Discuss a project