Urdu Outside the Guardrails: Why LLMs Miss Hate Speech in Unfamiliar Scripts

18 September 202612 views

Five well-known models — from GPT-4o to Llama-3.1 — reach different verdicts on the same text when it is written in Urdu rather than in English translation: the discrepancy reaches up to a third of cases, and some clearly hostile content remains unlabeled. An analysis of nine years of WOAH publications showed that not a single dedicated paper there has been devoted to this language.

Urdu Outside the Guardrails: Why LLMs Miss Hate Speech in Unfamiliar Scripts

In brief: the problem isn't the language, it's the script

Safety filters for large language models are typically evaluated in English. This gives rise to a convenient illusion: if a model reliably catches insults and calls to violence in English text, it will handle other languages roughly the same way. Urdu — a language with roughly 246 million speakers, the tenth most spoken in the world — barely figures in such testing. And this gap, as it turns out, has a measurable cost.

That's the subject of a recent paper with the identifier arXiv:2608.24191. The title "Ghaib in Translation" plays on the word ghaib, which in Urdu means "hidden, unseen": it refers to harm that goes unnoticed. The authors are Fawzia Zehra (Fuzzy) Kara-Isitt, Sonal Khosla, and Stephen Swift. The first version appeared on August 25, 2026, with a revised version on September 9 of the same year; the paper falls under cs.CL and cs.AI, DOI — 10.48550/arXiv.2608.24191.

How the experiment was designed

Models and data

The team ran five popular models through the same classification scenario: GPT-4o, Claude Sonnet 4.5, Gemini 2.5 Flash, Qwen-2.5, and Llama-3.1. The dataset consists of six datasets covering four different forms the language takes: Urdu in Nastaliq script, Urdu in Latin script (Roman Urdu), English, and mixed Urdu-English texts where the languages switch within a single message.

An important detail: what was tested was not so much quality as such, but the stability of the decision. The same meaning was presented to the model twice — in the original script and in English translation. If the filter only triggers in the second case, then the protection is tied not to the content but to the letters it happens to be written in.

Two metrics

The authors work with two simple indicators. The first is the discrepancy rate: how often the label changes between the original and the translation. The second is called "Missed-in-Urdu": content that was flagged as harmful in the English version but passed as harmless in the original script. In essence, this is a pure miss — exactly what moderation exists to prevent.

What the numbers showed

Across the five datasets where the text was written in Urdu script, the spread in results between classification of the original and the translation ranged from 15.9% to 31.6%. The best result was Gemini 2.5 Flash's, the worst Qwen-2.5's. In other words, in the worst case roughly every third filter decision depended on which alphabet the very same text was submitted in.

The share of missed harm ranged from 2.4% to 9.9%, with a median of 4.3%. At first glance that seems small, but it's worth translating into human scale: if a platform with an audience of millions of users processes tens of thousands of messages a day, even 4% means hundreds of pieces of toxic content slipping through daily. And these aren't borderline cases like sarcasm — they're material that the same model confidently flags as harmful once it's translated.

The general pattern the authors note: the smaller and more open the model, the more both metrics drop off. Flagship closed systems hold up more steadily, but even they don't have a zero gap.

Nine years of literature — and not a single paper

A separate line of research is bibliographic. The authors went through all 205 papers from nine editions of ALW/WOAH via the ACL Anthology API and found not a single one devoted specifically to Urdu. This doesn't mean the topic has never been studied — but in the corpus of materials from the main workshop on harm evaluation, it doesn't exist as a distinct area. It's a vicious circle: no benchmarks — no publications — no pressure on developers — again no benchmarks.

Why this happens

The paper doesn't give a definitive answer, but the mechanics are clear from general reasoning. Nastaliq is a cursive script where the form of a letter depends on its position in the word, which means tokenizers tuned to Latin script cut such text into unfavorable fragments. There is orders of magnitude less Urdu data in training corpora than English. Annotation for reinforcement learning (including on safety) is also gathered predominantly by native English speakers. On top of that, there's a convenient compromise: translate the input into English, run the check, return the answer. It saves resources but adds an intermediary — and that's what produces the discrepancy.

What product teams should do about it

  • Don't rely on translation as a proxy. If the filter makes its decision based on the English version, the difference needs to be measured, not swept under the rug.
  • Make the discrepancy a metric. It's useful to regularly count how many decisions change when the script changes: it's a cheap way to find blind spots without new annotation.
  • Test on the original script. Nastaliq, Roman Urdu, and code-switching are three different modes, and good results in one guarantee nothing in the other two.
  • Don't adjust thresholds blindly. Raising sensitivity on one script easily turns into a stream of false positives on another.
  • Take the audience into account. 246 million speakers is not a niche language but a significant share of users on any major platform.

The conclusion the authors lead to sounds dry but to the point: safety guarantees today are distributed unevenly across scripts. As long as that's the case, "the model passed safety tests" is a claim that always warrants clarification: in what language and in what letters those tests were written.

Frequently asked questions

Related materials

All materials
Urdu Outside the Guardrails: Why LLMs Miss Hate Speech in Unfamiliar Scripts