Urdu Outside the Guardrails: Why LLMs Miss Hate Speech in Unfamiliar Scripts
Five well-known models — from GPT-4o to Llama-3.1 — reach different verdicts on the same text when it is written in Urdu rather than in English translation: the discrepancy reaches up to a third of cases, and some clearly hostile content remains unlabeled. An analysis of nine years of WOAH publications showed that not a single dedicated paper there has been devoted to this language.








