Evaluating Multilingual Safety Benchmarks for Low-Resource Languages in the Majority World
Cover of the Journal of Online Trust & Safety, Volume 3, Issue 3, September 2026. The journal's logo—an open teal book with an upturned open palm beneath a floating key—sits left of the title in a white band across the top. The center is a dark teal photographic background of blurred light streaks and glowing bokeh circles, over which the issue's nine article titles are listed in white bold type: "Digital Intersectionality and Marginalization in the Majority World"; "Deepfake Abuse and Gendered Digital Creative Violence: Feminist AI Interventions from Mexico"; "Deepfakes in the Global South: Understanding Stakeholder Perceptions, Harms, and Governance Challenges in Sri Lanka, India, and Bangladesh"; "Evaluating Multilingual Safety Benchmarks for Low-Resource Languages in the Majority World"; "Platform Oversight in Practice: How Language Shapes Procedure, Not Outcome, in Meta's Oversight Board"; "Acceptance and Use of the National Digital Identity System in Indonesia"; "Security without Safety: Queering Cybersecurity in the Age of Digital Transnational Repression"; "Shaping Queerphobia: How X Amplifies Hate Speech against Iranian Queer Activists"; and "'Our Content Should Be Treated Differently': Race, Gender, Language, and the Experiences of Quechua Women on Social Media." A matching white band at the bottom carries the Creative Commons BY-NC-SA badge and the Stanford Tech Impact and Policy Center, Freeman Spogli Institute wordmark on the left, and "Volume 3, Issue 3 / September 2026 / ISSN: 2770-3142" on the right. Three muted gray-green accent bars sit at the lower left of the top band and at the upper and lower right of the photographic band.
PDF

Keywords

multilingual safety
large language models
benchmark evaluation
low-resource languages
content moderation

Categories

How to Cite

Mustafa, A., & Wu, C. . (2026). Evaluating Multilingual Safety Benchmarks for Low-Resource Languages in the Majority World. Journal of Online Trust and Safety, 3(3). https://doi.org/10.54501/jots.v3i3.330

Abstract

As large language models are deployed across multilingual environments, the benchmarks used to evaluate their safety remain largely designed for high-resource, English-dominant contexts. Trust & Safety systems increasingly rely on automated classifiers and generative models to moderate harmful content across dozens of languages, yet the tools used to assess them often fail to capture linguistic variation and culturally specific harm. This paper presents a structured review of 37 benchmarks relevant to multilingual safety evaluation, 27 of them safety benchmarks, published between 2019 and 2026. Using a nine-dimension taxonomy spanning linguistic authenticity, cultural grounding, and evaluation transparency, we analyze how benchmarks construct evaluation datasets and report safety performance. We identify recurring design patterns: reliance on translated English prompts rather than native-language data, aggregate metrics that obscure cross-language variation, and limited use of locally grounded harm categories or community-informed evaluation. These patterns show that broader language coverage alone does not ensure benchmarks capture how harm is expressed across cultural contexts. We therefore propose a design framework emphasizing native-authored data, disaggregated reporting, locally grounded harm taxonomies, symmetric evaluation of false positives, coverage of dialect and code-switching, and testing under deployment-relevant conditions.

https://doi.org/10.54501/jots.v3i3.330
PDF
Creative Commons License

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.

Copyright (c) 2026 Alisar Mustafa, Cherry Wu