Abstract
As large language models are deployed across multilingual environments, the benchmarks used to evaluate their safety remain largely designed for high-resource, English-dominant contexts. Trust & Safety systems increasingly rely on automated classifiers and generative models to moderate harmful content across dozens of languages, yet the tools used to assess them often fail to capture linguistic variation and culturally specific harm. This paper presents a structured review of 37 benchmarks relevant to multilingual safety evaluation, 27 of them safety benchmarks, published between 2019 and 2026. Using a nine-dimension taxonomy spanning linguistic authenticity, cultural grounding, and evaluation transparency, we analyze how benchmarks construct evaluation datasets and report safety performance. We identify recurring design patterns: reliance on translated English prompts rather than native-language data, aggregate metrics that obscure cross-language variation, and limited use of locally grounded harm categories or community-informed evaluation. These patterns show that broader language coverage alone does not ensure benchmarks capture how harm is expressed across cultural contexts. We therefore propose a design framework emphasizing native-authored data, disaggregated reporting, locally grounded harm taxonomies, symmetric evaluation of false positives, coverage of dialect and code-switching, and testing under deployment-relevant conditions.

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.
Copyright (c) 2026 Alisar Mustafa, Cherry Wu
