At an AI Festa keynote on October 6, 2026, Google DeepMind research scientist Lora Aroyo called for AI evaluations that account for local languages and cultural context. The distinction matters: translating a test question changes its language, but does not by itself make the evaluation reflect the context in which people use that language.

Why translation is not the same as local evaluation

A benchmark is a set of tests used to compare how AI systems perform. If its questions begin in one language and are translated, the wording changes; the cultural assumptions built into the questions may remain. Aroyo argued that evaluation material should be developed with the target language and cultural context in mind, rather than relying on translation alone.

That means testing more than fluency. A system can produce natural-sounding language while missing local meanings or norms. Aroyo described cultural errors across language and cultural norms, local religions and myths, history, and geopolitics.

What Global MMLU found

A 2024 study of MMLU reported that 28% of questions in the studied set required culturally sensitive knowledge. The same study found that 84.9% of its geography questions focused on North America or Europe. That second figure applies to geography questions, not to the entire benchmark.

The study also found that model rankings changed when researchers assessed culturally sensitive questions separately from the full set. In other words, a model’s overall position can depend on which questions an evaluation counts and how those questions are categorized.

Global MMLU extends the evaluation across 42 languages and labels questions as culturally sensitive or culturally agnostic. Its authors describe using compensated professional and community annotators to check translation quality, alongside a review of cultural bias in the original dataset.

The reported Malay and Tamil examples

Aroyo cited accuracy figures of 50% for English questions translated into Malay and 40% for English questions translated into Tamil. Those figures were presented as examples in her reported remarks, rather than as general performance rates for AI systems.

The examples underline the distinction at the heart of the debate: translating a question is one part of multilingual testing, while evaluating cultural context requires attention to what the question assumes and what knowledge it asks for. Global MMLU’s separate cultural subsets offer one way to examine that distinction alongside language coverage.

Four kinds of cultural context

Aroyo identified four areas for culturally attentive evaluation: language and cultural norms; local religions and myths; history; and geopolitics. Together, they point beyond whether a model can answer in a language to whether its answers handle the local context that gives those words meaning.