# Challenges and Recommendations for LLMs-as-a-Judge in Multilingual Settings and Low-Resource Languages
> Ficha editorial pública de Research IA. Estado: Lectura primaria completa. La interpretación editorial no sustituye la fuente primaria.

- Página canónica: https://luiseduardodemiguel.com/research-ia/papers/challenges-and-recommendations-for-llms-as-a-judge-in-multilingual-setti
- Fuente primaria: https://arxiv.org/html/2607.02235v1
- Versión leída: v1
- Fuente comprobada: 2026-08-19 · lectura primaria completa; extracción editorial automatizada, revisión humana pendiente
- Autores: A. Seza Doğruöz, Xixian Liao, Verena Blaschke, Jakob Prange, Senyu Li, David Ifeoluwa Adelani
- Fecha del corte: 2 JULIO 2026.
- Área: MULTIMODAL

## Tesis y contexto

Revisión crítica de LLM-as-a-Judge en evaluación multilingüe y lenguas de bajo recurso. De 650 papers que mencionan LLM-as-a-Judge, solo 33 se centran en estos escenarios; detecta sobreconfianza, resultados inconsistentes y uso habitual de un solo juez.

- Problema: Se está usando evaluación automática en idiomas donde el juez puede no ser competente.
- Por qué importa: Afecta a productos europeos/multilingües y evaluación en español, catalán, euskera, árabe, lenguas africanas, etc.

## Evidencia reportada

- **reported-result**: Among the 19/33 papers (58%) that include at least one low-resource language, coverage is still often skewed toward higher-resource ones: Sitaram et al. [localizador](https://arxiv.org/html/2607.02235#S4)
- **reported-result**: However, the check against human ratings is performed on an existing English-only benchmark of human-annotated dialogue responses. [localizador](https://arxiv.org/html/2607.02235#S4)
- **reported-result**: 2024 report that human-LLM agreement drops for direct assessment, particularly for Bengali and Odia. [localizador](https://arxiv.org/html/2607.02235#S4)
- **reported-result**: Fu and Liu 2025 report an average Fleiss’ Kappa of approximately 0.3 across 25 languages, with consistency being particularly poor in low-resource languages, and find that neither multilingual training nor model scale directly improves this result. [localizador](https://arxiv.org/html/2607.02235#S4)

## Lectura y límite

- Método: La lectura de 3 Literature Search and Annotation Methodology describe la intervención y su construcción: LLM-as-a-Judge is used broadly to cover both evaluator-oriented and annotator-oriented uses of LLMs in the literature. Evaluator -oriented settings use LLM judgments to measure, compare, or validate items (e.g., model responses, system outputs, retrieved evidence, or benchmark examples). Given task-specific context (e.g., an instruction, source text, candidate response, reference answer, rubric, or label definitions), the judge produces an evaluative output (e.g., a score, label, ranking, preference judgment, or textual assessment). Annotator -oriented settings use LLMs to produce labels, metadata, explanations,…
- Límite: La lectura primaria permite comprobar método y resultados en el HTML, pero no convierte sus conclusiones en validación independiente. La ficha no demuestra transferencia fuera de los datasets, modelos, herramientas y condiciones descritos en 4 LLM-as-a-Judge in Multilingual Research.
- Confianza editorial: Media
- Limitación: El cierre de la fuente señala: Consider real-world relevance and representativeness beyond language. The amount of available language data is not the only relevant dimension with respect to using LLM-as-a-Judge in low-resource settings, and failures on different dimensions should be distinguished. For instance, linguistic competence and cultural competence can be two independent dimensions. Watts et al. 2024 show that LLM evaluators agree less with humans on evaluating responses with…
- Limitación: La ficha no demuestra transferencia fuera de los datasets, modelos, herramientas y condiciones descritos en 4 LLM-as-a-Judge in Multilingual Research.

## Localizadores de evidencia
- [Fuente primaria · canonical](https://arxiv.org/html/2607.02235v1): tipo abstract
- [HTML · lectura completa](https://arxiv.org/html/2607.02235): tipo abstract
- [Método · 3 Literature Search and Annotation Methodology](https://arxiv.org/html/2607.02235#S3): tipo section
- [Evaluación · 4 LLM-as-a-Judge in Multilingual Research](https://arxiv.org/html/2607.02235#S4): tipo section
- [Cierre · 5 Discussion and Recommendations](https://arxiv.org/html/2607.02235#S5): tipo section
- [HTML · fuente navegable](https://arxiv.org/abs/2607.02235v1): tipo abstract

## Próxima prueba

- ¿La propuesta mejora QA multilingüe frente a la línea base actual?
- Métrica: Comparar la métrica principal de la fuente junto con calidad, coste, latencia y tasa de errores.
- Regla de parada: Parar si no aparece una mejora reproducible o si aumenta el riesgo, la complejidad o el coste sin compensación.

## Recursos reproducibles
- [Apache 2.0 license](https://github.com/acl-org/acl-anthology/blob/master/LICENSE)
- [370911e](https://github.com/acl-org/acl-anthology/tree/370911ef38556af764c17d456be9fce2d477b0bd)
- [CODEOFCONDUCT at multilingual counterspeech generation: A context-aware model…](https://aclanthology.org/2025.mcg-1.5/)
- [Evaluating the quality of benchmark datasets for low-resource languages: A case…](https://aclanthology.org/2025.gem-1.41/)

## Enlaces relacionados

- [MMDiff](https://luiseduardodemiguel.com/research-ia/markdown/papers/mmdiff)
- [MBA](https://luiseduardodemiguel.com/research-ia/markdown/papers/mbabench)
- [VibeLifeBench](https://luiseduardodemiguel.com/research-ia/markdown/papers/vibelifebench)