# Pluralis v0.1: Multicultural, Multimodal and Multilingual AI Risk Benchmark
> Ficha editorial pública de Research IA. Estado: Lectura primaria completa. La interpretación editorial no sustituye la fuente primaria.

- Página canónica: https://luiseduardodemiguel.com/research-ia/papers/pluralis-v0-1-multicultural-multimodal-and-multilingual-ai-risk-benchmar
- Fuente primaria: https://arxiv.org/abs/2607.06196
- Versión leída: v1
- Fuente comprobada: 2026-08-19 · lectura primaria completa; extracción editorial automatizada, revisión humana pendiente
- Autores: Alicia Parrish, Rajat Shinde, Sanket Badhe, Xinyi Bai, Sree Bhargavi Balija, Hua-Rong Chu, Emilio Ferrara, Armstrong Foundjem, Rajat Ghosh, Aakash Gupta, Xuanli He, Ong Chen Hui, Minji Jung, Madhangi Karimanal, Faiza Khan Khattak, Boryoung Kim, Eugenia Kim, Liliya Lavitas, Seok Min Lim, Victor Lu, Jim Moirangthem, Dhivya Nagasubramanian, Deepak Pandita, Sita Rajagopal, Geetha Raju, Evgeniia Razumovskaia, Aravind Reddy, Federico Ricciuti, Nobin Sarwar, Sungpil Shin, Sunayana Sitaram, Snehal Thorat, Tharindu Cyril Weerasooriya, Jasmijn Bastings, Joachim Baumann, Kongtao Chen, Murali Emani, Mariya Hendriksen, Jiho Jin, Jun Seong Kim, Younghoon Ko, Alicja Kwasniewska, Minjae Lee, Tom Wei-cyuan Lin Kashyap Ramanandula Manjusha, Junho Myung, Junyeong Park, Roma Patel, Shyam Ratan, Sudarsun Santhiappan, Priyanka Suresh, Tuesday, Ksheeraj Sai Vepuri Laura Amortegui-Ordonez, Claire Dennis, Minsuk Kahng, Chris Knotz, Alice Oh, Balaraman Ravindran, Soojung Ryu William Bartholomew, Hiwot Tesfaye, Lora Aroyo
- Fecha del corte: 7 JULIO 2026.
- Área: EVALUACIÓN · SEGURIDAD · MULTIMODAL

## Tesis y contexto

Benchmark de seguridad multimodal construido nativamente desde seis países de Asia-Pacífico y ocho idiomas. Evalúa casos donde texto e imagen parecen inocuos por separado, pero su combinación genera riesgos legales o culturales locales.

- Problema: Los benchmarks globales promedian culturas, idiomas y jurisdicciones, ocultando fallos regionales.
- Por qué importa: La seguridad de un producto internacional no puede depender únicamente de datasets occidentales traducidos. Pluralis incorpora 6.448 prompts y separa daño universal de adecuación cultural.

## Evidencia reportada

- **reported-result**: Scores are mapped into ordinal grade bands ( Good , Fair , Poor ) using thresholds calibrated against our reference baseline model. [localizador](https://arxiv.org/html/2607.06196#S6)
- **reported-result**: Table 5 and Figure 4 show the score’s tier, and the parenthetical bracket reports the range of bands the score crosses under judge-level uncertainty (the lower bound corresponds to the judge most lenient on this axis; the upper bound, the strictest). [localizador](https://arxiv.org/html/2607.06196#S6)
- **reported-result**: Broadly, we observe that model performance is better both in terms of whether the text makes sense grammatically and in terms of its accuracy for English and Traditional Chinese (the Non-English language in the Taiwan dataset) compared to the other seven languages in Pluralis . [localizador](https://arxiv.org/html/2607.06196#S6)
- **reported-result**: To better understand the underlying cause of SUT failures, both in terms of their safety violations and their culturally inappropriateness, we stratified results on dev set using the bottom-up cultural taxonomy. [localizador](https://arxiv.org/html/2607.06196#S6)

## Lectura y límite

- Método: La lectura de 3 Pluralis Dataset Creation Methods describe la intervención y su construcción: We constructed Pluralis through a coordinated, multi-regional effort involving paid annotators and volunteer researchers across multiple countries. To ensure cultural accuracy and methodological consistency, regional linguistic and safety experts managed the end-to-end data collection and validation process for each locale. Crucially, Pluralis follows a culture-first methodology. Unlike English-centric benchmarks that merely translate existing datasets created originally in English into target locales, Pluralis safety hazards were conceptualized natively by regional experts to capture localized legal, religious,…
- Límite: La lectura primaria permite comprobar método y resultados en el HTML, pero no convierte sus conclusiones en validación independiente. La ficha no demuestra transferencia fuera de los datasets, modelos, herramientas y condiciones descritos en 6 Preliminary Insights.
- Confianza editorial: Media
- Limitación: El cierre de la fuente señala: The massive variance and high false-negative rates we observed highlight that evaluating cultural alignment remains an open and complex challenge. Pluralis exposes the urgent need for much more reliable, efficient-to-develop multilingual evaluators and provides a framework for community innovation to deliver that technology. We call upon the research community to utilize this foundation to advance the science of multilingual, multicultural evaluation to…
- Limitación: La ficha no demuestra transferencia fuera de los datasets, modelos, herramientas y condiciones descritos en 6 Preliminary Insights.

## Localizadores de evidencia
- [Fuente primaria · canonical](https://arxiv.org/abs/2607.06196): tipo abstract
- [HTML · lectura completa](https://arxiv.org/html/2607.06196): tipo abstract
- [Método · 3 Pluralis Dataset Creation Methods](https://arxiv.org/html/2607.06196#S3): tipo section
- [Evaluación · 6 Preliminary Insights](https://arxiv.org/html/2607.06196#S6): tipo section
- [Cierre · 7 Limitations and Future Work](https://arxiv.org/html/2607.06196#S7): tipo section

## Próxima prueba

- ¿La propuesta mejora evaluación de VLMs frente a la línea base actual?
- Métrica: Comparar la métrica principal de la fuente junto con calidad, coste, latencia y tasa de errores.
- Regla de parada: Parar si no aparece una mejora reproducible o si aumenta el riesgo, la complejidad o el coste sin compensación.

## Recursos reproducibles
- [the following issues](https://github.com/arXiv/html_feedback/issues)
- [list of packages that need conversion](https://github.com/brucemiller/LaTeXML/wiki/Porting-LaTeX-packages-for-LaTeXML)
- [developer contributions](https://github.com/brucemiller/LaTeXML/issues)

## Enlaces relacionados

- [VAKRA](https://luiseduardodemiguel.com/research-ia/markdown/papers/vakra)
- [VibeLifeBench](https://luiseduardodemiguel.com/research-ia/markdown/papers/vibelifebench)
- [KnowHal](https://luiseduardodemiguel.com/research-ia/markdown/papers/knowhal)