# Benchmarking on Tasks That Matter: Dataset Selection for Preserving Model Rankings
> Ficha editorial pública de Research IA. Estado: Lectura primaria completa. La interpretación editorial no sustituye la fuente primaria.

- Página canónica: https://luiseduardodemiguel.com/research-ia/papers/benchmarking-on-tasks-that-matter-dataset-selection-for-preserving-model
- Fuente primaria: https://arxiv.org/abs/2606.27997
- Versión leída: v1
- Fuente comprobada: 2026-08-19 · lectura primaria completa; extracción editorial automatizada, revisión humana pendiente
- Autores: Rostislav Gusev, Alexey Zaytsev
- Fecha del corte: JUNIO 2026; KDD 2026.
- Área: EVALUACIÓN

## Tesis y contexto

Estudia cómo seleccionar datasets que preserven rankings de modelos. El listado de arXiv lo marca aceptado en KDD 2026.

- Problema: Evaluar modelos en demasiados datasets es caro, pero reducir mal el set distorsiona rankings.
- Por qué importa: Muy útil para equipos que comparan modelos de forma continua.

## Evidencia reportada

- **reported-result**: Figure 2 shows how ranking preservation improves as subset size increases. [localizador](https://arxiv.org/html/2606.27997#S6)
- **reported-result**: Across methods, we observe the largest gains at small k (typically k\leq 10 ), followed by diminishing returns as k approaches 20 . [localizador](https://arxiv.org/html/2606.27997#S6)
- **reported-result**: Diversity-oriented methods (Farthest-First with cosine or Euclidean distance) and K-Means generally outperform Random when the dataset representation captures ranking-relevant structure. [localizador](https://arxiv.org/html/2606.27997#S6)
- **reported-result**: Oracle-style representations derived from a posteriori information (e.g., rank features) provide an upper bound on achievable performance for geometry-based selection. [localizador](https://arxiv.org/html/2606.27997#S6)

## Lectura y límite

- Método: La lectura de 5. Benchmarking Protocol describe la intervención y su construcción: In this section, we describe an empirical evaluation protocol used to compare datasets’ subset selection strategies in terms of preserving the global ranking of models formally defined in the previous section and the representation of datasets in the considered domains. We adopt a unified and reusable benchmarking protocol designed to evaluate datasets subset selection strategies with respect to their ability to preserve the model ranking induced by the full benchmark. The protocol explicitly accounts for variability in dataset and model availability and provides statistically robust estimates of…
- Límite: La lectura primaria permite comprobar método y resultados en el HTML, pero no convierte sus conclusiones en validación independiente. La ficha no demuestra transferencia fuera de los datasets, modelos, herramientas y condiciones descritos en 6. Results.
- Confianza editorial: Media
- Limitación: El cierre de la fuente señala: Overall, this work provides a practical framework for rank-preserving benchmark reduction. Instead of treating benchmark subset choice as an informal or convenience-driven decision, we pose it as a measurable dataset selection problem with explicit rank-preservation metrics and statistical significance comparisons. The resulting protocol can be reused in a multi-dataset benchmark with more than 50 tasks where evaluation costs are substantial and…
- Limitación: La ficha no demuestra transferencia fuera de los datasets, modelos, herramientas y condiciones descritos en 6. Results.

## Localizadores de evidencia
- [Fuente primaria · canonical](https://arxiv.org/abs/2606.27997): tipo abstract
- [HTML · lectura completa](https://arxiv.org/html/2606.27997): tipo abstract
- [Método · 5. Benchmarking Protocol](https://arxiv.org/html/2606.27997#S5): tipo section
- [Evaluación · 6. Results](https://arxiv.org/html/2606.27997#S6): tipo section
- [Cierre · 7. Conclusions and discussion](https://arxiv.org/html/2606.27997#S7): tipo section

## Próxima prueba

- ¿La propuesta mejora model selection frente a la línea base actual?
- Métrica: Comparar la métrica principal de la fuente junto con calidad, coste, latencia y tasa de errores.
- Regla de parada: Parar si no aparece una mejora reproducible o si aumenta el riesgo, la complejidad o el coste sin compensación.

## Recursos reproducibles
- [https://github.com/hangover137/efficient_benchmarking](https://github.com/hangover137/efficient_benchmarking)
- [the following issues](https://github.com/arXiv/html_feedback/issues)
- [list of packages that need conversion](https://github.com/brucemiller/LaTeXML/wiki/Porting-LaTeX-packages-for-LaTeXML)
- [developer contributions](https://github.com/brucemiller/LaTeXML/issues)

## Enlaces relacionados

- [VAKRA](https://luiseduardodemiguel.com/research-ia/markdown/papers/vakra)
- [VibeLifeBench](https://luiseduardodemiguel.com/research-ia/markdown/papers/vibelifebench)
- [KnowHal](https://luiseduardodemiguel.com/research-ia/markdown/papers/knowhal)