---
title: "LMArena (Chatbot Arena) (FR)"
source: "https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)"
wiki: "systems-analysis.info/int"
article: "LMArena_(Chatbot_Arena)_(FR)"
language: "fr"
categories:
  - "Category:French"
  - "Category:Large language models"
  - "Category:LLM benchmarks"
  - "Category:Machine learning"
revision_id: 3702
wiki_created_at: 2026-09-06T23:24:42Z
wiki_modified_at: 2026-09-06T23:24:42Z
downloaded_at: 2026-09-07T22:58:24Z
---

# LMArena (Chatbot Arena) (FR)

**Arena** (jusqu'au 28 janvier 2026 : **LMArena** (*Large Model Arena*), auparavant : **Chatbot Arena**) est une plateforme web ouverte et participative (*crowdsourcing*) destinée à l'évaluation et au classement comparatif des [grands modèles de langage](https://systems-analysis.info/int/Grands_mod%C3%A8les_de_langage "Grands modèles de langage") (LLM) et des modèles multimodaux (texte, image, vidéo) d'après les préférences humaines réelles. Elle repose sur des comparaisons anonymes par paires (*blind battles*) et sur un système de classement de type Elo ; sur cette base sont publiés des classements publics (*leaderboards*) transparents, considérés comme l'un des bancs d'essai (*benchmarks*) indépendants les plus reconnus pour les modèles d'IA de pointe (*frontier*).<sup>[\[1\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_note-chiang2024-1)[\[2\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_note-hello_2025-2)</sup>

La plateforme est née en 2023 comme projet de recherche académique de l'organisation **LMSYS Org** (Large Model Systems Organization) à l'université de Californie à Berkeley (Sky Computing Lab). En septembre 2024, le projet a pris son indépendance (« graduation ») en migrant vers son propre domaine, lmarena.ai<sup>[\[3\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_note-new_site_2024-3)</sup>. En avril 2025, la société indépendante **Arena Intelligence Inc.** a été créée<sup>[\[4\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_note-new_beta_2025-4)</sup> ; le 21 mai 2025, elle a levé **100 millions de dollars** lors d'un tour d'amorçage (*seed*), à une valorisation de 600 millions de dollars (principaux investisseurs : a16z et UC Investments)<sup>[\[5\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_note-seed_prn-5)[\[6\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_note-tc_seed-6)</sup>. Le 6 janvier 2026, la société a bouclé un tour de série A de **150 millions de dollars**, portant sa valorisation post-money à **1,7 milliard de dollars**<sup>[\[7\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_note-series_a_2026-7)</sup>. Le 28 janvier 2026 a eu lieu le changement de marque final : la plateforme a été rebaptisée **Arena** et a migré vers le domaine <a href="https://arena.ai/" class="external text" rel="nofollow">arena.ai</a><sup>[\[8\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_note-arena_rebrand_2026-8)</sup>.

## Histoire

La plateforme a été lancée en avril 2023 sous le nom de **Chatbot Arena**, comme projet de recherche de l'organisation LMSYS Org (Large Model Systems Organization) à l'université de Californie à Berkeley (Sky Computing Lab). Elle est devenue l'un des premiers outils d'évaluation participative des grands modèles de langage au moyen de comparaisons anonymes par paires (*blind battles*) et d'un système de classement Elo fondé sur les préférences réelles des utilisateurs.

- **24 avril 2023** — lancement technique de Chatbot Arena.
- **3 mai 2023** — lancement public officiel et publication du premier classement.
- **2023** — publication des premiers jeux de données ouverts : 33 000 dialogues par paires (juillet) et **LMSYS‑Chat‑1M** (septembre, environ 1 million de dialogues réels).
- **1er mars 2024** — publication de la politique officielle de la plateforme et formalisation de sa mission comme système d'évaluation ouvert et piloté par la communauté.
- **27 juin 2024** — ajout de la prise en charge des images et début de l'extension aux tâches multimodales.
- **20 septembre 2024** — prise d'indépendance (« graduation ») : migration vers le domaine indépendant **lmarena.ai**.
- **Fin 2024 – printemps 2025** — lancement d'arènes spécialisées (Arena-Hard, WebDev Arena, RepoChat Arena, Style/Sentiment Control, etc.).
- **17 avril 2025** — constitution officielle en société indépendante, **Arena Intelligence Inc.**, et lancement d'une version bêta de la plateforme rénovée sous la marque LMArena.
- **21 mai 2025** — annonce de la création de la société et levée de fonds d'amorçage de **100 millions de dollars** (valorisation : 600 millions de dollars).
- **31 juillet 2025** — publication d'un jeu de données ouvert de **140 000** dialogues récents de la Text Arena.
- **6 janvier 2026** — levée de **150 millions de dollars** lors d'un tour de série A, portant la valorisation post-money à **1,7 milliard de dollars**.
- **Janvier 2026** — lancement de la **Video Arena** (prise en charge complète de la modalité vidéo).
- **28 janvier 2026** — changement de marque final : la plateforme devient **Arena** et migre vers le domaine **arena.ai**.

En mars 2026, Arena (arena.ai) compte plus de 5 millions d'utilisateurs actifs mensuels répartis dans plus de 150 pays, a accumulé des dizaines de millions de votes et demeure l'un des outils indépendants les plus reconnus pour l'évaluation des modèles d'IA de pointe d'après les préférences humaines réelles.

## Fonctionnement de l'évaluation

L'utilisateur saisit une requête et reçoit deux réponses de modèles anonymes choisis au hasard (« A » et « B »), puis vote pour la meilleure réponse (ou signale une égalité ou une réponse insatisfaisante). Le classement est fondé sur le modèle statistique de **Bradley-Terry** (une régression logistique sur les préférences par paires), dont le principe est proche du système de classement Elo<sup>[\[1\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_note-chiang2024-1)</sup>. La plateforme publie un **Arena Score** et des intervalles de confiance, et applique une repondération (*re-weighting*) pour maintenir l'impartialité en cas d'échantillonnage non uniforme<sup>[\[9\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_note-policy-9)</sup>.

**Transparence et ouverture.** Les pipelines initiaux d'évaluation et de classement sont disponibles en open source dans le dépôt **FastChat**<sup>[\[10\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_note-10)</sup> ; une partie des données brutes est publiée périodiquement à des fins de vérification et de recherche (par ex., la publication de 140 000 dialogues en juillet 2025)<sup>[\[9\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_note-policy-9)[\[11\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_note-opendata_2025-11)</sup>. Selon la *FAQ* et les avertissements sur la page d'accueil, les requêtes des utilisateurs peuvent être transmises aux fournisseurs de modèles et partiellement publiées à des fins de recherche — il est conseillé de ne pas soumettre de données sensibles<sup>[\[12\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_note-12)[\[13\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_note-13)</sup>.

**Règles de sélection et d'échantillonnage.** Les classements incluent des modèles accessibles au public (poids ouverts/API publique/service accessible au public). Pour fiabiliser une évaluation, au moins **1 000** votes sont généralement requis ; au moins **20 %** des batailles ont lieu uniquement entre des modèles publics ; la probabilité d'échantillonnage augmente avec le rang et l'incertitude, et la régression avec repondération garantit l'impartialité des scores finaux<sup>[\[9\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_note-policy-9)</sup>.

**Métriques automatiques et contrôle du style.** Pour accélérer l'évaluation et réduire les effets des préférences stylistiques, des techniques auxiliaires sont utilisées : *MT-Bench* (LLM-as-a-judge)<sup>[\[14\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_note-14)</sup>, *Arena-Hard* (génération automatique de questions difficiles)<sup>[\[15\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_note-arena_hard-15)</sup>, ainsi que le **Style/Sentiment Control** (modélisation et « correction » de l'effet du ton/des émotions sur les préférences)<sup>[\[16\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_note-sentiment_control-16)</sup>. Pour *Arena-Hard-Auto*, une très forte concordance avec les votes humains « en direct » a été observée (jusqu'à **≈98,6 %** dans des conditions contrôlées)<sup>[\[17\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_note-17)</sup>.

## Arènes et domaines d'évaluation

La plateforme a évolué pour devenir un ensemble d'« arènes » spécialisées par type de tâches :

- **Text Arena** — dialogues/tâches générales, classement principal<sup>[\[18\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_note-text_stats-18)</sup>.
- **Vision Arena** — modèles multimodaux (compréhension d'image/vidéo, analyse visuelle)<sup>[\[19\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_note-19)</sup>.
- **Text-to-Image** et **Image Edit** — génération et édition d'images (y compris le cas de *nano-banana*)<sup>[\[20\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_note-tti_page-20)[\[21\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_note-nanobanana_blog-21)</sup>.
- **Text‑/Image‑to‑Video** — génération de vidéos<sup>[\[22\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_note-22)</sup>.
- **WebDev Arena** — création d'applications web à partir de descriptions<sup>[\[23\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_note-webdev_arena-23)</sup>.
- **RepoChat Arena** — tâches d'ingénierie IA sur du code/des dépôts<sup>[\[24\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_note-repochat_arena-24)</sup>.
- **Search Arena** — modèles connectés à la recherche web ; lancée initialement en avril 2025 (legacy), puis intégrée au site principal, accompagnée d'un jeu de données et d'une publication<sup>[\[25\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_note-25)[\[26\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_note-26)[\[27\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_note-27)</sup>.
- **BiomedArena.AI** — évaluation spécifique au domaine pour les tâches biomédicales (en partenariat avec DataTecnica)<sup>[\[28\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_note-28)</sup>.

## Applications et influence

- **Vitrine de l'industrie.** Les plus grands fournisseurs (OpenAI, Anthropic, Google, etc.) testent et présentent régulièrement leurs modèles sur la plateforme ; les médias spécialisés la décrivent comme une référence importante<sup>[\[6\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_note-tc_seed-6)[\[29\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_note-29)</sup>. Dans une communication de la session Industrie de NAACL 2025, l'évaluation Elo de Chatbot Arena est qualifiée de « **gold industry-standard** »<sup>[\[30\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_note-30)</sup>.
- **Tests de pré-lancement.** La politique autorise les aperçus anonymes de modèles non publiés, avec notification à la communauté et publication ultérieure des évaluations publiques après le lancement ; un minimum de ≈1 000 votes est requis pour la stabilisation<sup>[\[9\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_note-policy-9)</sup>.
- **Faits marquants.** Au printemps 2025, le modèle anonyme *Llama-4 Maverick-03-26-Experimental* a fait l'objet de discussions (incident concernant sa comparaison avec des versions publiques), attirant une large attention médiatique et entraînant des mises à jour des règles et de la communication<sup>[\[31\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_note-31)[\[32\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_note-32)</sup>. En août 2025, « nano-banana » s'est révélé être **Gemini 2.5 Flash Image** et a pris la tête des classements visuels<sup>[\[21\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_note-nanobanana_blog-21)[\[20\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_note-tti_page-20)</sup>.

## Limites et critiques

Malgré son ampleur et sa popularité, l'approche présente des limites :

- **Subjectivité et effets de style.** Les préférences de vote dépendent du ton et du style de la réponse ; l'équipe met en œuvre le **Style/Sentiment Control** pour découpler le « style » du « contenu »<sup>[\[16\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_note-sentiment_control-16)</sup>.
- **Non-représentativité de l'audience.** Le noyau actif est composé de passionnés de technologie et de développeurs ; pour les scénarios spécifiques à un domaine, des arènes spécialisées sont créées (Search, WebDev, Biomed, etc.)<sup>[\[33\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_note-33)</sup>.
- **Vulnérabilité aux manipulations et aux biais.** Des recherches de 2025 montrent qu'en l'absence de protections strictes, des stratégies de « truquage des votes » avec des centaines ou des milliers de votes sont possibles ; cependant, la collaboration entre les chercheurs et LMArena a conduit à la mise en place de mesures de protection (CAPTCHA/connexion/protection anti-bot/détection d'anomalies) et à l'augmentation du « coût de l'attaque »<sup>[\[34\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_note-34)[\[35\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_note-35)[\[36\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_note-36)</sup>.
- **Critiques méthodologiques.** L'étude *The Leaderboard Illusion* (avril 2025) souligne des facteurs systémiques et institutionnels susceptibles de fausser la compétition ; LMArena a publié une réponse détaillée et tient à jour un *changelog* public de sa méthodologie<sup>[\[37\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_note-37)[\[38\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_note-38)[\[39\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_note-changelog-39)</sup>.

## Liens externes

- <a href="https://arena.ai/" class="external text" rel="nofollow">Site officiel d'Arena (anciennement lmarena.ai)</a>
- <a href="https://news.lmarena.ai/" class="external text" rel="nofollow">Blog, politiques et mises à jour</a>
- <a href="https://lmsys.org/" class="external text" rel="nofollow">Site du groupe de recherche LMSYS (incubateur original du projet)</a>

## Bibliographie

- Chiang, W.-L. et al. (2024). *Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference*. <a href="https://arxiv.org/abs/2403.04132" class="external text" rel="nofollow">arXiv:2403.04132</a>.
- Zheng, L. et al. (2023). *Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena*. <a href="https://arxiv.org/abs/2306.05685" class="external text" rel="nofollow">arXiv:2306.05685</a>.
- Li, T. et al. (2024). *From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline*. <a href="https://arxiv.org/abs/2406.11939" class="external text" rel="nofollow">arXiv:2406.11939</a>.
- Ameli, S.; Zhuang, S.; Stoica, I.; Mahoney, M. W. (2024). *A Statistical Framework for Ranking LLM-Based Chatbots*. <a href="https://arxiv.org/abs/2412.18407" class="external text" rel="nofollow">arXiv:2412.18407</a>.
- Boubdir, M. et al. (2023). *Elo Uncovered: Robustness and Best Practices in Language Model Evaluation*. <a href="https://arxiv.org/abs/2311.17295" class="external text" rel="nofollow">arXiv:2311.17295</a>.
- Huang, J. Y.; Shen, Y.; Wei, D.; Broderick, T. (2025). *Dropping Just a Handful of Preferences Can Change Top Large Language Model Rankings*. <a href="https://arxiv.org/abs/2508.11847" class="external text" rel="nofollow">arXiv:2508.11847</a>.
- Xu, Y.; Ruis, L.; Rocktäschel, T.; Kirk, R. (2025). *Investigating Non-Transitivity in LLM-as-a-Judge*. <a href="https://arxiv.org/abs/2502.14074" class="external text" rel="nofollow">arXiv:2502.14074</a>.
- Li, H. et al. (2024). *LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods*. <a href="https://arxiv.org/abs/2412.05579" class="external text" rel="nofollow">arXiv:2412.05579</a>.
- Zheng, L. et al. (2024). *LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset*. <a href="https://arxiv.org/abs/2309.11998" class="external text" rel="nofollow">arXiv:2309.11998</a>.
- Dubois, Y. et al. (2024). *Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators*. <a href="https://arxiv.org/abs/2404.04475" class="external text" rel="nofollow">arXiv:2404.04475</a>.
- Singh, S. et al. (2025). *The Leaderboard Illusion*. <a href="https://arxiv.org/abs/2504.20879" class="external text" rel="nofollow">arXiv:2504.20879</a>.
- Min, R.; Pang, T.; Du, C.; Liu, Q.; Cheng, M.; Lin, M. (2025). *Improving Your Model Ranking on Chatbot Arena by Vote Rigging*. <a href="https://arxiv.org/abs/2501.17858" class="external text" rel="nofollow">arXiv:2501.17858</a>.

## Références

1.  <span id="cite_note-chiang2024-1">↑ <sup>[1.0](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_ref-chiang2024_1-0)</sup> <sup>[1.1](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_ref-chiang2024_1-1)</sup> Chiang, W.-L. et al. « Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference ». *arXiv:2403.04132*, 2024.</span>
2.  <span id="cite_note-hello_2025-2">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_ref-hello_2025_2-0) « Hello from LMArena: The Community Platform for Exploring Frontier AI ». *LMArena Blog*, 23 juin 2025.</span>
3.  <span id="cite_note-new_site_2024-3">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_ref-new_site_2024_3-0) « Announcing a New Site for Chatbot Arena ». *LMSYS Blog*, 20 septembre 2024.</span>
4.  <span id="cite_note-new_beta_2025-4">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_ref-new_beta_2025_4-0) « LMArena is Growing to Support our Community Platform ». *LMArena Blog*, 17 avril 2025.</span>
5.  <span id="cite_note-seed_prn-5">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_ref-seed_prn_5-0) « LMArena Secures \$100M in Seed Funding… ». *PR Newswire*, 21 mai 2025. <a href="https://www.prnewswire.com/news-releases/lmarena-secures-100m-in-seed-funding-to-bring-scientific-rigor-to-ai-reliability-302462025.html" class="external autonumber" rel="nofollow">[1]</a></span>
6.  <span id="cite_note-tc_seed-6">↑ <sup>[6.0](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_ref-tc_seed_6-0)</sup> <sup>[6.1](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_ref-tc_seed_6-1)</sup> Wiggers, K. « LM Arena… lands \$100M ». *TechCrunch*, 21 mai 2025. <a href="https://techcrunch.com/2025/05/21/lm-arena-the-organization-behind-popular-ai-leaderboards-lands-100m/" class="external autonumber" rel="nofollow">[2]</a></span>
7.  <span id="cite_note-series_a_2026-7">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_ref-series_a_2026_7-0) « Arena Intelligence Closes \$150M Series A at \$1.7B Valuation ». *Arena Intelligence Blog*, 6 janvier 2026. <a href="https://news.lmarena.ai/series-a-2026/" class="external autonumber" rel="nofollow">[3]</a></span>
8.  <span id="cite_note-arena_rebrand_2026-8">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_ref-arena_rebrand_2026_8-0) « LMArena is now Arena ». *Arena Blog*, 28 janvier 2026. <a href="https://arena.ai/blog/lmarena-is-now-arena/" class="external autonumber" rel="nofollow">[4]</a></span>
9.  <span id="cite_note-policy-9">↑ <sup>[9.0](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_ref-policy_9-0)</sup> <sup>[9.1](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_ref-policy_9-1)</sup> <sup>[9.2](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_ref-policy_9-2)</sup> <sup>[9.3](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_ref-policy_9-3)</sup> *LMArena Leaderboard Policy*. *LMArena Blog*, éd. 8 septembre 2025. <a href="https://news.lmarena.ai/policy/" class="external autonumber" rel="nofollow">[5]</a></span>
10. <span id="cite_note-10">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_ref-10) lm‑sys/FastChat (GitHub). <a href="https://github.com/lm-sys/FastChat" class="external autonumber" rel="nofollow">[6]</a></span>
11. <span id="cite_note-opendata_2025-11">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_ref-opendata_2025_11-0) Y. Song. « A Deep Dive into Recent Arena Data ». *LMArena Blog*, 31 juillet 2025. <a href="https://news.lmarena.ai/opendata-july2025/" class="external autonumber" rel="nofollow">[7]</a></span>
12. <span id="cite_note-12">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_ref-12) *FAQ*. *LMArena*. <a href="https://lmarena.ai/faq" class="external autonumber" rel="nofollow">[8]</a></span>
13. <span id="cite_note-13">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_ref-13) Page d'accueil de *LMArena* (avertissement sur la publication possible des données et leur transmission aux fournisseurs). <a href="https://lmarena.ai/" class="external autonumber" rel="nofollow">[9]</a></span>
14. <span id="cite_note-14">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_ref-14) Zheng, L. et al. « Judging LLM‑as‑a‑Judge with MT‑Bench and Chatbot Arena ». *arXiv:2306.05685*, 2023. <a href="https://arxiv.org/abs/2306.05685" class="external autonumber" rel="nofollow">[10]</a></span>
15. <span id="cite_note-arena_hard-15">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_ref-arena_hard_15-0) Li, T. et al. « From Crowdsourced Data to High‑Quality Benchmarks: Arena‑Hard and BenchBuilder Pipeline ». *arXiv:2406.11939*, 2024. <a href="https://arxiv.org/abs/2406.11939" class="external autonumber" rel="nofollow">[11]</a></span>
16. <span id="cite_note-sentiment_control-16">↑ <sup>[16.0](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_ref-sentiment_control_16-0)</sup> <sup>[16.1](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_ref-sentiment_control_16-1)</sup> « Does Sentiment Matter Too? Introducing Sentiment Control ». *LMArena Blog*, 22 avril 2025. <a href="https://news.lmarena.ai/sentiment-control/" class="external autonumber" rel="nofollow">[12]</a></span>
17. <span id="cite_note-17">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_ref-17) Li, T. et al. « From Crowdsourced Data… » *arXiv:2406.11939* (tableaux de concordance). <a href="https://arxiv.org/abs/2406.11939" class="external autonumber" rel="nofollow">[13]</a></span>
18. <span id="cite_note-text_stats-18">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_ref-text_stats_18-0) *Text Arena (English)*. *LMArena*. <a href="https://lmarena.ai/leaderboard/text/english" class="external autonumber" rel="nofollow">[14]</a></span>
19. <span id="cite_note-19">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_ref-19) *Vision Arena*. *LMArena*, mis à jour le 2 septembre 2025. <a href="https://lmarena.ai/leaderboard/vision" class="external autonumber" rel="nofollow">[15]</a></span>
20. <span id="cite_note-tti_page-20">↑ <sup>[20.0](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_ref-tti_page_20-0)</sup> <sup>[20.1](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_ref-tti_page_20-1)</sup> *Text‑to‑Image Arena*. *LMArena*, mis à jour le 25 août 2025. <a href="https://lmarena.ai/leaderboard/text-to-image" class="external autonumber" rel="nofollow">[16]</a></span>
21. <span id="cite_note-nanobanana_blog-21">↑ <sup>[21.0](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_ref-nanobanana_blog_21-0)</sup> <sup>[21.1](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_ref-nanobanana_blog_21-1)</sup> « Nano Banana (Gemini 2.5 Flash Image): Try it on LMArena ». *LMArena Blog*, 27 août 2025. <a href="https://news.lmarena.ai/nano-banana/" class="external autonumber" rel="nofollow">[17]</a></span>
22. <span id="cite_note-22">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_ref-22) *Text‑to‑Video* et *Image‑to‑Video Leaderboards*. *LMArena*, août 2025. <a href="https://lmarena.ai/leaderboard/text-to-video" class="external autonumber" rel="nofollow">[18]</a> <a href="https://lmarena.ai/leaderboard/image-to-video" class="external autonumber" rel="nofollow">[19]</a></span>
23. <span id="cite_note-webdev_arena-23">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_ref-webdev_arena_23-0) « WebDev Arena: A Live LLM Leaderboard for Web App Development ». *LMArena Blog*, 10 mars 2025. <a href="https://lmarena.github.io/blog/2025/webdev-arena/" class="external autonumber" rel="nofollow">[20]</a></span>
24. <span id="cite_note-repochat_arena-24">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_ref-repochat_arena_24-0) « RepoChat Arena: A Live Benchmark for AI Software Engineers ». *LMArena Blog*, 9 avril 2025. <a href="https://lmarena.github.io/blog/2025/repochat-arena/" class="external autonumber" rel="nofollow">[21]</a></span>
25. <span id="cite_note-25">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_ref-25) « Introducing the Search Arena ». *LMArena Blog*, 14 avril 2025. <a href="https://news.lmarena.ai/search-arena/" class="external autonumber" rel="nofollow">[22]</a></span>
26. <span id="cite_note-26">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_ref-26) « Search Arena & What We're Learning About Human Preference ». *LMArena Blog*, 23 juillet 2025. <a href="https://news.lmarena.ai/search-arena-update/" class="external autonumber" rel="nofollow">[23]</a></span>
27. <span id="cite_note-27">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_ref-27) Frick, E. et al. « Search Arena: Analyzing Search‑Augmented LLMs ». *arXiv:2506.05334*, 2025. <a href="https://arxiv.org/abs/2506.05334" class="external autonumber" rel="nofollow">[24]</a></span>
28. <span id="cite_note-28">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_ref-28) « Introducing BiomedArena.AI ». *LMArena Blog*, 19 août 2025. <a href="https://news.lmarena.ai/introducing-biomedarena/" class="external autonumber" rel="nofollow">[25]</a></span>
29. <span id="cite_note-29">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_ref-29) Google. « Gemma 3… », 12 mars 2025 (lien vers les résultats de LMArena). <a href="https://blog.google/technology/developers/gemma-3/" class="external autonumber" rel="nofollow">[26]</a></span>
30. <span id="cite_note-30">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_ref-30) Spangher, L. et al. « Chatbot Arena Estimate… ». *NAACL Industry*, 2025. <a href="https://aclanthology.org/2025.naacl-industry.77.pdf" class="external autonumber" rel="nofollow">[27]</a></span>
31. <span id="cite_note-31">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_ref-31) « Meta's experimental Llama 4 model briefly topped AI leaderboard… ». *The Register*, 7 avril 2025. <a href="https://www.theregister.com/2025/04/07/meta_llama_arena/" class="external autonumber" rel="nofollow">[28]</a></span>
32. <span id="cite_note-32">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_ref-32) Clarifications/posts officiels de LMArena sur X concernant l'incident (avril 2025). <a href="https://x.com/lmarena_ai/status/1776881220809552182" class="external autonumber" rel="nofollow">[29]</a></span>
33. <span id="cite_note-33">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_ref-33) « Search Arena & What We're Learning… ». *LMArena Blog*, 23 juillet 2025. <a href="https://news.lmarena.ai/search-arena-update/" class="external autonumber" rel="nofollow">[30]</a></span>
34. <span id="cite_note-34">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_ref-34) Min, R. et al. « Improving Your Model Ranking on Chatbot Arena by Vote Rigging ». *arXiv:2501.17858*, 2025. <a href="https://arxiv.org/abs/2501.17858" class="external autonumber" rel="nofollow">[31]</a></span>
35. <span id="cite_note-35">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_ref-35) Huang, Y. et al. « Exploring and Mitigating Adversarial Manipulation of Voting‑Based Leaderboards ». *arXiv:2501.07493*, 2025. <a href="https://arxiv.org/abs/2501.07493" class="external autonumber" rel="nofollow">[32]</a></span>
36. <span id="cite_note-36">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_ref-36) « Hundreds of rigged votes can skew… ». *Fast Company*, 6 février 2025. <a href="https://www.fastcompany.com/91273226/rigged-votes-ai-model-rankings-chatbot-arena" class="external autonumber" rel="nofollow">[33]</a></span>
37. <span id="cite_note-37">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_ref-37) Singh, S. et al. « The Leaderboard Illusion ». *arXiv:2504.20879*, 2025. <a href="https://arxiv.org/abs/2504.20879" class="external autonumber" rel="nofollow">[34]</a></span>
38. <span id="cite_note-38">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_ref-38) « Our Response to 'The Leaderboard Illusion' ». *LMArena Blog*, 9 mai 2025. <a href="https://news.lmarena.ai/our-response/" class="external autonumber" rel="nofollow">[35]</a></span>
39. <span id="cite_note-changelog-39">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(FR)#cite_ref-changelog_39-0) *Leaderboard Changelog*. *LMArena Blog*, entrées d'août 2025. <a href="https://news.lmarena.ai/leaderboard-changelog/" class="external autonumber" rel="nofollow">[36]</a></span>
