---
title: "LMArena (Chatbot Arena) (PL)"
source: "https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)"
wiki: "systems-analysis.info/int"
article: "LMArena_(Chatbot_Arena)_(PL)"
language: "pl"
categories:
  - "Category:Large language models"
  - "Category:LLM benchmarks"
  - "Category:Machine learning"
  - "Category:Polish"
revision_id: 3710
wiki_created_at: 2026-09-06T23:24:49Z
wiki_modified_at: 2026-09-06T23:24:49Z
downloaded_at: 2026-09-07T22:58:28Z
---

# LMArena (Chatbot Arena) (PL)

**Arena** (do 28 stycznia 2026 roku — **LMArena** (*Large Model Arena*), wcześniej — **Chatbot Arena**) — otwarta, oparta na crowdsourcingu platforma internetowa służąca do oceny i porównawczego rankingowania dużych modeli językowych (LLM) oraz modeli multimodalnych (tekst, obraz, wideo) na podstawie rzeczywistych preferencji ludzkich. Podstawę platformy stanowią anonimowe porównania parowe (blind battles) modeli oraz system rankingowy Elo; na ich podstawie publikowane są przejrzyste, publiczne tablice liderów (leaderboardy), uznawane za jeden z najbardziej miarodajnych niezależnych benchmarków frontier-modeli sztucznej inteligencji.<sup>[\[1\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_note-chiang2024-1)[\[2\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_note-hello_2025-2)</sup>

Platforma powstała w 2023 roku jako akademicki projekt badawczy organizacji **LMSYS Org** (Large Model Systems Organization) przy Uniwersytecie Kalifornijskim w Berkeley (Sky Computing Lab). We wrześniu 2024 roku projekt „ukończył inkubację" i przeniósł się na samodzielną domenę lmarena.ai („Graduation")<sup>[\[3\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_note-new_site_2024-3)</sup>. W kwietniu 2025 roku utworzono niezależną spółkę **Arena Intelligence Inc.**, a 21 maja 2025 roku pozyskano rundę seed w wysokości **100 mln USD** (wycena 600 mln USD, główni inwestorzy — a16z i UC Investments)<sup>[\[4\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_note-seed_prn-4)[\[5\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_note-tc_seed-5)</sup>. 6 stycznia 2026 roku spółka zamknęła rundę Series A o wartości **150 mln USD** przy wycenie post-money wynoszącej **1,7 mld USD**<sup>[\[6\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_note-series_a_2026-6)</sup>. 28 stycznia 2026 roku przeprowadzono finalny rebranding — platforma przyjęła nazwę **Arena** i przeniosła się na domenę arena.ai<sup>[\[7\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_note-arena_rebrand_2026-7)[\[8\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_note-new_beta_2025-8)</sup>.

## Historia

Platforma została uruchomiona w kwietniu 2023 roku pod nazwą **Chatbot Arena** jako projekt badawczy organizacji LMSYS Org (Large Model Systems Organization) przy Uniwersytecie Kalifornijskim w Berkeley (Sky Computing Lab). Stała się jednym z pierwszych narzędzi crowdsourcingowej oceny dużych modeli językowych za pomocą anonimowych porównań parowych (blind battles) oraz systemu rankingowego Elo, opartego na rzeczywistych preferencjach użytkowników.

- **24 kwietnia 2023** — techniczne uruchomienie Chatbot Arena.
- **3 maja 2023** — oficjalne publiczne wydanie i publikacja pierwszego leaderboardu.
- **2023** — wydanie pierwszych otwartych datasetów: 33 tys. dialogów parowych (lipiec) oraz **LMSYS‑Chat‑1M** (wrzesień, ok. 1 mln rzeczywistych dialogów).
- **1 marca 2024** — publikacja oficjalnej polityki platformy, formalizacja misji jako otwartego, community-driven systemu oceny.
- **27 czerwca 2024** — dodanie obsługi obrazów i rozpoczęcie rozszerzenia o zadania multimodalne.
- **20 września 2024** — „Graduation": przeniesienie na samodzielną domenę **lmarena.ai**.
- **Koniec 2024 — wiosna 2025** — uruchomienie wyspecjalizowanych aren (Arena-Hard, WebDev Arena, RepoChat Arena, Style/Sentiment Control i in.).
- **17 kwietnia 2025** — oficjalna inkorporacja jako niezależna spółka **Arena Intelligence Inc.**, uruchomienie wersji beta zaktualizowanej platformy pod marką LMArena.
- **21 maja 2025** — ogłoszenie powołania spółki i pozyskanie rundy seed w wysokości **100 mln USD** (wycena — 600 mln USD).
- **31 lipca 2025** — publikacja otwartego datasetu zawierającego **140 tys.** najnowszych dialogów Text Arena.
- **6 stycznia 2026** — pozyskanie rundy Series A w wysokości **150 mln USD** przy wycenie post-money **1,7 mld USD**.
- **Styczeń 2026** — uruchomienie **Video Arena** (pełna obsługa modalności wideo).
- **28 stycznia 2026** — finalny rebranding: platforma przyjęła nazwę **Arena** i przeniosła się na domenę **arena.ai**.

Do marca 2026 roku Arena (arena.ai) obsługuje ponad 5 mln miesięcznie aktywnych użytkowników ze 150+ krajów, zgromadziła dziesiątki milionów głosów i pozostaje jednym z najbardziej miarodajnych niezależnych narzędzi oceny frontier-modeli AI na podstawie rzeczywistych preferencji ludzkich.

## Jak działa ocena

Użytkownik wprowadza zapytanie i otrzymuje dwie odpowiedzi od losowo wybranych anonimowych modeli („A" i „B"), po czym głosuje na lepszą odpowiedź (lub odnotowuje remis/niezadowolenie). Ranking oparty jest na statystycznym modelu **Bradleya–Terry'ego** (regresja logistyczna na preferencjach parowych), intuicyjnie zbliżonym do Elo<sup>[\[1\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_note-chiang2024-1)</sup>. Platforma publikuje **Arena Score** wraz z przedziałami ufności, a także stosuje korekty próbkowania (re‑weighting) w celu zachowania nieobciążoności przy nierównomiernym samplingiem<sup>[\[9\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_note-policy-9)</sup>.

**Przejrzystość i otwartość.** Źródłowe pipeline'y oceny i rankingowania są otwarte w repozytorium **FastChat**<sup>[\[10\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_note-10)</sup>; okresowo publikowane są fragmenty surowych danych do weryfikacji i badań (np. wydanie 140K dialogów w lipcu 2025)<sup>[\[9\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_note-policy-9)[\[11\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_note-opendata_2025-11)</sup>. Zgodnie z *FAQ* i ostrzeżeniami na stronie głównej, zapytania użytkowników mogą być ujawniane dostawcom modeli i częściowo publikowane w celach badawczych — nie należy przesyłać danych wrażliwych<sup>[\[12\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_note-12)[\[13\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_note-13)</sup>.

**Zasady doboru i samplingowania.** Do leaderboardów włączane są modele ogólnodostępne (otwarte wagi/publiczne API/publiczny serwis). Dla stabilizacji oceny wymagane jest zazwyczaj ≥**1000** głosów; co najmniej **20%** starć — wyłącznie między modelami publicznymi; prawdopodobieństwo samplingowania rośnie wraz z rankingiem i niepewnością, a regresja z przeważaniem zapewnia nieobciążoność końcowych ocen<sup>[\[9\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_note-policy-9)</sup>.

**Auto‑metryki i kontrola stylu.** W celu przyspieszenia oceny i ograniczenia efektów „stylowych" preferencji stosowane są pomocnicze metody: *MT‑Bench* (LLM‑as‑a‑judge)<sup>[\[14\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_note-14)</sup>, *Arena‑Hard* (autogeneracja trudnych pytań)<sup>[\[15\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_note-arena_hard-15)</sup> oraz **Style/Sentiment Control** (modelowanie i „leczenie" wpływu tonu/emocji na preferencje)<sup>[\[16\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_note-sentiment_control-16)</sup>. Dla *Arena‑Hard‑Auto* odnotowano bardzo wysoką zgodność z „żywymi" głosami ludzkimi (do **≈98,6%** w warunkach kontrolowanych)<sup>[\[17\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_note-17)</sup>.

## Areny i domeny oceny

Platforma rozwinęła się w zestaw „aren" według typów zadań:

- **Text Arena** — ogólne dialogi/zadania, główna tabela<sup>[\[18\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_note-text_stats-18)</sup>.
- **Vision Arena** — modele multimodalne „tekst→obraz/wideo/analiza obrazów"<sup>[\[19\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_note-19)</sup>.
- **Text‑to‑Image** i **Image Edit** — generowanie i edycja obrazów (m.in. przypadek *nano‑banana*)<sup>[\[20\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_note-tti_page-20)[\[21\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_note-nanobanana_blog-21)</sup>.
- **Text‑/Image‑to‑Video** — generowanie wideo<sup>[\[22\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_note-22)</sup>.
- **WebDev Arena** — budowanie aplikacji webowych na podstawie opisów<sup>[\[23\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_note-webdev_arena-23)</sup>.
- **RepoChat Arena** — zadania z zakresu inżynierii AI dotyczące kodu/repozytoriów<sup>[\[24\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_note-repochat_arena-24)</sup>.
- **Search Arena** — modele z podłączeniem do wyszukiwania internetowego; początkowo uruchomiona w kwietniu 2025 (legacy), następnie przeniesiona na główną stronę, towarzyszy jej dataset i publikacja<sup>[\[25\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_note-25)[\[26\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_note-26)[\[27\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_note-27)</sup>.
- **BiomedArena.AI** — dziedzinowa ocena dla zadań biomedycznych (partnerstwo z DataTecnica)<sup>[\[28\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_note-28)</sup>.

## Zastosowanie i wpływ

- **Witryna przemysłowa.** Najwięksi dostawcy (OpenAI, Anthropic, Google i in.) regularnie testują i prezentują modele na LMArena; branżowe media opisują platformę jako ważny punkt odniesienia<sup>[\[5\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_note-tc_seed-5)[\[29\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_note-29)</sup>. W branżowej publikacji NAACL‑2025 ocena Elo Chatbot Arena została nazwana „**gold industry‑standard**"<sup>[\[30\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_note-30)</sup>.
- **Testowanie przedwydaniowe.** Polityka dopuszcza anonimowe podglądy „niewydanych" modeli z powiadomieniem społeczności i późniejszą publikacją ocen publicznych po wydaniu; minimum ≈1000 głosów dla stabilizacji<sup>[\[9\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_note-policy-9)</sup>.
- **Znane epizody.** Wiosną 2025 roku dyskutowano o anonimowym modelu *Llama‑4 Maverick‑03‑26‑Experimental* (incydent dotyczący porównania z wersjami publicznymi), co przyciągnęło szeroką uwagę prasy i sprowokowało aktualizacje zasad/komunikacji<sup>[\[31\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_note-31)[\[32\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_note-32)</sup>. W sierpniu 2025 „nano‑banana" okazał się **Gemini 2.5 Flash Image** i zajął czołowe pozycje w arenach wizualnych<sup>[\[21\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_note-nanobanana_blog-21)[\[20\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_note-tti_page-20)</sup>.

## Ograniczenia i krytyka

Pomimo skali i popularności, podejście ma swoje ograniczenia:

- **Subiektywność i efekty stylowe.** Preferencje głosów zależą od tonu/sposobu odpowiedzi; zespół wdrożył **Style/Sentiment Control** w celu oddzielenia „stylu" od „treści"<sup>[\[16\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_note-sentiment_control-16)</sup>.
- **Niereprezentacyjność odbiorców.** Aktywne jądro to entuzjaści technologii/deweloperzy; dla scenariuszy dziedzinowych tworzone są wyspecjalizowane areny (Search, WebDev, Biomed i in.)<sup>[\[33\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_note-33)</sup>.
- **Podatność na manipulacje i obciążenia.** Badania z 2025 roku wykazują, że przy braku ścisłych zabezpieczeń możliwe są strategie „nabijania głosów" z setkami–tysiącami głosów; przy czym współpraca badaczy z LMArena doprowadziła do wdrożenia środków ochrony (CAPTCHA/logowanie/ochrona przed botami/wykrywanie anomalii) i wzrostu „kosztu ataku"<sup>[\[34\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_note-34)[\[35\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_note-35)[\[36\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_note-36)</sup>.
- **Krytyka metodologiczna.** Praca *The Leaderboard Illusion* (kwiecień 2025) wskazuje na systematyczne i instytucjonalne czynniki mogące wypaczać pole rywalizacji; LMArena opublikowała obszerną odpowiedź i prowadzi publiczny *changelog* metodologii<sup>[\[37\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_note-37)[\[38\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_note-38)[\[39\]](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_note-changelog-39)</sup>.

## Linki

- Oficjalna strona LMArena
- Blog/polityki i aktualizacje LMArena
- Strona grupy badawczej LMSYS (oryginalny inkubator projektu)

## Literatura

- Chiang, W.-L. et al. (2024). *Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference*. arXiv:2403.04132.
- Zheng, L. et al. (2023). *Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena*. arXiv:2306.05685.
- Li, T. et al. (2024). *From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline*. arXiv:2406.11939.
- Ameli, S.; Zhuang, S.; Stoica, I.; Mahoney, M. W. (2024). *A Statistical Framework for Ranking LLM-Based Chatbots*. arXiv:2412.18407.
- Boubdir, M. et al. (2023). *Elo Uncovered: Robustness and Best Practices in Language Model Evaluation*. arXiv:2311.17295.
- Huang, J. Y.; Shen, Y.; Wei, D.; Broderick, T. (2025). *Dropping Just a Handful of Preferences Can Change Top Large Language Model Rankings*. arXiv:2508.11847.
- Xu, Y.; Ruis, L.; Rocktäschel, T.; Kirk, R. (2025). *Investigating Non-Transitivity in LLM-as-a-Judge*. arXiv:2502.14074.
- Li, H. et al. (2024). *LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods*. arXiv:2412.05579.
- Zheng, L. et al. (2024). *LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset*. arXiv:2309.11998.
- Dubois, Y. et al. (2024). *Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators*. arXiv:2404.04475.
- Singh, S. et al. (2025). *The Leaderboard Illusion*. arXiv:2504.20879.
- Min, R.; Pang, T.; Du, C.; Liu, Q.; Cheng, M.; Lin, M. (2025). *Improving Your Model Ranking on Chatbot Arena by Vote Rigging*. arXiv:2501.17858.

## Przypisy

1.  <span id="cite_note-chiang2024-1">↑ <sup>[1.0](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_ref-chiang2024_1-0)</sup> <sup>[1.1](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_ref-chiang2024_1-1)</sup> Chiang, W.-L. et al. «Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference». *arXiv:2403.04132*, 2024.</span>
2.  <span id="cite_note-hello_2025-2">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_ref-hello_2025_2-0) «Hello from LMArena: The Community Platform for Exploring Frontier AI». *LMArena Blog*, 23 июня 2025.</span>
3.  <span id="cite_note-new_site_2024-3">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_ref-new_site_2024_3-0) «Announcing a New Site for Chatbot Arena». *LMSYS Blog*, 20 сентября 2024.</span>
4.  <span id="cite_note-seed_prn-4">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_ref-seed_prn_4-0) «Arena Intelligence Raises \$100M Seed Round to Build the Standard for AI Evaluation». *PR Newswire / Arena Intelligence*, 21 мая 2025. <a href="https://www.prnewswire.com/news-releases/arena-intelligence-raises-100m-seed-round.html" class="external autonumber" rel="nofollow">[1]</a></span>
5.  <span id="cite_note-tc_seed-5">↑ <sup>[5.0](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_ref-tc_seed_5-0)</sup> <sup>[5.1](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_ref-tc_seed_5-1)</sup> Wiggers, K. «LMArena, the Chatbot Arena spinoff, raises \$100M seed at a \$600M valuation». *TechCrunch*, 21 мая 2025. <a href="https://techcrunch.com/2025/05/21/lmarena-chatbot-arena-spinoff-raises-100m-seed/" class="external autonumber" rel="nofollow">[2]</a></span>
6.  <span id="cite_note-series_a_2026-6">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_ref-series_a_2026_6-0) «Arena Intelligence Closes \$150M Series A at \$1.7B Valuation». *Arena Intelligence Blog*, 6 января 2026. <a href="https://news.lmarena.ai/series-a-2026/" class="external autonumber" rel="nofollow">[3]</a></span>
7.  <span id="cite_note-arena_rebrand_2026-7">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_ref-arena_rebrand_2026_7-0) «We Are Now Arena». *Arena Blog*, 28 января 2026. <a href="https://news.arena.ai/we-are-now-arena/" class="external autonumber" rel="nofollow">[4]</a></span>
8.  <span id="cite_note-new_beta_2025-8">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_ref-new_beta_2025_8-0) «LMArena is Growing to Support our Community Platform». *LMArena Blog*, 17 апреля 2025.</span>
9.  <span id="cite_note-policy-9">↑ <sup>[9.0](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_ref-policy_9-0)</sup> <sup>[9.1](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_ref-policy_9-1)</sup> <sup>[9.2](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_ref-policy_9-2)</sup> <sup>[9.3](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_ref-policy_9-3)</sup> *LMArena Leaderboard Policy*. *LMArena Blog*, ред. 8 сентября 2025. <a href="https://news.lmarena.ai/policy/" class="external autonumber" rel="nofollow">[5]</a></span>
10. <span id="cite_note-10">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_ref-10) lm‑sys/FastChat (GitHub). <a href="https://github.com/lm-sys/FastChat" class="external autonumber" rel="nofollow">[6]</a></span>
11. <span id="cite_note-opendata_2025-11">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_ref-opendata_2025_11-0) «Releasing 140K Text Arena Conversations». *LMArena Blog*, 31 июля 2025. <a href="https://news.lmarena.ai/140k-text-arena-conversations/" class="external autonumber" rel="nofollow">[7]</a></span>
12. <span id="cite_note-12">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_ref-12) *FAQ*. *LMArena*. <a href="https://lmarena.ai/faq" class="external autonumber" rel="nofollow">[8]</a></span>
13. <span id="cite_note-13">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_ref-13) Главная страница *LMArena* (дисклеймер о возможной публикации данных и передачи провайдерам). <a href="https://lmarena.ai/" class="external autonumber" rel="nofollow">[9]</a></span>
14. <span id="cite_note-14">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_ref-14) Zheng, L. et al. «Judging LLM‑as‑a‑Judge with MT‑Bench and Chatbot Arena». *arXiv:2306.05685*, 2023. <a href="https://arxiv.org/abs/2306.05685" class="external autonumber" rel="nofollow">[10]</a></span>
15. <span id="cite_note-arena_hard-15">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_ref-arena_hard_15-0) Li, T. et al. «From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline». *arXiv:2406.11939*, 2024. <a href="https://arxiv.org/abs/2406.11939" class="external autonumber" rel="nofollow">[11]</a></span>
16. <span id="cite_note-sentiment_control-16">↑ <sup>[16.0](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_ref-sentiment_control_16-0)</sup> <sup>[16.1](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_ref-sentiment_control_16-1)</sup> «Style and Sentiment Control in the Arena». *LMArena Blog*, 2025. <a href="https://news.lmarena.ai/style-sentiment-control/" class="external autonumber" rel="nofollow">[12]</a></span>
17. <span id="cite_note-17">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_ref-17) Li, T. et al. «From Crowdsourced Data…» *arXiv:2406.11939* (таблицы согласованности). <a href="https://arxiv.org/abs/2406.11939" class="external autonumber" rel="nofollow">[13]</a></span>
18. <span id="cite_note-text_stats-18">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_ref-text_stats_18-0) *Text Arena (English)*. *LMArena*. <a href="https://lmarena.ai/leaderboard/text/english" class="external autonumber" rel="nofollow">[14]</a></span>
19. <span id="cite_note-19">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_ref-19) *Vision Arena*. *LMArena*, обновлено 2 сентября 2025. <a href="https://lmarena.ai/leaderboard/vision" class="external autonumber" rel="nofollow">[15]</a></span>
20. <span id="cite_note-tti_page-20">↑ <sup>[20.0](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_ref-tti_page_20-0)</sup> <sup>[20.1](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_ref-tti_page_20-1)</sup> *Text-to-Image Leaderboard*. *LMArena*. <a href="https://lmarena.ai/leaderboard/text-to-image" class="external autonumber" rel="nofollow">[16]</a></span>
21. <span id="cite_note-nanobanana_blog-21">↑ <sup>[21.0](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_ref-nanobanana_blog_21-0)</sup> <sup>[21.1](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_ref-nanobanana_blog_21-1)</sup> «The nano-banana Story: Gemini 2.5 Flash Image Tops the Visual Arena». *LMArena Blog*, август 2025. <a href="https://news.lmarena.ai/nano-banana-reveal/" class="external autonumber" rel="nofollow">[17]</a></span>
22. <span id="cite_note-22">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_ref-22) *Text‑to‑Video* и *Image‑to‑Video Leaderboards*. *LMArena*, август 2025. <a href="https://lmarena.ai/leaderboard/text-to-video" class="external autonumber" rel="nofollow">[18]</a> <a href="https://lmarena.ai/leaderboard/image-to-video" class="external autonumber" rel="nofollow">[19]</a></span>
23. <span id="cite_note-webdev_arena-23">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_ref-webdev_arena_23-0) «Introducing WebDev Arena». *LMArena Blog*, 2025. <a href="https://news.lmarena.ai/webdev-arena/" class="external autonumber" rel="nofollow">[20]</a></span>
24. <span id="cite_note-repochat_arena-24">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_ref-repochat_arena_24-0) «Introducing RepoChat Arena». *LMArena Blog*, 2025. <a href="https://news.lmarena.ai/repochat-arena/" class="external autonumber" rel="nofollow">[21]</a></span>
25. <span id="cite_note-25">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_ref-25) «Introducing the Search Arena». *LMArena Blog*, 14 апреля 2025. <a href="https://news.lmarena.ai/search-arena/" class="external autonumber" rel="nofollow">[22]</a></span>
26. <span id="cite_note-26">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_ref-26) «Search Arena & What We're Learning About Human Preference». *LMArena Blog*, 23 июля 2025. <a href="https://news.lmarena.ai/search-arena-update/" class="external autonumber" rel="nofollow">[23]</a></span>
27. <span id="cite_note-27">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_ref-27) Frick, E. et al. «Search Arena: Analyzing Search‑Augmented LLMs». *arXiv:2506.05334*, 2025. <a href="https://arxiv.org/abs/2506.05334" class="external autonumber" rel="nofollow">[24]</a></span>
28. <span id="cite_note-28">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_ref-28) «Introducing BiomedArena.AI». *LMArena Blog*, 19 августа 2025. <a href="https://news.lmarena.ai/introducing-biomedarena/" class="external autonumber" rel="nofollow">[25]</a></span>
29. <span id="cite_note-29">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_ref-29) Google. «Gemma 3…», 12 марта 2025 (ссылка на результаты LMArena). <a href="https://blog.google/technology/developers/gemma-3/" class="external autonumber" rel="nofollow">[26]</a></span>
30. <span id="cite_note-30">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_ref-30) Spangher, L. et al. «Chatbot Arena Estimate…». *NAACL Industry*, 2025. <a href="https://aclanthology.org/2025.naacl-industry.77.pdf" class="external autonumber" rel="nofollow">[27]</a></span>
31. <span id="cite_note-31">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_ref-31) «Meta's experimental Llama 4 model briefly topped AI leaderboard…». *The Register*, 7 апреля 2025. <a href="https://www.theregister.com/2025/04/07/meta_llama_arena/" class="external autonumber" rel="nofollow">[28]</a></span>
32. <span id="cite_note-32">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_ref-32) Официальные разъяснения/посты LMArena в X по инциденту (апрель 2025). <a href="https://x.com/lmarena_ai/status/1776881220809552182" class="external autonumber" rel="nofollow">[29]</a></span>
33. <span id="cite_note-33">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_ref-33) «Search Arena & What We're Learning…». *LMArena Blog*, 23 июля 2025. <a href="https://news.lmarena.ai/search-arena-update/" class="external autonumber" rel="nofollow">[30]</a></span>
34. <span id="cite_note-34">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_ref-34) Min, R. et al. «Improving Your Model Ranking on Chatbot Arena by Vote Rigging». *arXiv:2501.17858*, 2025. <a href="https://arxiv.org/abs/2501.17858" class="external autonumber" rel="nofollow">[31]</a></span>
35. <span id="cite_note-35">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_ref-35) Huang, Y. et al. «Exploring and Mitigating Adversarial Manipulation of Voting‑Based Leaderboards». *arXiv:2501.07493*, 2025. <a href="https://arxiv.org/abs/2501.07493" class="external autonumber" rel="nofollow">[32]</a></span>
36. <span id="cite_note-36">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_ref-36) «Hundreds of rigged votes can skew…». *Fast Company*, 6 февраля 2025. <a href="https://www.fastcompany.com/91273226/rigged-votes-ai-model-rankings-chatbot-arena" class="external autonumber" rel="nofollow">[33]</a></span>
37. <span id="cite_note-37">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_ref-37) Singh, S. et al. «The Leaderboard Illusion». *arXiv:2504.20879*, 2025. <a href="https://arxiv.org/abs/2504.20879" class="external autonumber" rel="nofollow">[34]</a></span>
38. <span id="cite_note-38">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_ref-38) «Our Response to 'The Leaderboard Illusion'». *LMArena Blog*, 9 мая 2025. <a href="https://news.lmarena.ai/our-response/" class="external autonumber" rel="nofollow">[35]</a></span>
39. <span id="cite_note-changelog-39">[↑](https://systems-analysis.info/int/LMArena_(Chatbot_Arena)_(PL)#cite_ref-changelog_39-0) «Methodology Changelog». *LMArena Blog*. <a href="https://news.lmarena.ai/changelog/" class="external autonumber" rel="nofollow">[36]</a></span>
