# Models > The leader of AI leaderboards: every public model leaderboard it can read (28 so far: LMArena, Artificial Analysis, Scale SEAL, Epoch AI, ARC Prize, LiveBench, SWE-bench, BFCL, MTEB, OpenRouter and more), who is #1 on each, and a transparent composite that names the leader of leaders. It grows: new models join the day a board ranks them (15 in the last 30 days), and new leaderboards are found each run (41 found, waiting for a reader). Every board is captured daily at 04:30 UTC straight from the board's own page, data file or API. Each capture stores, per board, every model's rank and score as published, with the capture date; variants of one model (effort levels, thinking modes, dated snapshots) collapse to the model and its best variant stands for it. History comes from these daily captures plus LMArena's and LiveBench's own published snapshots. The composite gives 10 points for #1 down to 1 for #10 on each fresh language-model board (chat, reasoning, coding, agents, vision) and reports points over points available as a podium score. Full method: https://models.fru.dev/method Last successful capture: 2026-09-26 23:00:25 UTC. 25 fresh boards, 823 models, 1735 ranked entries. Current leader of leaders: 1. Claude Fable 5.1 (podium score 57.3, 1 board wins); 2. GPT-6 Astra (podium score 48, 4 board wins); 3. Claude Opus 5 (podium score 38, 2 board wins); 4. Claude Fable 5 (podium score 37.3, 1 board wins); 5. Claude Opus 5.5 (podium score 32, 3 board wins). ## About this data Scores come from each leaderboard's own published results, captured every day and shown exactly as each board published them; we run no evaluations. Model facts come from Hugging Face, OpenRouter and the fru.dev pricing and regions sites. The list grows as boards add new models. Logos via logo.dev; trademarks belong to their owners. Anonymous usage stats via Google Analytics. ## Pages - [Leader of leaders and the podium wall](https://models.fru.dev/) - [All boards](https://models.fru.dev/boards) - [The model catalog: parameters, modalities, license, context, price, hosting, rank](https://models.fru.dev/models) - [Benchmarks: what each AI benchmark measures, with paper, versions and captured scores](https://models.fru.dev/benchmarks) - [Vendors](https://models.fru.dev/vendors) - [Compare two to four models or labs side by side on every board](https://models.fru.dev/compare): /compare?models=,, /compare?vendors=,, /compare?weights=open,closed (JSON: https://models.fru.dev/api/compare with the same query) - [Compare: GPT vs Claude](https://models.fru.dev/compare?models=gpt-6-astra,claude-fable-5-1): GPT-6 Astra and Claude Fable 5.1, the top-ranked of each - [Compare: Claude vs Gemini](https://models.fru.dev/compare?models=claude-fable-5-1,gemini-3-8-flash): Claude Fable 5.1 and Gemini 3.8 Flash, the top-ranked of each - [Compare: Anthropic vs OpenAI vs Meta](https://models.fru.dev/compare?vendors=anthropic,openai,meta): Each lab's best-ranked model on every board - [Compare: Open-weight vs closed](https://models.fru.dev/compare?weights=open,closed): The best open-weight model on every board against the best closed one - [Method](https://models.fru.dev/method) - [Changes: models entering, moving, leaving; new board versions](https://models.fru.dev/changes) - [Chat leaders](https://models.fru.dev/categories/chat) - [Reasoning leaders](https://models.fru.dev/categories/reasoning) - [Coding leaders](https://models.fru.dev/categories/coding) - [Agents and tools leaders](https://models.fru.dev/categories/agents) - [Vision leaders](https://models.fru.dev/categories/vision) - [Image leaders](https://models.fru.dev/categories/image) - [Video leaders](https://models.fru.dev/categories/video) - [Speech leaders](https://models.fru.dev/categories/speech) - [Embeddings leaders](https://models.fru.dev/categories/embeddings) - [Speed leaders](https://models.fru.dev/categories/speed) - [Price leaders](https://models.fru.dev/categories/price) - [Usage leaders](https://models.fru.dev/categories/usage) - [LMArena Text Arena](https://models.fru.dev/boards/lmarena-text): #1 Claude Opus 5.5 (1,509), updated 2026-09-25 - [Scale SEAL MultiChallenge](https://models.fru.dev/boards/scale-multichallenge): #1 Muse Spark (75.5%), updated 2026-09-26 - [Artificial Analysis Intelligence Index](https://models.fru.dev/boards/aa-intelligence): #1 Claude Opus 5.5 (57.6), updated 2026-09-26 - [Humanity's Last Exam](https://models.fru.dev/boards/scale-hle): #1 GPT-6 Astra (54.8%), updated 2026-09-26 - [GPQA Diamond (Epoch AI runs)](https://models.fru.dev/boards/epoch-gpqa): #1 GPT-6 Astra (95.8%), updated 2026-09-02 - [Epoch Capabilities Index](https://models.fru.dev/boards/epoch-eci): #1 GPT-6 Astra (166.6), updated 2026-09-26 - [ARC-AGI-2](https://models.fru.dev/boards/arc-agi-2): #1 GPT-6 Astra (95.0%), updated 2026-09-24 - [LiveBench](https://models.fru.dev/boards/livebench): #1 Claude Fable 5.1 (83.4%), updated 2026-06-25 - [MMLU-Pro](https://models.fru.dev/boards/mmlu-pro): #1 Gemini 3.1 Pro Preview (91.2%), updated 2026-03-11 - [LMArena WebDev Arena](https://models.fru.dev/boards/lmarena-webdev): #1 Claude Opus 5.5 (1,827), updated 2026-09-25 - [SWE-Bench Pro](https://models.fru.dev/boards/scale-swe-pro): #1 Claude Opus 5 (98.0%), updated 2026-09-26 - [SWE-bench Verified](https://models.fru.dev/boards/swe-bench-verified): #1 Claude Opus 4.5 (76.8%), updated 2026-02-26 - [Aider Polyglot](https://models.fru.dev/boards/aider-polyglot): #1 GPT-5 (88.0%), updated 2025-10-03 - [Berkeley Function Calling Leaderboard](https://models.fru.dev/boards/bfcl): #1 Claude Opus 4.5 (77.5%), updated 2026-04-12 - [MCP Atlas](https://models.fru.dev/boards/scale-mcp-atlas): #1 Muse Spark 1.1 (88.1%), updated 2026-09-26 - [LMArena Search Arena](https://models.fru.dev/boards/lmarena-search): #1 GPT-5.6 Sol (1,257), updated 2026-08-24 - [LMArena Vision Arena](https://models.fru.dev/boards/lmarena-vision): #1 Claude Fable 5 (1,310), updated 2026-09-13 - [LMArena Document Arena](https://models.fru.dev/boards/lmarena-document): #1 Claude Opus 5 (1,516), updated 2026-09-13 - [LMArena Text-to-Image](https://models.fru.dev/boards/lmarena-text-to-image): #1 GPT Image 2.5 Sunburst (1,424), updated 2026-09-24 - [LMArena Image Edit](https://models.fru.dev/boards/lmarena-image-edit): #1 GPT Image 2.5 Sunburst (1,526), updated 2026-09-21 - [LMArena Text-to-Video](https://models.fru.dev/boards/lmarena-text-to-video): #1 Gemini Omni 1.1 Flash (1,516), updated 2026-09-21 - [LMArena Image-to-Video](https://models.fru.dev/boards/lmarena-image-to-video): #1 MiniMax H3 (1,495), updated 2026-09-21 - [Artificial Analysis Speech Arena](https://models.fru.dev/boards/aa-tts): #1 Sonic 3.6 (1,277), updated 2026-09-26 - [Open ASR Leaderboard](https://models.fru.dev/boards/open-asr): #1 scribe v2 pro (zoom) (3.59 WER), updated 2026-09-25 - [MTEB Multilingual v2](https://models.fru.dev/boards/mteb-multilingual): #1 harrier-oss-v1-27b (microsoft) (74.3%), updated 2026-09-22 - [Artificial Analysis Output Speed](https://models.fru.dev/boards/aa-speed): #1 Gemini 3.8 Flash (328 tok/s), updated 2026-09-26 - [Artificial Analysis Price](https://models.fru.dev/boards/aa-price): #1 GPT-6 Luna ($0.20), updated 2026-09-26 - [OpenRouter weekly usage](https://models.fru.dev/boards/openrouter-usage): #1 DeepSeek V4.1 Flash (19.2T tok), updated 2026-09-25 - [Claude Fable 5.1](https://models.fru.dev/models/claude-fable-5-1) - [GPT-6 Astra](https://models.fru.dev/models/gpt-6-astra) - [Claude Opus 5](https://models.fru.dev/models/claude-opus-5) - [Claude Fable 5](https://models.fru.dev/models/claude-fable-5) - [Claude Opus 5.5](https://models.fru.dev/models/claude-opus-5-5) - [Anthropic](https://models.fru.dev/vendors/anthropic): 8 board wins - [OpenAI](https://models.fru.dev/vendors/openai): 8 board wins - [Google](https://models.fru.dev/vendors/google): 2 board wins - [Meta](https://models.fru.dev/vendors/meta): 2 board wins - [DeepSeek](https://models.fru.dev/vendors/deepseek): 1 board wins - [Microsoft](https://models.fru.dev/vendors/microsoft): 1 board wins - [MiniMax](https://models.fru.dev/vendors/minimax): 1 board wins - [Zoom](https://models.fru.dev/vendors/zoom): 1 board wins ## Benchmarks - [MMLU](https://models.fru.dev/benchmarks/mmlu): Knowledge; Answer four-option multiple-choice questions across 57 subjects, from elementary math to law. Paper: https://arxiv.org/abs/2009.03300 - [MMLU-Pro](https://models.fru.dev/benchmarks/mmlu-pro): Knowledge; Answer ten-option multiple-choice questions that need more reasoning than MMLU. Paper: https://arxiv.org/abs/2406.01574 - [GPQA (Diamond)](https://models.fru.dev/benchmarks/gpqa): Knowledge; Answer graduate-level biology, physics and chemistry questions that non-experts cannot look up. Paper: https://arxiv.org/abs/2311.12022 - [Humanity's Last Exam](https://models.fru.dev/benchmarks/hle): Knowledge; Answer expert-written closed-ended questions at the frontier of human knowledge, some with images. Paper: https://arxiv.org/abs/2501.14249 - [SimpleQA](https://models.fru.dev/benchmarks/simpleqa): Knowledge; Answer short fact-seeking questions with a single correct answer, or decline. Paper: https://arxiv.org/abs/2411.04368 - [TruthfulQA](https://models.fru.dev/benchmarks/truthfulqa): Knowledge; Answer questions where a common misconception tempts a false answer. Paper: https://arxiv.org/abs/2109.07958 - [ARC-AGI-1](https://models.fru.dev/benchmarks/arc-agi-1): Reasoning; Infer the rule behind a few input-output grid pairs and apply it to a new grid. Paper: https://arxiv.org/abs/1911.01547 - [ARC-AGI-2](https://models.fru.dev/benchmarks/arc-agi-2): Reasoning; Solve harder abstract grid puzzles that are easy for people and hard for AI. Paper: https://arxiv.org/abs/2505.11831 - [ARC-AGI-3](https://models.fru.dev/benchmarks/arc-agi-3): Reasoning; Play unfamiliar interactive games with no instructions: explore, find the goal and win as efficiently as a person. - [BIG-Bench Hard](https://models.fru.dev/benchmarks/bbh): Reasoning; Solve 23 BIG-Bench tasks where earlier models fell short of average human raters. Paper: https://arxiv.org/abs/2210.09261 - [BIG-bench](https://models.fru.dev/benchmarks/big-bench): Reasoning; A broad collaborative suite of tasks believed to be beyond language models at the time. Paper: https://arxiv.org/abs/2206.04615 - [HellaSwag](https://models.fru.dev/benchmarks/hellaswag): Reasoning; Pick the most plausible continuation of an everyday scenario. Paper: https://arxiv.org/abs/1905.07830 - [AI2 Reasoning Challenge (ARC)](https://models.fru.dev/benchmarks/ai2-arc): Reasoning; Answer grade-school science multiple-choice questions; the Challenge set defeats simple retrieval. Paper: https://arxiv.org/abs/1803.05457 - [WinoGrande](https://models.fru.dev/benchmarks/winogrande): Reasoning; Resolve which of two options a pronoun-like blank refers to, using commonsense. Paper: https://arxiv.org/abs/1907.10641 - [MuSR](https://models.fru.dev/benchmarks/musr): Reasoning; Solve long natural-language murder mysteries, object placements and team allocations step by step. Paper: https://arxiv.org/abs/2310.16049 - [LiveBench](https://models.fru.dev/benchmarks/livebench): Reasoning; Answer fresh questions in math, coding, reasoning, data analysis, language and instruction following. Paper: https://arxiv.org/abs/2406.19314 - [GSM8K](https://models.fru.dev/benchmarks/gsm8k): Math; Solve grade-school math word problems that take two to eight steps. Paper: https://arxiv.org/abs/2110.14168 - [MATH (and MATH-500)](https://models.fru.dev/benchmarks/math): Math; Solve competition mathematics problems and give the final answer. Paper: https://arxiv.org/abs/2103.03874 - [AIME](https://models.fru.dev/benchmarks/aime): Math; Solve the American Invitational Mathematics Examination, a qualifier for the US Math Olympiad. - [FrontierMath](https://models.fru.dev/benchmarks/frontiermath): Math; Solve original, unpublished research-level mathematics problems with automatically checkable answers. Paper: https://arxiv.org/abs/2411.04872 - [MGSM](https://models.fru.dev/benchmarks/mgsm): Math; Solve grade-school math word problems in ten typologically diverse languages. Paper: https://arxiv.org/abs/2210.03057 - [HumanEval](https://models.fru.dev/benchmarks/humaneval): Coding; Write a Python function from its signature and docstring. Paper: https://arxiv.org/abs/2107.03374 - [MBPP](https://models.fru.dev/benchmarks/mbpp): Coding; Write short Python programs from a one-sentence description. Paper: https://arxiv.org/abs/2108.07732 - [SWE-bench](https://models.fru.dev/benchmarks/swe-bench): Coding; Resolve a real GitHub issue by editing the repository so the hidden tests pass. Paper: https://arxiv.org/abs/2310.06770 - [SWE-bench Verified](https://models.fru.dev/benchmarks/swe-bench-verified): Coding; Resolve real GitHub issues; the subset people confirmed is well specified and fairly tested. Paper: https://arxiv.org/abs/2310.06770 - [SWE-Bench Pro](https://models.fru.dev/benchmarks/swe-bench-pro): Coding; Solve long-horizon, enterprise-style software engineering tasks in real repositories. Paper: https://arxiv.org/abs/2509.16941 - [LiveCodeBench](https://models.fru.dev/benchmarks/livecodebench): Coding; Solve competitive programming problems published after a model was trained; also self-repair and test-output prediction. Paper: https://arxiv.org/abs/2403.07974 - [Aider Polyglot](https://models.fru.dev/benchmarks/aider-polyglot): Coding; Edit code to solve hard programming exercises in six languages, through the Aider coding assistant. - [Terminal-Bench](https://models.fru.dev/benchmarks/terminal-bench): Coding; Complete hard, realistic tasks in a computer terminal: build, debug, configure, analyse. Paper: https://arxiv.org/abs/2601.11868 - [SciCode](https://models.fru.dev/benchmarks/scicode): Coding; Write code for research problems from 16 natural-science subfields. Paper: https://arxiv.org/abs/2407.13168 - [BigCodeBench](https://models.fru.dev/benchmarks/bigcodebench): Coding; Write Python that calls many library functions to complete practical tasks. Paper: https://arxiv.org/abs/2406.15877 - [MLE-bench](https://models.fru.dev/benchmarks/mle-bench): Coding; Act as a machine-learning engineer on Kaggle competitions: prepare data, train, submit. Paper: https://arxiv.org/abs/2410.07095 - [SWE-Lancer](https://models.fru.dev/benchmarks/swe-lancer): Coding; Complete real freelance software jobs from Upwork, both implementation and choosing between proposals. Paper: https://arxiv.org/abs/2502.12115 - [τ-bench (tau-bench)](https://models.fru.dev/benchmarks/tau-bench): Agents and tools; Serve a simulated customer through tools while following a domain policy. Paper: https://arxiv.org/abs/2406.12045 - [Berkeley Function Calling Leaderboard (BFCL)](https://models.fru.dev/benchmarks/bfcl): Agents and tools; Call functions and tools correctly: single, parallel, multi-turn, and agentic web search and memory. - [MCP Atlas](https://models.fru.dev/benchmarks/mcp-atlas): Agents and tools; Complete realistic tasks by calling tools on real MCP servers, often across several servers. - [GAIA](https://models.fru.dev/benchmarks/gaia): Agents and tools; Answer real-world questions that need browsing, tools, files and several steps. Paper: https://arxiv.org/abs/2311.12983 - [WebArena](https://models.fru.dev/benchmarks/webarena): Agents and tools; Complete tasks on realistic self-hosted websites: shopping, forums, code hosting, maps, content management. Paper: https://arxiv.org/abs/2307.13854 - [OSWorld](https://models.fru.dev/benchmarks/osworld): Agents and tools; Operate a real computer (Ubuntu, Windows, macOS) through screenshots and input to finish open-ended tasks. Paper: https://arxiv.org/abs/2404.07972 - [BrowseComp](https://models.fru.dev/benchmarks/browsecomp): Agents and tools; Find hard-to-locate, entangled facts by persistently browsing the web. Paper: https://arxiv.org/abs/2504.12516 - [GDPval](https://models.fru.dev/benchmarks/gdpval): Agents and tools; Produce real work products (documents, slides, spreadsheets) for tasks from 44 occupations. Paper: https://arxiv.org/abs/2510.04374 - [Vending-Bench](https://models.fru.dev/benchmarks/vending-bench): Agents and tools; Run a simulated vending machine business for a long time: inventory, orders, prices, fees. Paper: https://arxiv.org/abs/2502.15840 - [IFEval](https://models.fru.dev/benchmarks/ifeval): Chat and instructions; Follow verifiable formatting instructions such as "write more than 400 words". Paper: https://arxiv.org/abs/2311.07911 - [MultiChallenge](https://models.fru.dev/benchmarks/multichallenge): Chat and instructions; Hold realistic multi-turn conversations: retain instructions, remember context, stay self-coherent. Paper: https://arxiv.org/abs/2501.17399 - [RULER](https://models.fru.dev/benchmarks/ruler): Long context; Retrieve, trace and aggregate information across long synthetic contexts, beyond one needle. Paper: https://arxiv.org/abs/2404.06654 - [LongBench](https://models.fru.dev/benchmarks/longbench): Long context; Understand long real documents in English and Chinese: QA, summarisation, code, few-shot. Paper: https://arxiv.org/abs/2308.14508 - [∞Bench (InfiniteBench)](https://models.fru.dev/benchmarks/infinitebench): Long context; Answer questions over contexts longer than 100,000 tokens: books, code, math, dialogue. Paper: https://arxiv.org/abs/2402.13718 - [MRCR (multi-round coreference)](https://models.fru.dev/benchmarks/mrcr): Long context; Find the right one of several near-identical requests hidden in a long conversation and reproduce its answer. Paper: https://arxiv.org/abs/2409.12640 - [MMMU](https://models.fru.dev/benchmarks/mmmu): Multimodal; Answer college-level questions that need an image: charts, diagrams, maps, music sheets, chemical structures. Paper: https://arxiv.org/abs/2311.16502 - [MathVista](https://models.fru.dev/benchmarks/mathvista): Multimodal; Solve math problems set in images: plots, tables, geometry, puzzles. Paper: https://arxiv.org/abs/2310.02255 - [ChartQA](https://models.fru.dev/benchmarks/chartqa): Multimodal; Answer questions about charts that need visual and logical reasoning. Paper: https://arxiv.org/abs/2203.10244 - [DocVQA](https://models.fru.dev/benchmarks/docvqa): Multimodal; Answer questions about scanned document images. Paper: https://arxiv.org/abs/2007.00398 - [Video-MME](https://models.fru.dev/benchmarks/video-mme): Multimodal; Answer questions about videos from 11 seconds to an hour long, with or without subtitles and audio. Paper: https://arxiv.org/abs/2405.21075 - [MMBench](https://models.fru.dev/benchmarks/mmbench): Multimodal; Answer multiple-choice questions testing 20 vision-language abilities. Paper: https://arxiv.org/abs/2307.06281 - [AI2D](https://models.fru.dev/benchmarks/ai2d): Multimodal; Answer questions about grade-school science diagrams. Paper: https://arxiv.org/abs/1603.07396 - [GenEval](https://models.fru.dev/benchmarks/geneval): Image and video generation; Generate images that get objects, counts, colours and positions right. Paper: https://arxiv.org/abs/2310.11513 - [VBench](https://models.fru.dev/benchmarks/vbench): Image and video generation; Generate videos that score well on 16 separate quality and consistency dimensions. Paper: https://arxiv.org/abs/2311.17982 - [Open ASR Leaderboard](https://models.fru.dev/benchmarks/open-asr): Speech; Transcribe English short- and long-form speech and multilingual speech. Paper: https://arxiv.org/abs/2510.06961 - [LibriSpeech](https://models.fru.dev/benchmarks/librispeech): Speech; Transcribe read English audiobook speech. - [Common Voice](https://models.fru.dev/benchmarks/common-voice): Speech; Transcribe crowd-sourced read speech in many languages. Paper: https://arxiv.org/abs/1912.06670 - [FLEURS](https://models.fru.dev/benchmarks/fleurs): Speech; Recognise, identify and translate speech in 102 languages. Paper: https://arxiv.org/abs/2205.12446 - [Artificial Analysis Speech Arena](https://models.fru.dev/benchmarks/aa-speech-arena): Speech; Read text aloud in a voice people prefer. - [MTEB (and MMTEB)](https://models.fru.dev/benchmarks/mteb): Embeddings; Embed text for retrieval, clustering, classification, reranking and similarity. Paper: https://arxiv.org/abs/2210.07316 - [HarmBench](https://models.fru.dev/benchmarks/harmbench): Safety and honesty; Refuse harmful behaviours under automated red-teaming attacks. Paper: https://arxiv.org/abs/2402.04249 - [AgentHarm](https://models.fru.dev/benchmarks/agentharm): Safety and honesty; Refuse explicitly malicious multi-step agent tasks that use tools. Paper: https://arxiv.org/abs/2410.09024 - [StrongREJECT](https://models.fru.dev/benchmarks/strongreject): Safety and honesty; Measure how much a jailbreak really helps: refusals and the usefulness of any harmful answer. Paper: https://arxiv.org/abs/2402.10260 - [XSTest](https://models.fru.dev/benchmarks/xstest): Safety and honesty; Answer safe prompts that only look risky, and refuse the unsafe contrasts. Paper: https://arxiv.org/abs/2308.01263 - [BBQ](https://models.fru.dev/benchmarks/bbq): Safety and honesty; Answer questions where a social stereotype could bias the answer. Paper: https://arxiv.org/abs/2110.08193 - [WMDP](https://models.fru.dev/benchmarks/wmdp): Safety and honesty; Measure hazardous knowledge in biosecurity, cybersecurity and chemical security (lower can be safer). Paper: https://arxiv.org/abs/2403.03218 - [MASK](https://models.fru.dev/benchmarks/mask): Safety and honesty; Tell whether a model states what it believes when pressured to lie, separately from whether it is accurate. Paper: https://arxiv.org/abs/2503.03750 - [LMArena Text Arena (Chatbot Arena)](https://models.fru.dev/benchmarks/lmarena-text): Arenas; Chat: people ask anything and vote for the better of two anonymous answers. Paper: https://arxiv.org/abs/2403.04132 - [LMArena WebDev Arena](https://models.fru.dev/benchmarks/lmarena-webdev): Arenas; Build a web app from a prompt; people vote for the better of two. - [LMArena Vision Arena](https://models.fru.dev/benchmarks/lmarena-vision): Arenas; Chat about images; people vote for the better answer. - [LMArena Document Arena](https://models.fru.dev/benchmarks/lmarena-document): Arenas; Answer questions about uploaded documents and PDFs; people vote. - [LMArena Search Arena](https://models.fru.dev/benchmarks/lmarena-search): Arenas; Answer with web search grounding; people vote for the better answer. - [LMArena Text-to-Image](https://models.fru.dev/benchmarks/lmarena-text-to-image): Arenas; Generate an image from a prompt; people vote for the better of two. - [LMArena Image Edit](https://models.fru.dev/benchmarks/lmarena-image-edit): Arenas; Edit an image from an instruction; people vote for the better edit. - [LMArena Text-to-Video](https://models.fru.dev/benchmarks/lmarena-text-to-video): Arenas; Generate a video from a prompt; people vote for the better of two. - [LMArena Image-to-Video](https://models.fru.dev/benchmarks/lmarena-image-to-video): Arenas; Animate an image into a video; people vote for the better of two. - [Artificial Analysis Intelligence Index](https://models.fru.dev/benchmarks/aa-intelligence-index): Composite indexes; One number from ten independent evaluations that Artificial Analysis runs itself. - [Epoch Capabilities Index](https://models.fru.dev/benchmarks/epoch-eci): Composite indexes; One capability score stitched together from dozens of benchmarks. - [Open LLM Leaderboard](https://models.fru.dev/benchmarks/hf-open-llm): Composite indexes; Rank open-weights models on a fixed set of academic benchmarks run the same way. - [Artificial Analysis Output Speed](https://models.fru.dev/benchmarks/aa-speed): Speed, price and usage; Measure how fast a model streams tokens on its first-party API. - [Artificial Analysis Price](https://models.fru.dev/benchmarks/aa-price): Speed, price and usage; Compare what a model costs on its first-party API. - [OpenRouter usage rankings](https://models.fru.dev/benchmarks/openrouter-usage): Speed, price and usage; Show which models developers actually send tokens to. ## Data - [Sitemap](https://models.fru.dev/sitemap.xml) - [Full plain-text dump](https://models.fru.dev/llms-full.txt): every board's top 25 - JSON, the whole computed summary: https://models.fru.dev/api/summary ## API - [For AI agents](https://models.fru.dev/agents): how to use this data in an agent (system prompt line, tool definition, code) - [OpenAPI 3.1 spec](https://models.fru.dev/openapi.json): the public read endpoints below, ready to load as tools - [GET /api/leaders](https://models.fru.dev/api/leaders?limit=3): The composite ranking: every model that earns points on the fresh language-model boards; returns models in rank order with podium score (0 to 100), #1 finishes, top-10 places and median rank - [GET /api/boards](https://models.fru.dev/api/boards?category=coding&limit=2): Every leaderboard tracked, with its current #1; returns boards with category, metric, status (fresh, stale, retired), publish date, entry count and the #1 model and score - [GET /api/boards/{id}](https://models.fru.dev/api/boards/lmarena-text?limit=3): One board: the full ranking as published, best variant per model; returns the board and up to 100 entries with rank, model, name as published, score and verified flag - [GET /api/models/{id}](https://models.fru.dev/api/models/gpt-6-astra): One model: its rank and score on every board that lists it; returns the model (name, vendor, open weights, release date), its composite standing and one row per board - [GET /api/compare](https://models.fru.dev/api/compare?vendors=openai,anthropic): Compare two to four models, vendors, or open against closed weights on every board where one is ranked; returns the sides with boards ranked and boards placed best, then one row per board (current capture) with each side's rank, score and model, and which side placed best - [GET /api/changes](https://models.fru.dev/api/changes?limit=3): Every rank change, newest first (append-only history); returns changes with board, model, kind (entered, moved, rescored, extra = votes/CI/cost per task changed at the same rank with extra_json, left, corrected, board_updated, model_*), old and new rank and score, and the day - [GET /api/search](https://models.fru.dev/api/search?q=claude): Search models, boards, labs and pages; returns up to 20 ranked results with title, link and one line Free to read. Cite "Models (models.fru.dev)" with a link. Responses are cached at the edge; keep to about 60 requests a minute. Rows marked unverified were read from a board page by a model because the parser failed, and nobody has checked them by hand yet.