https://ai-stats.phaseo.app/benchmarks/nyt-connections
NYT Connections - Benchmark Leaderboard & Model Performance | AI Stats
NYT Connections benchmark leaderboard on AI Stats. See 49 scored models, track historical performance, and inspect the underlying methodology. Current top...
nyt connectionsbenchmark leaderboardperformance aimodelstats
https://ai-stats.phaseo.app/benchmarks/simpleqa
SimpleQA - Benchmark Leaderboard & Model Performance | AI Stats
SimpleQA benchmark leaderboard on AI Stats. See 91 scored models, track historical performance, and inspect the underlying methodology. Current top model:...
benchmark leaderboardperformance aisimpleqamodelstats
https://llm-stats.com/benchmarks/bfcl-v3-multiturn
BFCL_v3_MultiTurn Benchmark Leaderboard
May 12, 2026 - Berkeley Function Calling Leaderboard (BFCL) V3 MultiTurn benchmark that evaluates large language models' ability to handle multi-turn and multi-step function...
multiturnbenchmarkleaderboard
https://llm-stats.com/benchmarks/mmau-sound
MMAU Sound Benchmark Leaderboard
May 16, 2026 - A subset of the MMAU benchmark focused specifically on environmental sound understanding and reasoning tasks. Part of a comprehensive multimodal audio...
soundbenchmarkleaderboard
https://llm-stats.com/benchmarks/mmbench-video
MMBench-Video Benchmark Leaderboard
May 15, 2026 - A long-form multi-shot benchmark for holistic video understanding that incorporates approximately 600 web videos from YouTube spanning 16 major categories,...
videobenchmarkleaderboard
https://llm-stats.com/benchmarks/qmsum
QMSum Benchmark Leaderboard
May 16, 2026 - QMSum is a benchmark for query-based multi-domain meeting summarization consisting of 1,808 query-summary pairs over 232 meetings across academic, product, and...
benchmarkleaderboard
https://llm-stats.com/benchmarks/gpqa-chemistry
GPQA Chemistry Benchmark Leaderboard
May 16, 2026 - Chemistry subset of GPQA, containing challenging multiple-choice questions written by domain experts in chemistry. These Google-proof questions require...
gpqachemistrybenchmarkleaderboard
https://llm-stats.com/benchmarks/finsearchcomp-t3
FinSearchComp-T3 Benchmark Leaderboard
May 20, 2026 - FinSearchComp-T3 is a benchmark for evaluating financial search and reasoning capabilities, testing models' ability to retrieve and analyze financial...
benchmarkleaderboard
https://llm-stats.com/benchmarks/cruxeval-output-cot
CRUXEval-Output-CoT Benchmark Leaderboard
May 16, 2026 - CRUXEval-O (output prediction) with Chain-of-Thought prompting. Part of the CRUXEval benchmark consisting of 800 Python functions (3-13 lines) designed to...
outputcotbenchmarkleaderboard
https://llm-stats.com/benchmarks/livecodebench-v6
LiveCodeBench v6 Benchmark Leaderboard
May 20, 2026 - LiveCodeBench is a holistic and contamination-free evaluation benchmark for large language models for code. It continuously collects new problems from...
livecodebenchbenchmarkleaderboard
https://llm-stats.com/benchmarks/activitynet
ActivityNet Benchmark Leaderboard
May 20, 2026 - A large-scale video benchmark for human activity understanding. Provides samples from 203 activity classes with an average of 137 untrimmed videos per class...
benchmarkleaderboard
https://www.sobigdata.eu/blog/building-cd-lead-benchmark-and-leaderboard-community-detection
Building CD-LEAD: a benchmark and leaderboard for community detection | SoBigData.eu
https://llm-stats.com/benchmarks/cnmo-2024
CNMO 2024 Benchmark Leaderboard
May 16, 2026 - China Mathematical Olympiad 2024 - A challenging mathematics competition.
cnmobenchmarkleaderboard
https://ha-vln-project.vercel.app/
An Open Benchmark and Leaderboard for Human-Aware Navigation in Discrete and Continuous...
Project page of paper 'HA-VLN 2.0: An Open Benchmark and Leaderboard for Human-Aware Navigation in Discrete and Continuous Environments with Dynamic...
https://llm-stats.com/benchmarks/mask
MASK Benchmark Leaderboard
May 12, 2026 - MASK is a collection of 1000 questions measuring whether models faithfully report their beliefs when pressured to lie. It operationalizes deception as the rate...
maskbenchmarkleaderboard
https://llm-stats.com/benchmarks/vision2web
Vision2Web Benchmark Leaderboard
May 20, 2026 - Vision2Web evaluates multimodal models on converting visual designs and screenshots into functional web pages, measuring end-to-end design-to-code capability.
benchmarkleaderboard
https://llm-stats.com/benchmarks/arc
Arc Benchmark Leaderboard
May 8, 2026 - The Abstraction and Reasoning Corpus (ARC) is a benchmark designed to measure human-like general fluid intelligence through grid-based reasoning tasks. It...
arcbenchmarkleaderboard
https://llm-stats.com/benchmarks/humaneval-plus
HumanEval Plus Benchmark Leaderboard
May 12, 2026 - Enhanced version of HumanEval that extends the original test cases by 80x using EvalPlus framework for rigorous evaluation of LLM-synthesized code functional...
humanevalplusbenchmarkleaderboard
https://llm-stats.com/benchmarks/chexpert-cxr
CheXpert CXR Benchmark Leaderboard
May 12, 2026 - CheXpert is a large dataset of 224,316 chest radiographs from 65,240 patients for automated chest X-ray interpretation. The dataset includes uncertainty labels...
benchmarkleaderboard
https://llm-stats.com/benchmarks/ruler-1000k
RULER 1000K Benchmark Leaderboard
May 12, 2026 - RULER 1000K evaluates the official 13-task RULER v1 suite at a 1048576-token (1M) context budget.
rulerbenchmarkleaderboard