Robuta

https://ai-stats.phaseo.app/benchmarks/nyt-connections NYT Connections - Benchmark Leaderboard & Model Performance | AI Stats NYT Connections benchmark leaderboard on AI Stats. See 49 scored models, track historical performance, and inspect the underlying methodology. Current top... nyt connectionsbenchmark leaderboardperformance aimodelstats https://ai-stats.phaseo.app/benchmarks/simpleqa SimpleQA - Benchmark Leaderboard & Model Performance | AI Stats SimpleQA benchmark leaderboard on AI Stats. See 91 scored models, track historical performance, and inspect the underlying methodology. Current top model:... benchmark leaderboardperformance aisimpleqamodelstats https://llm-stats.com/benchmarks/bfcl-v3-multiturn BFCL_v3_MultiTurn Benchmark Leaderboard May 12, 2026 - Berkeley Function Calling Leaderboard (BFCL) V3 MultiTurn benchmark that evaluates large language models' ability to handle multi-turn and multi-step function... multiturnbenchmarkleaderboard https://llm-stats.com/benchmarks/mmau-sound MMAU Sound Benchmark Leaderboard May 16, 2026 - A subset of the MMAU benchmark focused specifically on environmental sound understanding and reasoning tasks. Part of a comprehensive multimodal audio... soundbenchmarkleaderboard https://llm-stats.com/benchmarks/mmbench-video MMBench-Video Benchmark Leaderboard May 15, 2026 - A long-form multi-shot benchmark for holistic video understanding that incorporates approximately 600 web videos from YouTube spanning 16 major categories,... videobenchmarkleaderboard https://llm-stats.com/benchmarks/qmsum QMSum Benchmark Leaderboard May 16, 2026 - QMSum is a benchmark for query-based multi-domain meeting summarization consisting of 1,808 query-summary pairs over 232 meetings across academic, product, and... benchmarkleaderboard https://llm-stats.com/benchmarks/gpqa-chemistry GPQA Chemistry Benchmark Leaderboard May 16, 2026 - Chemistry subset of GPQA, containing challenging multiple-choice questions written by domain experts in chemistry. These Google-proof questions require... gpqachemistrybenchmarkleaderboard https://llm-stats.com/benchmarks/finsearchcomp-t3 FinSearchComp-T3 Benchmark Leaderboard May 20, 2026 - FinSearchComp-T3 is a benchmark for evaluating financial search and reasoning capabilities, testing models' ability to retrieve and analyze financial... benchmarkleaderboard https://llm-stats.com/benchmarks/cruxeval-output-cot CRUXEval-Output-CoT Benchmark Leaderboard May 16, 2026 - CRUXEval-O (output prediction) with Chain-of-Thought prompting. Part of the CRUXEval benchmark consisting of 800 Python functions (3-13 lines) designed to... outputcotbenchmarkleaderboard https://llm-stats.com/benchmarks/livecodebench-v6 LiveCodeBench v6 Benchmark Leaderboard May 20, 2026 - LiveCodeBench is a holistic and contamination-free evaluation benchmark for large language models for code. It continuously collects new problems from... livecodebenchbenchmarkleaderboard https://llm-stats.com/benchmarks/activitynet ActivityNet Benchmark Leaderboard May 20, 2026 - A large-scale video benchmark for human activity understanding. Provides samples from 203 activity classes with an average of 137 untrimmed videos per class... benchmarkleaderboard https://www.sobigdata.eu/blog/building-cd-lead-benchmark-and-leaderboard-community-detection Building CD-LEAD: a benchmark and leaderboard for community detection | SoBigData.eu https://llm-stats.com/benchmarks/cnmo-2024 CNMO 2024 Benchmark Leaderboard May 16, 2026 - China Mathematical Olympiad 2024 - A challenging mathematics competition. cnmobenchmarkleaderboard https://ha-vln-project.vercel.app/ An Open Benchmark and Leaderboard for Human-Aware Navigation in Discrete and Continuous... Project page of paper 'HA-VLN 2.0: An Open Benchmark and Leaderboard for Human-Aware Navigation in Discrete and Continuous Environments with Dynamic... https://llm-stats.com/benchmarks/mask MASK Benchmark Leaderboard May 12, 2026 - MASK is a collection of 1000 questions measuring whether models faithfully report their beliefs when pressured to lie. It operationalizes deception as the rate... maskbenchmarkleaderboard https://llm-stats.com/benchmarks/vision2web Vision2Web Benchmark Leaderboard May 20, 2026 - Vision2Web evaluates multimodal models on converting visual designs and screenshots into functional web pages, measuring end-to-end design-to-code capability. benchmarkleaderboard https://llm-stats.com/benchmarks/arc Arc Benchmark Leaderboard May 8, 2026 - The Abstraction and Reasoning Corpus (ARC) is a benchmark designed to measure human-like general fluid intelligence through grid-based reasoning tasks. It... arcbenchmarkleaderboard https://llm-stats.com/benchmarks/humaneval-plus HumanEval Plus Benchmark Leaderboard May 12, 2026 - Enhanced version of HumanEval that extends the original test cases by 80x using EvalPlus framework for rigorous evaluation of LLM-synthesized code functional... humanevalplusbenchmarkleaderboard https://llm-stats.com/benchmarks/chexpert-cxr CheXpert CXR Benchmark Leaderboard May 12, 2026 - CheXpert is a large dataset of 224,316 chest radiographs from 65,240 patients for automated chest X-ray interpretation. The dataset includes uncertainty labels... benchmarkleaderboard https://llm-stats.com/benchmarks/ruler-1000k RULER 1000K Benchmark Leaderboard May 12, 2026 - RULER 1000K evaluates the official 13-task RULER v1 suite at a 1048576-token (1M) context budget. rulerbenchmarkleaderboard