https://aclanthology.org/2025.findings-emnlp.835/
ULTRABENCH: Benchmarking LLMs under Extreme Fine-grained Text Generation - ACL Anthology
Longfei Yun, Letian Peng, Jingbo Shang. Findings of the Association for Computational Linguistics: EMNLP 2025. 2025.
benchmarking llmsfine grainedtext generationextreme
https://arxiv.org/abs/2505.20139
[2505.20139] StructEval: Benchmarking LLMs' Capabilities to Generate Structural Outputs
Abstract page for arXiv paper 2505.20139: StructEval: Benchmarking LLMs' Capabilities to Generate Structural Outputs
benchmarking llmscapabilitiesgeneratestructuraloutputs
https://openreview.net/forum?id=L0oSfTroNE
Benchmarking LLMs via Uncertainty Quantification | OpenReview
The proliferation of open-source Large Language Models (LLMs) from various institutions has highlighted the urgent need for comprehensive evaluation methods....
benchmarking llmsuncertainty quantificationviaopenreview
https://arxiv.org/html/2406.10290v1
MobileAIBench: Benchmarking LLMs and LMMs for On-Device Use Cases
benchmarking llmson devicelmmsusecases
https://research.google/blog/benchmarking-llms-for-global-health/
Benchmarking LLMs for global health
benchmarking llmsglobalhealth
https://arxiv.org/abs/2411.07127v1
[2411.07127v1] Benchmarking LLMs' Judgments with No Gold Standard
Abstract page for arXiv paper 2411.07127v1: Benchmarking LLMs' Judgments with No Gold Standard
benchmarking llmsno gold2411judgmentsstandard
https://alphabench.cc/
AlphaBench - Benchmarking LLMs in Formulaic Alpha Factor Mining
The First Comprehensive Evaluation Framework for Formulaic Factor Mining by Large Language Models - ICLR 2026
benchmarking llmsalpha factorformulaicmining
https://openreview.net/forum?id=4diKTLmg2y&referrer=%5Bthe%20profile%20of%20%C3%89tienne%20Marcotte%5D(%2Fprofile%3Fid%3D~%C3%89tienne_Marcotte1)
RepLiQA: A Question-Answering Dataset for Benchmarking LLMs on Unseen Reference Content | OpenReview
Large Language Models (LLMs) are trained on vast amounts of data, most of which is automatically scraped from the internet. This data includes encyclopedic...
https://aclanthology.org/2025.acl-long.445/
White Men Lead, Black Women Help? Benchmarking and Mitigating Language Agency Social Biases in LLMs...
Yixin Wan, Kai-Wei Chang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
https://arxiv.org/abs/2601.21070v1
[2601.21070v1] Towards Comprehensive Benchmarking Infrastructure for LLMs In Software Engineering
Abstract page for arXiv paper 2601.21070v1: Towards Comprehensive Benchmarking Infrastructure for LLMs In Software Engineering
for llms2601towardscomprehensivebenchmarking
https://openreview.net/forum?id=7ccWCDbdM1
VidHal: Benchmarking Hallucinations in Vision LLMs | OpenReview
Vision Large Language Models (VLLMs) are widely acknowledged to be prone to hallucinations. Existing research addressing this problem has primarily been...
in visionbenchmarkinghallucinationsllmsopenreview
https://openreview.net/forum?id=ULEHJkolxB&referrer=%5Bthe%20profile%20of%20Fenrong%20Liu%5D(%2Fprofile%3Fid%3D~Fenrong_Liu1)
LogiConBench: Benchmarking Logical Consistencies of LLMs | OpenReview
Logical consistency, the requirement that statements remain non-contradictory under logical rules, is fundamental for trustworthy reasoning, yet current LLMs...
benchmarkinglogicalconsistenciesllmsopenreview
https://arxiv.org/abs/2505.19819v1
[2505.19819v1] FinLoRA: Benchmarking LoRA Methods for Fine-Tuning LLMs on Financial Datasets
Abstract page for arXiv paper 2505.19819v1: FinLoRA: Benchmarking LoRA Methods for Fine-Tuning LLMs on Financial Datasets
https://aclanthology.org/2024.acl-long.604/
Benchmarking Chinese Commonsense Reasoning of LLMs: From Chinese-Specifics to...
Jiaxing Sun, Weiquan Huang, Jiang Wu, Chenya Gu, Wei Li, Songyang Zhang, Hang Yan, Conghui He. Proceedings of the 62nd Annual Meeting of the Association for...
commonsense reasoningbenchmarkingchinesellmsspecifics
https://research.google/pubs/omnia-de-egotempo-benchmarking-temporal-understanding-of-multi-modal-llms-in-egocentric-videos/
Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric Videos