Robuta

https://aclanthology.org/2025.findings-emnlp.835/ ULTRABENCH: Benchmarking LLMs under Extreme Fine-grained Text Generation - ACL Anthology Longfei Yun, Letian Peng, Jingbo Shang. Findings of the Association for Computational Linguistics: EMNLP 2025. 2025. benchmarking llmsfine grainedtext generationextreme https://arxiv.org/abs/2505.20139 [2505.20139] StructEval: Benchmarking LLMs' Capabilities to Generate Structural Outputs Abstract page for arXiv paper 2505.20139: StructEval: Benchmarking LLMs' Capabilities to Generate Structural Outputs benchmarking llmscapabilitiesgeneratestructuraloutputs https://openreview.net/forum?id=L0oSfTroNE Benchmarking LLMs via Uncertainty Quantification | OpenReview The proliferation of open-source Large Language Models (LLMs) from various institutions has highlighted the urgent need for comprehensive evaluation methods.... benchmarking llmsuncertainty quantificationviaopenreview https://arxiv.org/html/2406.10290v1 MobileAIBench: Benchmarking LLMs and LMMs for On-Device Use Cases benchmarking llmson devicelmmsusecases https://research.google/blog/benchmarking-llms-for-global-health/ Benchmarking LLMs for global health benchmarking llmsglobalhealth https://arxiv.org/abs/2411.07127v1 [2411.07127v1] Benchmarking LLMs' Judgments with No Gold Standard Abstract page for arXiv paper 2411.07127v1: Benchmarking LLMs' Judgments with No Gold Standard benchmarking llmsno gold2411judgmentsstandard https://alphabench.cc/ AlphaBench - Benchmarking LLMs in Formulaic Alpha Factor Mining The First Comprehensive Evaluation Framework for Formulaic Factor Mining by Large Language Models - ICLR 2026 benchmarking llmsalpha factorformulaicmining https://openreview.net/forum?id=4diKTLmg2y&referrer=%5Bthe%20profile%20of%20%C3%89tienne%20Marcotte%5D(%2Fprofile%3Fid%3D~%C3%89tienne_Marcotte1) RepLiQA: A Question-Answering Dataset for Benchmarking LLMs on Unseen Reference Content | OpenReview Large Language Models (LLMs) are trained on vast amounts of data, most of which is automatically scraped from the internet. This data includes encyclopedic... https://aclanthology.org/2025.acl-long.445/ White Men Lead, Black Women Help? Benchmarking and Mitigating Language Agency Social Biases in LLMs... Yixin Wan, Kai-Wei Chang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. https://arxiv.org/abs/2601.21070v1 [2601.21070v1] Towards Comprehensive Benchmarking Infrastructure for LLMs In Software Engineering Abstract page for arXiv paper 2601.21070v1: Towards Comprehensive Benchmarking Infrastructure for LLMs In Software Engineering for llms2601towardscomprehensivebenchmarking https://openreview.net/forum?id=7ccWCDbdM1 VidHal: Benchmarking Hallucinations in Vision LLMs | OpenReview Vision Large Language Models (VLLMs) are widely acknowledged to be prone to hallucinations. Existing research addressing this problem has primarily been... in visionbenchmarkinghallucinationsllmsopenreview https://openreview.net/forum?id=ULEHJkolxB&referrer=%5Bthe%20profile%20of%20Fenrong%20Liu%5D(%2Fprofile%3Fid%3D~Fenrong_Liu1) LogiConBench: Benchmarking Logical Consistencies of LLMs | OpenReview Logical consistency, the requirement that statements remain non-contradictory under logical rules, is fundamental for trustworthy reasoning, yet current LLMs... benchmarkinglogicalconsistenciesllmsopenreview https://arxiv.org/abs/2505.19819v1 [2505.19819v1] FinLoRA: Benchmarking LoRA Methods for Fine-Tuning LLMs on Financial Datasets Abstract page for arXiv paper 2505.19819v1: FinLoRA: Benchmarking LoRA Methods for Fine-Tuning LLMs on Financial Datasets https://aclanthology.org/2024.acl-long.604/ Benchmarking Chinese Commonsense Reasoning of LLMs: From Chinese-Specifics to... Jiaxing Sun, Weiquan Huang, Jiang Wu, Chenya Gu, Wei Li, Songyang Zhang, Hang Yan, Conghui He. Proceedings of the 62nd Annual Meeting of the Association for... commonsense reasoningbenchmarkingchinesellmsspecifics https://research.google/pubs/omnia-de-egotempo-benchmarking-temporal-understanding-of-multi-modal-llms-in-egocentric-videos/ Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric Videos