Robuta

https://openreview.net/forum?id=kvjbFVHpny&referrer=%5Bthe%20profile%20of%20Binhua%20Li%5D(%2Fprofile%3Fid%3D~Binhua_Li1) EvoCodeBench: An Evolving Code Generation Benchmark with Domain-Specific Evaluations | OpenReview How to evaluate Large Language Models (LLMs) in code generation remains an open question. Many benchmarks have been proposed, but they have two limitations,... code generationdomain specificevolving https://arxiv.org/abs/2602.10171 [2602.10171] EvoCodeBench: A Human-Performance Benchmark for Self-Evolving LLM-Driven Coding Systems Abstract page for arXiv paper 2602.10171: EvoCodeBench: A Human-Performance Benchmark for Self-Evolving LLM-Driven Coding Systems