Robuta

https://www.2077ai.com/datasets/dataset-criticlean CriticLeanBench: A Benchmark for Evaluating Mathematical Formalization Critics - 2077AI CriticLeanBench is a specialized benchmark designed to evaluate the critical reasoning of AI models, specifically on the task of validating the translation of... benchmarkevaluatingmathematicalformalizationcritics