https://www.2077ai.com/datasets/dataset-criticlean
CriticLeanBench: A Benchmark for Evaluating Mathematical Formalization Critics - 2077AI
CriticLeanBench is a specialized benchmark designed to evaluate the critical reasoning of AI models, specifically on the task of validating the translation of...
benchmarkevaluatingmathematicalformalizationcritics