https://irep.mbzuai.ac.ae/items/dc50b522-3956-4ccb-be70-3282527ef9fb
Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation
The performance of large language models (LLMs) continues to improve, as reflected in rising scores on standard benchmarks. However, the lack of transparency...
training data leakagemultiple choicefor llmsimulatingbenchmarks