Robuta

https://openreview.net/forum?id=q2DmkZ1wVe CofCA: A STEP-WISE Counterfactual Multi-hop QA benchmark | OpenReview While Large Language Models (LLMs) excel in question-answering (QA) tasks, their real reasoning abilities on multiple evidence retrieval and integration on... a stepwisemultihopqa