https://openreview.net/forum?id=q2DmkZ1wVe
CofCA: A STEP-WISE Counterfactual Multi-hop QA benchmark | OpenReview
While Large Language Models (LLMs) excel in question-answering (QA) tasks, their real reasoning abilities on multiple evidence retrieval and integration on...
a stepwisemultihopqa