https://openreview.net/forum?id=SEFSkn4l6d
CausalGame: Benchmarking Causal Thinking of LLM Agents in Games | OpenReview
Recently, it has received growing attention in building AI Scientist agents with Large Language Models (LLMs). Since scientific discovery fundamentally relies...
llm agentsin gamesbenchmarkingcausalthinking