Robuta

https://openreview.net/forum?id=SEFSkn4l6d CausalGame: Benchmarking Causal Thinking of LLM Agents in Games | OpenReview Recently, it has received growing attention in building AI Scientist agents with Large Language Models (LLMs). Since scientific discovery fundamentally relies... llm agentsin gamesbenchmarkingcausalthinking