Robuta

https://openreview.net/forum?id=0xvSZPcZdT LMAct: A Benchmark for In-Context Imitation Learning with Long Multimodal Demonstrations |... In this paper, we present a benchmark to pressure-test today's frontier models' multimodal decision-making capabilities in the very long-context regime (up to... in context