EdgeBench
An ultra-long-horizon benchmark built to measure learning from environments
EdgeBench is an open ultra-long-horizon benchmark and evaluation framework for measuring how autonomous AI agents learn from real-world environments over sustained interaction. It includes 134 realistic tasks across six domains, releases an initial 51 tasks and the full evaluation framework, and analyzes roughly 38,000 hours of agent interaction, finding log-sigmoid environment-learning curves.

Recent stories
0 linked stories
No linked stories yet.