An agent-automation benchmark and evaluation harness for assessing whether AI agents can complete automation tasks, including the reliability of their trajectories and verifiers.

Recent stories
1 linked story
An agent-automation benchmark and evaluation harness for assessing whether AI agents can complete automation tasks, including the reliability of their trajectories and verifiers.
