Cua-Bench
A benchmark for computer-use agents on professional software
A computer-use benchmark for evaluating agents on professional GUI software, initially comprising 25 expert-authored KiCad tasks graded using exact netlist matches.

Recent stories
CUA released Cua-S1-4B-0.2, a multimodal decision model trained with supervised learning and task-completion RL in live environments. CUA reports 92.9% on a frozen GUI-360 split and released adapters and training code.
TryCua and Snorkel opened Cua-Bench, a computer-use benchmark with 25 expert-authored KiCad tasks graded by exact netlist matches. The early results show frontier models still struggle with GUI execution, wiring completion, and self-checking, so treat benchmark wins as incomplete for real computer-use work.