AA-Briefcase
Agentic Knowledge Work Benchmark
An Artificial Analysis benchmark for evaluating agentic knowledge-work performance on multi-file office tasks, reported with metrics including score and Elo.

Recent stories
Artificial Analysis reports Kimi K3 averages 56.4 minutes, 83 turns, and 120k output tokens per AA-Briefcase task. Kilo also found UI-build outputs close to Claude Fable 5 at 29% of the cost.
Posts put Muse Spark 1.1 ahead of GPT-5.6 Sol and Gemini 3.1 on Radiology's Last Exam, third on Debate Benchmark, and at 863 Elo on AA-Briefcase. The same posts noted weaker presentation quality.
Genspark and OpenClaw added Grok 4.5 after xAI's launch, extending the model into more coding-agent workflows. Follow-up evidence covered AA-Briefcase and Terminal-Bench results, a Composio credential-audit run, and SuperGrok usage-meter reports.