AA-Briefcase
Agentic Knowledge Work Benchmark
AA-Briefcase is Artificial Analysis' private agentic knowledge-work benchmark for evaluating AI agents and models on long-horizon business workflows that require deliverables such as spreadsheets, presentations, memos, financial models, and design mock-ups.

Recent stories
Artificial Analysis reports Kimi K3 averages 56.4 minutes, 83 turns, and 120k output tokens per AA-Briefcase task. Kilo also found UI-build outputs close to Claude Fable 5 at 29% of the cost.
Posts put Muse Spark 1.1 ahead of GPT-5.6 Sol and Gemini 3.1 on Radiology's Last Exam, third on Debate Benchmark, and at 863 Elo on AA-Briefcase. The same posts noted weaker presentation quality.
Genspark and OpenClaw added Grok 4.5 after xAI's launch, extending the model into more coding-agent workflows. Follow-up evidence covered AA-Briefcase and Terminal-Bench results, a Composio credential-audit run, and SuperGrok usage-meter reports.