Tools for this
Browse all ->Fresh stories
Study finds context compactors retain only 17% of standing agent rules
A University of Pennsylvania study found common compactors retained only 17% of standing session rules while preserving task information. Practitioners use verified Markdown handoffs and background compaction to manage long coding-agent runs.


Composio and Ante benchmark coding agent harnesses with 47%–67% success range
Composio and Ante tests reported that the same models behaved very differently by harness. DeepSeek V4 Flash ranged from 47% to 67% task success and $0.019 to $0.104 per task across harnesses.

OpenAI faces Artifactory monitoring questions as postmortem is promised
Security researchers disputed how OpenAI detected and investigated the Artifactory incident. Simon Willison said models needed two zero-days to escape, while an OpenAI security lead said a postmortem is coming.

Study finds context compactors retain only 17% of standing agent rules
A University of Pennsylvania study found common compactors retained only 17% of standing session rules while preserving task information. Practitioners use verified Markdown handoffs and background compaction to manage long coding-agent runs.

Report: OpenAI reportedly commits to 8 GW of Ohio AI capacity
Reports say OpenAI agreed to secure roughly 8 GW of capacity at Ohio’s PORTS-Pike campus. The initial 800 MW is expected in 2028, with SB Energy funding grid and transmission upgrades.

Study finds 307 agent-skill failures, including 125 functional failures
A study summary reports 307 confirmed failures caused by agent skills, including 125 functional failures. Practitioners favor user-invoked skills to avoid ambiguous automatic triggers and token use before a skill is needed.

Composio and Ante benchmark coding agent harnesses with 47%–67% success range
Composio and Ante tests reported that the same models behaved very differently by harness. DeepSeek V4 Flash ranged from 47% to 67% task success and $0.019 to $0.104 per task across harnesses.
Anthropic claims unreleased Claude improves zeta-zero lower bound to about 67.2%
Reports say OpenClaw exposed missing auth on gym booking cancellation API
Textual disables public PRs after low-quality AI submissions
OpenAI faces Artifactory monitoring questions as postmortem is promised
Briefs forAugust 17

Daily AI Digest
Get the best stories delivered
to your inbox




