Tools for this
Browse all ->Fresh stories
Study finds 307 agent-skill failures, including 125 functional failures
A study summary reports 307 confirmed failures caused by agent skills, including 125 functional failures. Practitioners favor user-invoked skills to avoid ambiguous automatic triggers and token use before a skill is needed.

Qwen releases Qwen3.8 27B multimodal model under Apache 2.0
Qwen released its open-weight Qwen3.8 27B vision-language model with 262K native context and adjustable reasoning. In a 484-sample test, enabled-thinking scores fell from above 92% through 64K to 74.3–81.8% at 128K.

Anthropic claims unreleased Claude improves zeta-zero lower bound to about 67.2%
Anthropic says an unreleased Claude did not solve the Riemann hypothesis but improved a related zeta-zero lower bound from 41.6% to about 67.2%. Posts describe subagents, expert prompting, and Lean formalization.


GitHub users report outage disrupting commits and Actions
Users reported failures retrieving commits, running Actions, and syncing GitHub-hosted repositories. Some developers said the disruption blocked pull-request merges amid broader reliability complaints.

Study finds 307 agent-skill failures, including 125 functional failures
A study summary reports 307 confirmed failures caused by agent skills, including 125 functional failures. Practitioners favor user-invoked skills to avoid ambiguous automatic triggers and token use before a skill is needed.

Anthropic plans Claude text-watermark detection API
Anthropic plans global Claude watermarking based on statistical word-choice patterns rather than hidden characters or user identifiers. A MIT-licensed removal repository reportedly reached roughly 11,000 GitHub stars shortly after the announcement.

Reports: GLM-5.3 post-training lifts Terminal-Bench from 4.6 to 28.3
Reports say GLM-5.3 retained GLM-5.2's base model while post-training raised Terminal-Bench from 4.6 to 28.3 and DeepSWE from 46.2 to 66.9. A technical account attributes the gains to RL infrastructure changes.
Qwen releases Qwen3.8 27B multimodal model under Apache 2.0
Qwen 3.8 Max 2.4T open weights ship as text-only, users say
Composio and Ante benchmark coding agent harnesses with 47%–67% success range
Anthropic claims unreleased Claude improves zeta-zero lower bound to about 67.2%
Briefs forAugust 17

Daily AI Digest
Get the best stories delivered
to your inbox



