Tools for this
Browse all ->Fresh stories
Study finds context compactors retain only 17% of standing agent rules
A University of Pennsylvania study found common compactors retained only 17% of standing session rules while preserving task information. Practitioners use verified Markdown handoffs and background compaction to manage long coding-agent runs.

Codex opens 1M-token GPT-5.6 Sol context for ChatGPT subscribers
Codex now lets ChatGPT subscribers enable a 1 million-token context window for GPT-5.6 Sol. Automatic compaction starts at 900,000 tokens, and tokens beyond the default window count double against limits.

Agent builders test task-specific harnesses; AGENTS.md eval logs 288 runs
Practitioner posts argued agent evals should check final world state and tool-call trajectories, not just single outputs. A 288-run AGENTS.md test found context files did not improve correctness.


Study finds context compactors retain only 17% of standing agent rules
A University of Pennsylvania study found common compactors retained only 17% of standing session rules while preserving task information. Practitioners use verified Markdown handoffs and background compaction to manage long coding-agent runs.

Codex opens 1M-token GPT-5.6 Sol context for ChatGPT subscribers
Codex now lets ChatGPT subscribers enable a 1 million-token context window for GPT-5.6 Sol. Automatic compaction starts at 900,000 tokens, and tokens beyond the default window count double against limits.

Speculative decoding tests report acceptance drop from 0.71 to 0.18 after ~32K context
A practitioner report found speculative-decoding acceptance fell from 0.71 to 0.18 beyond about 32K context. Separate DSpark and mlx-dspark tests reported speedups on RTX and Apple Silicon setups.

Echo Gap paper reports agents endorsed 31%–54% of their own wrong answers
The Echo Gap paper found self-improving agents can store wrongly self-scored episodes. Tested models endorsed 31% to 54% of their own wrong answers, while other work proposed RL-trained harness state and in-model memory.
Agent builders compare thin harnesses with large skill files for coding agents
Agent builders test task-specific harnesses; AGENTS.md eval logs 288 runs
LightOn releases mDenseOn and mLateOn retrieval models for 8 languages
Agent skills cause regressions in nearly 6,000 paired office-automation runs
Briefs forAugust 17

Daily AI Digest
Get the best stories delivered
to your inbox




