Agent Skills
Stories, products, and related signals connected to this tag in Explore.
Stories
Filter storiesNext.js 16.4 made Cache Components the default for new apps and added APIs for static output and partial navigation rendering. The release also adds agent-guided upgrades and cuts Turbopack disk-cache size by 20–25%.
Matt Pocock's skills v1.3 adds /retro to find workflow improvements in old agent transcripts. Its migration prompt compares installed skills, renames CONTEXT.md to GLOSSARY.md and reviews recent skill usage.
Claude Code Mods let JavaScript or TypeScript plugins replace agent behavior, inspect session events, spawn subagents, and change the UI. Practitioners have shared a mod-building skill and hot-reload workflows.
A case study reports reducing a roughly 30,000-token AGENTS.md file by almost 50% while improving instruction quality. Session transcripts informed which rules remained and which moved into a skill.
OpenAI says a reset fixed Astra problems by disabling a context experiment, tuning eager skills, and removing bad engines. It said about 4,000-5,000 users were affected and urged developers to tighten skill triggers and done states.
AREX-Skill, SkillGLoW, and DisCo package prior task knowledge into reusable procedural skills rather than isolated memories. DisCo reports MLE-bench rising from 31.11% to 72.89% with the same model.
Across 60 research projects, ASI-Bench found full procedures averaged 50.91, versus 29.10 for prompts naming only a method. Other evaluations similarly measure whether procedural skills improve execution rather than merely adding more instructions.
A study of 557 coding-agent sessions finds instruction files and working notes account for most of what agents read. Related work puts installed skills' standing prompt cost at 50–280 tokens, while Backpass turns past sessions into reviewable AGENTS.M files.
A study summary reports 307 confirmed failures caused by agent skills, including 125 functional failures. Practitioners favor user-invoked skills to avoid ambiguous automatic triggers and token use before a skill is needed.
Google DeepMind introduced SkillSmith, a method that treats prefix weights or KV-cache states as an input modality so a frozen Gemma 3 4B model can synthesize new skill prefixes at inference time. Reported Composite-SNI Elo improved when cache composition was combined with text descriptions, making it a research artifact rather than a deployable runtime.
New guides, plugins, and reusable libraries show the Agent Skills format moving beyond Claude Code into multiple coding-agent clients and runtimes. That matters because workflows are becoming portable artifacts instead of one-off prompts tied to a single harness.