All Stories
737 storiesSort:
Time:
5th October

Epoch estimates OpenAI researchers’ coding-agent use doubled every 34 days
🧠Coding Agents5th October

ContextQA ships Ship to test deployed user flows
Release🛡️Reliability5th October

Pi Durable persists agent state beyond the transcript
Workflow⚙️Durable Execution5th October

Devin adds persistent memory with overnight cleanup
Release🔎Coding Agents5th October

gdp-ts makes authorization proofs compile-time requirements
Release⌨️Security5th October

Reflection releases 501B-parameter Beam model
Release🧠Open Models5th October

Era generates simulated companies for agent testing
Release⚙️Evals5th October

SemiAnalysis reports 4–5× more API-equivalent value from Anthropic subscriptions
🧠Claude5th October
4th October

Claude Code Mods walkthrough says plugins run unsandboxed with user permissions
Workflow⌨️Claude Code4th October

Teknium says Hermes checks nearly 500 plugins for malware, not security hardening
⌨️Hermes Agent4th October

NerfBench publishes method for testing Claude Opus 5.5 performance changes
🧠Claude4th October

Pi Durable agent resumes after redeploying itself with SQLite memory
Workflow⚙️Durable Execution4th October

Geoffrey Huntley demos an SBCL kernel that modifies itself while running
Workflow⚙️Agent runtime infrastructure4th October

Pi Durable Android prototype runs without Termux
Workflow🧠Durable Execution4th October

Matt Pocock releases skills v1.3 with /retro transcript reviews
Release⚙️Agent Skills4th October

Peter Gostev reports Opus gains roughly 250 Elo in agent chess test
🧠Claude4th October
3rd October

T3 Code ships Orchestrator V2 with cross-harness agent delegation
Release⚙️Coding Agents3rd October

Claude Code 2.1.289 fixes Read-deny bypasses via file references and symlinks
Release⌨️Claude Code3rd October

Aleph Alpha releases Kolibri with a 1M-token context window
Release🧠Open Models3rd October

Kevin Kern delegates coding tasks from Opus 5.5 to GPT-6.1 Sol
Workflow🧠Claude Code3rd October

Astra's Elo falls in Peter Gostev's repeated-play agent chess test
🧠Astra3rd October

Pi Durable prototype resumes Android agent work after worker restarts
Workflow🧠Durable Execution3rd October

SemiAnalysis reports configuration failures in Vultr's ClusterMAX test
🧠GPU Infrastructure3rd October
2nd October

Cua launches Spaces for agent-controlled desktops
Release⚙️Computer Use2nd October

Meta reports Muse Spark solutions to six open mathematics problems
🧠Muse2nd October

NerfBench says Claude Opus 5.5's 94.2% score is within normal variance
🧠Claude2nd October

xAI releases an experimental TypeScript SDK for Grok
Release🧠Grok2nd October

OpenAI describes Dot coordinating Codex tasks across apps
⌨️Codex2nd October

T3 Code adds cross-provider child agents in its upcoming nightly overhaul
Release⚙️Coding Agents2nd October

Meta releases Muse Gadgets with ESP32 firmware and a Linux SDK
Release🧠Muse2nd October
1st October

Anthropic adds Claude Code Mods for JavaScript and TypeScript plugins
Release⚙️Claude Code1st October

Modal launches globally available RDMA clusters
Release⚙️Agent runtime infrastructure1st October

Kev releases 1.0 open-weight decision models with a 64k document window
Release🧠Open Models1st October

Amp investigates ChatGPT subscription connection errors
⚙️Coding Agents1st October

Pi releases version 1.0 with durable sessions backed by SQLite
Release⚙️Coding Agents1st October

Rivet reports 0.82 MB per session for its Pi integration
Release⚙️Agent runtime infrastructure1st October

Black Forest Labs launches FLUX 3 Image with native 4K output
Release📈Multimodal1st October
30th September

GPT-6.1 Sol scores 86.3% on MathArena BrokenArXiv
🧠GPT-6.1 Sol30th September

OpenAI adds capacity after GPT-6.1 Sol overloads ChatGPT
🧠GPT-6.1 Sol30th September

Google rolls out Gemini 4 Argon to cyber defenders
Release🧠Gemini30th September

Perplexity open-sources pplx-embed-v2-context-9b-preview
Release🔎RAG30th September

Reports rank Gemini 4 Argon highly on four engineering benchmarks
🧠Gemini30th September
29th September

OpenAI releases GPT-6.1 Sol at $2 per million input tokens
Release🧠GPT-6.1 Sol29th September

Anthropic reports GLM-5.3 built browser exploits in 50 of 410 attempts
🧠GLM29th September

Developers report Firebase payload crashed iOS apps
⌨️Reliability29th September

OpenAI lets partner apps use ChatGPT subscription allowances
Release⚙️OAuth29th September

OpenAI launches GPT-6 Astra Ultrafast at up to 8× Standard speed
Release🧠GPT-6 Astra29th September