All Stories
798 storiesSort:
Time:
4th October

Claude Code Mods walkthrough says plugins run unsandboxed with user permissions
Workflow⌨️Claude Code4th October

Teknium says Hermes checks nearly 500 plugins for malware, not security hardening
⌨️Hermes Agent4th October

Pi Durable agent resumes after redeploying itself with SQLite memory
Workflow⚙️Durable Execution4th October

Geoffrey Huntley demos an SBCL kernel that modifies itself while running
Workflow⚙️Agent runtime infrastructure4th October

Pi Durable Android prototype runs without Termux
Workflow🧠Durable Execution4th October

NerfBench publishes method for testing Claude Opus 5.5 performance changes
🧠Claude4th October

Matt Pocock releases skills v1.3 with /retro transcript reviews
Release⚙️Agent Skills4th October

Peter Gostev reports Opus gains roughly 250 Elo in agent chess test
🧠Claude4th October
3rd October

Aleph Alpha releases Kolibri with a 1M-token context window
Release🧠Open Models3rd October

Claude Code 2.1.289 fixes Read-deny bypasses via file references and symlinks
Release⌨️Claude Code3rd October

T3 Code ships Orchestrator V2 with cross-harness agent delegation
Release⚙️Coding Agents3rd October

SemiAnalysis reports configuration failures in Vultr's ClusterMAX test
🧠GPU Infrastructure3rd October

Astra's Elo falls in Peter Gostev's repeated-play agent chess test
🧠Astra3rd October

Kevin Kern delegates coding tasks from Opus 5.5 to GPT-6.1 Sol
Workflow🧠Claude Code3rd October

Pi Durable prototype resumes Android agent work after worker restarts
Workflow🧠Durable Execution3rd October
2nd October

Cua launches Spaces for agent-controlled desktops
Release⚙️Computer Use2nd October

NerfBench says Claude Opus 5.5's 94.2% score is within normal variance
🧠Claude2nd October

xAI releases an experimental TypeScript SDK for Grok
Release🧠Grok2nd October

Meta reports Muse Spark solutions to six open mathematics problems
🧠Muse2nd October

Meta releases Muse Gadgets with ESP32 firmware and a Linux SDK
Release🧠Muse2nd October

OpenAI describes Dot coordinating Codex tasks across apps
⌨️Codex2nd October

T3 Code adds cross-provider child agents in its upcoming nightly overhaul
Release⚙️Coding Agents2nd October
1st October

Modal launches globally available RDMA clusters
Release⚙️Agent runtime infrastructure1st October

Anthropic adds Claude Code Mods for JavaScript and TypeScript plugins
Release⚙️Claude Code1st October

Amp investigates ChatGPT subscription connection errors
⚙️Coding Agents1st October

Pi releases version 1.0 with durable sessions backed by SQLite
Release⚙️Coding Agents1st October

Rivet reports 0.82 MB per session for its Pi integration
Release⚙️Agent runtime infrastructure1st October

Kev releases 1.0 open-weight decision models with a 64k document window
Release🧠Open Models1st October

Black Forest Labs launches FLUX 3 Image with native 4K output
Release📈Multimodal1st October
30th September

Perplexity open-sources pplx-embed-v2-context-9b-preview
Release🔎RAG30th September

GPT-6.1 Sol scores 86.3% on MathArena BrokenArXiv
🧠GPT-6.1 Sol30th September

OpenAI adds capacity after GPT-6.1 Sol overloads ChatGPT
🧠GPT-6.1 Sol30th September

Google rolls out Gemini 4 Argon to cyber defenders
Release🧠Gemini30th September

Reports rank Gemini 4 Argon highly on four engineering benchmarks
🧠Gemini30th September
29th September

Anthropic reports GLM-5.3 built browser exploits in 50 of 410 attempts
🧠GLM29th September

Developers report Firebase payload crashed iOS apps
⌨️Reliability29th September

OpenAI lets partner apps use ChatGPT subscription allowances
Release⚙️OAuth29th September

OpenAI releases GPT-6.1 Sol at $2 per million input tokens
Release🧠GPT-6.1 Sol29th September

OpenAI launches GPT-6 Astra Ultrafast at up to 8× Standard speed
Release🧠GPT-6 Astra29th September
28th September

NVIDIA launches OpenShell to isolate AI agents
Release⚙️Sandboxing28th September

Anthropic releases Claude Sonnet 5.5 at unchanged token prices
Release🧠Claude28th September

H Company releases 27B and 35B-A3B Holo4 computer-use models
Release🧠Computer Use28th September

ARC Prize reports lower matched-effort scores for GPT-6 Sol than GPT-5.6 Sol
🧠GPT28th September

Independent tests put Sonnet 5.5 at $7.60 per max-effort task
🧠Claude28th September

Manus launches its 2.0 agent platform with cloud computers and automations
Release⚙️Agent product updates28th September