All Stories
634 storiesSort:
Time:
8th October

Anthropic launches free vulnerability scans for opted-in open-source projects
Release🧠Claude8th October

GPT-6.1 Sol adds Ultrafast mode at up to 8× Standard speed
Release🧠GPT-6.1 Sol8th October

Arena Alignment Index: safety failures roughly double as conversations double in length
🛡️Benchmarks8th October

StepFun releases Step 5 Preview with a 1M-token context window
Release💳Coding Agents8th October

Claude Haiku 5.5 scores 1,587 in Code Arena, about 260 points above Haiku 4.5
Release💳Benchmarks8th October

ChatGPT UI teardown says server-compiled DIL renders client components
🧠GPT8th October

Harvey LAB-AA v1.1 requires hallucination-free answers for benchmark credit
🛡️Benchmarks8th October

Epoch launches Automation Reports to evaluate models on open-ended research tasks
Release🧠GPT-6 Astra8th October
7th October

Vals AI finds recoverable fixes in 67% of MiMo coding training tasks
🧠Reinforcement Learning7th October

Anthropic SDKs add computer-use action loops for Python and TypeScript
Release🧠Claude7th October

OpenAI rolls out Intelligent UI interactive answers in ChatGPT
Release📈GPT7th October

SemiAnalysis reports outages and GPU errors at IREN's Canadian sites
🧠GPU Infrastructure7th October

Theo releases tsc-rs, a Rust TypeScript compiler with claimed compatibility
Release🧠Claude7th October

Perplexity releases pplx-embed-v2-late models for shared text-image retrieval
Release🔎RAG7th October
6th October

OpenAI publishes 722 math manuscripts from an internal model
🛡️Formal Mathematics6th October

Next.js 16.4 enables Cache Components by default in new apps
Release⌨️Developer tools6th October

METR patches Inspect JavaScript bug that could alter reviewed transcripts
🛡️Security6th October

Google releases EmbeddingGemma 2, a 740M-parameter local retrieval model
Release🧠Gemma6th October
5th October

Era generates simulated companies for agent testing
Release⚙️Evals5th October

Epoch estimates OpenAI researchers’ coding-agent use doubled every 34 days
🧠Coding Agents5th October

ContextQA ships Ship to test deployed user flows
Release🛡️Reliability5th October

Pi Durable persists agent state beyond the transcript
Workflow⚙️Durable Execution5th October

Devin adds persistent memory with overnight cleanup
Release🔎Coding Agents5th October

gdp-ts makes authorization proofs compile-time requirements
Release⌨️Security5th October

Reflection releases 501B-parameter Beam model
Release🧠Open Models5th October

SemiAnalysis reports 4–5× more API-equivalent value from Anthropic subscriptions
🧠Claude5th October
4th October

Claude Code Mods walkthrough says plugins run unsandboxed with user permissions
Workflow⌨️Claude Code4th October

Teknium says Hermes checks nearly 500 plugins for malware, not security hardening
⌨️Hermes Agent4th October

NerfBench publishes method for testing Claude Opus 5.5 performance changes
🧠Claude4th October

Pi Durable agent resumes after redeploying itself with SQLite memory
Workflow⚙️Durable Execution4th October

Geoffrey Huntley demos an SBCL kernel that modifies itself while running
Workflow⚙️Agent runtime infrastructure4th October

Pi Durable Android prototype runs without Termux
Workflow🧠Durable Execution4th October

Matt Pocock releases skills v1.3 with /retro transcript reviews
Release⚙️Agent Skills4th October

Peter Gostev reports Opus gains roughly 250 Elo in agent chess test
🧠Claude4th October
3rd October

Aleph Alpha releases Kolibri with a 1M-token context window
Release🧠Open Models3rd October

Claude Code 2.1.289 fixes Read-deny bypasses via file references and symlinks
Release⌨️Claude Code3rd October

T3 Code ships Orchestrator V2 with cross-harness agent delegation
Release⚙️Coding Agents3rd October

Kevin Kern delegates coding tasks from Opus 5.5 to GPT-6.1 Sol
Workflow🧠Claude Code3rd October

Pi Durable prototype resumes Android agent work after worker restarts
Workflow🧠Durable Execution3rd October

Astra's Elo falls in Peter Gostev's repeated-play agent chess test
🧠Astra3rd October

SemiAnalysis reports configuration failures in Vultr's ClusterMAX test
🧠GPU Infrastructure3rd October
2nd October

Cua launches Spaces for agent-controlled desktops
Release⚙️Computer Use2nd October

Meta reports Muse Spark solutions to six open mathematics problems
🧠Muse2nd October

xAI releases an experimental TypeScript SDK for Grok
Release🧠Grok2nd October

NerfBench says Claude Opus 5.5's 94.2% score is within normal variance
🧠Claude2nd October

Meta releases Muse Gadgets with ESP32 firmware and a Linux SDK
Release🧠Muse2nd October

OpenAI describes Dot coordinating Codex tasks across apps
⌨️Codex2nd October

T3 Code adds cross-provider child agents in its upcoming nightly overhaul
Release⚙️Coding Agents2nd October