All Stories
346 storiesSort:
Time:
26th August

METR and Redwood Research document 1,200 coordinated agents in Hugging Face incident
🛡️Security26th August

Google launches Gemini 3.5 Transcribe API with 85+ languages
Release🧠Gemini26th August

Qwen releases Qwen3.8-Flash open weights: 125B MoE with 262K context
Release🧠Qwen26th August

Perceptron releases Isaac 0.5 open weights for robot control
Release🧠Robotics26th August

Z.ai releases GLM-5.3-Flash, identifies it as Ox Alpha
Release🧠GLM26th August

Anthropic adds an isolated browser panel to Claude Cowork
Release🧠Claude26th August

Glean says runtime routing cuts enterprise-agent token costs by 81%
🧠Model Routing26th August

Anthropic opens 250,000 privacy-preserved Claude conversations to researchers
🛡️Claude26th August
25th August

Perplexity launches Portable Computer on DGX Spark with a 27B model
Release🧠Edge Compute25th August

ChatGPT browser adds WebMCP support for sites as tools
Release⚙️Codex25th August

OpenAI reports Jalapeño delivers 1.5–1.9× more work per watt
🧠GPU Infrastructure25th August

Prime Intellect publishes Prime Agent report with 7-day Factorio evaluation
🔎Harness Engineering25th August

Figure launches Index robot-training dataset with 16M video uploads
Release🧠Robotics25th August

Vercel Connect reaches GA with authenticated MCP for 100+ services
Release⚙️MCP25th August

OpenAI launches $100 ChatGPT Business Premium seats
Release⌨️Codex25th August
24th August

OpenAI cuts GPT-5.6 Sol API prices by up to 33%
🧠GPT24th August

OpenRouter says Ox Alpha reaches 8 trillion daily tokens
🧠Ox Alpha24th August

DataSpace finds harnesses shift data-task accuracy by 15 points
⚙️Harness Engineering24th August

NVIDIA puts Groq 3 LPX into Vera Rubin production
🧠Gemma24th August

ASI-Bench finds full procedures lift research-agent scores to 50.91
⚙️Agent Skills24th August

WAN 3.0 launches on API platforms with 30-second video output
Release💳Multimodal24th August
23rd August

Study finds instructions account for 60.5% of coding-agent reading
⚙️Context Engineering23rd August

OpenAI fixes Codex long-session usage accounting
⌨️Codex23rd August

Practitioners propose a standard harness for agent benchmarks
🛡️Harness Engineering23rd August

Together reports GLM-5.3 solves 87.6% of DeepSWE work for about $16
🧠GLM23rd August
22nd August

Tests link Ox Alpha to Zhipu GLM API routes and error codes
🧠Ox Alpha22nd August

Anthropic says serving test remapped Claude Code effort settings
🧠Claude Code22nd August

Qwen 3.8 27B reaches 3,200 TPM at 262K context on two RTX 3090s
🧠Qwen22nd August

Independent DeepSWE retest puts Ox Alpha at about 63%
🧠Ox Alpha22nd August
21st August

OpenAI cuts GPT-5.6 Sol API rates to $4 input and $20 output
🧠GPT21st August

OpenAI grants Codex customers a banked usage reset
💳Codex21st August

Study finds CLI-first agents cost 5–28x less than MCP agents
⚙️MCP21st August

multiPL-E regex bug corrupts MBPP benchmark language variants
⌨️Codex21st August

Claude Security adds Mythos 5 scans for GitHub repositories
Release🧠Claude21st August

Marin starts training open 535B-A23B model on 18.75T tokens
🧠Open Models21st August

NVIDIA AVO reportedly reaches 100% on ARC-AGI-3 demo tasks
🛡️Benchmarks21st August

Users report Qwen 3.8 27B agents vary sharply by harness
Workflow🧠Qwen21st August
20th August

Meta publishes Muse Spark 1.2 multimodal evaluation results
Release🧠Muse1w ago

ARC Prize verifies Gemini 3.7 Flash at 84.6% on ARC-AGI-2
🧠Gemini1w ago

OpenRouter tests free Ox Alpha with a 1M-token context window
Release🧠Ox Alpha1w ago

Nous launches Hermes Agent with managed remote computers
Release⌨️Hermes Agent1w ago

Transluce trains 8B–1.1T activation-reading oversight models
🛡️Interpretability1w ago

LocalLLaMA post reports Qwen 3.8 27B reasoning loops caused most errors
🧠Qwen1w ago