All Stories
212 storiesSort:
Time:
23rd September

CUA releases Cua-S1-4B-0.2 for computer use
Release🧠Computer Use23rd September

Transluce releases 30,000 logs it says document suspected rogue-agent attacks
⚙️Security23rd September

OpenAI releases MentalHealthBench, an open AI mental-health benchmark
Release📈Evals23rd September

Black Forest Labs releases open 7B FLUX 3 Action model
Release🧠Robotics23rd September

ChatGPT Voice adds email, calendar, and Slack plugins through ChatGPT Work
Release📈GPT23rd September

Meta adds computer use to Muse for Mac
Release🧠Muse23rd September
22nd September

Anthropic releases Claude Opus 5.5 at $4 per million input and $20 per million output tokens
Release🧠Claude22nd September

Firecrawl launches Alexandria API/MCP interface for 100+ external data providers
Release🔎Agent runtime infrastructure22nd September

Anthropic reports Claude Opus 5.5 generated unprompted malicious instructions
🛡️Claude22nd September

OpenAI releases GPT-6 Sol and GPT-6 Luna at roughly half GPT-5.6 API prices
Release🧠GPT22nd September

Xiaomi releases MiMo V2.6 Pro and Flash model weights
Release🧠Open Models22nd September

DigitalOcean launches Managed Agents with persistent runtimes in isolated microVMs
Release🔎Agent runtime infrastructure22nd September
21st September

LangSmith adds Jev as a production trace judge
Workflow🛡️Evals21st September

OpenAI says an internal model solved more than 100 long-standing math problems
🛡️Foundation Model Research21st September

xAI releases Grok 4.7 through coding tools and APIs
Release🧠Grok21st September

Parakeet Redux cuts NVIDIA's Parakeet from 1.2 GB to 178 MB
Release🧠Voice AI21st September

Xiaomi releases open-weight MiMo-V2.6 models with 1M-token context
Release🧠Open Models21st September

Decode's harness verifies coding agents with hidden tests in fresh sandboxes
Workflow⚙️Evals21st September
20th September

Qwen releases Qwen-Image-2.1 with open weights
Release🧠Qwen20th September

Developers test Jev as a low-cost judge for agent evaluations
Workflow🧠Jev20th September

Kev releases open decision models built on Qwen3
Release🧠Qwen20th September

Developers are adopting Jev for document classification
Workflow🧠Jev20th September
19th September
18th September

Claude helped exploit a Discourse flaw affecting OpenAI, researchers report
🧠Claude18th September

Jev cuts Stagehand's median Act latency from 1.97s to 0.46s, Stagehand says
Workflow🧠Jev18th September

CUA releases open-source CUA-S1-FORMS for bounded web-form actions
Release🧠Computer Use18th September

Gemini accessed three real companies during Google's May security tests, Google says
🧠Gemini18th September

Sherpa scores 89.8% on 176 cross-chapter memory questions, Pocket FM says
Release🔎Creative Tools18th September
17th September

OpenAI launches Astra for Law with GPT-6 Astra via Trusted Access
Release🧠GPT-6 Astra1w ago

Goodfire says activation probes could flag reward hacking in real time
🛡️Red Teaming1w ago

Google adds the Antigravity harness to Gemini managed agents
Release⚙️Gemini1w ago

Ramp's 137-task Accounting Bench finds the best model fully solves 21% of tasks
🧠Claude1w ago

Figure says Helix 2.5 raises zero-shot household-task success from 9% to 56% across 30 unseen homes
⚙️Robotics1w ago

Exa launches a historical web index with 400 billion snapshots
Release🛡️Search1w ago

Raindrop opens Simulations for agent changes on every pull request
Release🛡️Evals1w ago

