Where people deep in AI come to stay current.
Category
Tags
Independent AI model and API provider benchmarking
The platform for agent engineering.
Benchmark Generative AI for Enterprise Applications.
From scan to fix, done seamlessly
Secure, elastic infrastructure for AI agents.
Agentic Knowledge Work Benchmark
One API for all AI models.
Build & Test with AI Coding Models
Cua Spaces provides agent-controlled desktops on local or remote machines. The free, source-available release includes approved app-session transfers, a locked local Keyvault, and separate human and agent cursors.
Meta published six research collaborations reporting solutions to open mathematics problems with Muse Spark. Mathematicians guided the work through its regular chat interface, independently reviewed results, and marked human contributions.
NerfBench reports Claude Opus 5.5 at 94.2% of its launch baseline but says the drop remains within normal variance. Theo disputes the degradation framing and offers to fund independently audited testing.
OpenAI DevDay notes highlight four levers for agent cost control
Vik Paruchuri shows five ways PDFs misrepresent text positions
Sharpening Tax measures lost test-time scalability after post-training
OpenClaw v2026.9.8 adds GPT-6.1 Sol and fixes agent reply handling
A documentation folder rename lifts an agent's evaluation score by five points
Hamel Husain recommends binary checks over 1-to-5 LLM eval ratings
Meta's SWE-sweep tests whether agents can find bugs without issue hints
Researchers identify a ChatGPT macOS flaw that could expose local data
CloudSEK traces MALFEX malware to compromised npm packages
AgentBreaker tests context-aware indirect prompt injection in web agents
Apple plans additional macOS Full Disk Access controls for AI agents
A study finds reasoning models can automate multi-turn jailbreaks
AgentXploit rediscovers known agent-framework flaws at 59.3% success
A review of 1,123 live Bolt apps finds exposed keys and open databases
Google Research shifts federated-learning computation server-side with verifiable privacy
Anthropic prevents new Claude shared chats from appearing in search
A malicious dedh-devops-automation package steals AWS credentials
BeyondTrust details a Vertex AI Agent Engine sandbox response-channel hijack
Vercel AI SDK fixes an ACP harness flaw allowing unrequested host-tool calls
GitLab patches a critical Duo AI Gateway sandbox escape
ProVer uses rollout comparisons to assign credit during agent training
Arena reports Sonnet 5.5 displaces GPT-6.1 Sol from WebDev cost frontier
Google's Cogentic uses separate agents to explore and verify proofs
Jason Zhou reports Jev was 10× cheaper and 18× faster than GPT6-luna in team tests
Google reports Cogentic agents proved five unsolved math problems
Cursor Rollouts links detected regressions to one-click cloud-agent fixes
Google Research introduces TEE-based federated learning for verifiable differential privacy
OpenAI fixes the missed allowance reset for Pro 500 users
LlamaIndex reports 93–96%+ accuracy for Extract v2.5 on long-list extraction
Factory adds model-level consumption analytics for coding agents
OpenRouter benchmarks seven model routers across six tests
NVIDIA reports 62.8% lower accuracy on 128K-token tasks than 4K-token tasks
OpenClaw runs AI-Infra-Guard on every ClawHub upload
Reports place Gemini 4 Argon at 77.9% on DeepSWE, 57.6% on Terminal-Bench, 77.5% on AutomationBench-AA, and 68% on CWE-Bench. The reports also cite lower cost or token use than rivals.
Google is rolling out Gemini 4 Argon through Fairwind to government users, vetted cyber defenders, and trusted testers. The model supports up to 1M output tokens and costs $2/M input and $10/M output.
GPT-6.1 Sol scored 86.3% on MathArena BrokenArXiv, 78.6% on Braintrust problem solving, and third in Code Arena WebDev. A user also reports that it outperformed its mini-SWE result in Codex.
Perplexity released pplx-embed-v2-context-9b-preview, which encodes chunks using whole-document context. Perplexity reports leading results on ConTEB and Turbopuffer's context benchmark.
Developers traced widespread iOS app crashes to a malformed server-side Firebase payload, which was rolled back. Cached payloads reportedly kept some apps failing afterward.
Anthropic reports that GLM-5.3 produced working browser exploits in 50 of 410 controlled attempts. In a separate binary-exploitation test, it achieved control-flow hijacks in 4% of trials.
GPT-6.1 Sol is available in the API, Codex and ChatGPT Work. It costs $2 per million input tokens and $10 per million output tokens; OpenAI claims near-Astra coding results.
Holo4 ships as a dense 27B model and a 35B-A3B mixture-of-experts model, with API access and downloadable weights. H Company reports a 61.7% score on OSWorld 2.0 and publishes replayable benchmark trajectories.
NVIDIA's Open Agent Safety Platform uses OpenShell to run agents in isolated environments. A technical report describes a BlueField-4 Sentry layer that can quarantine agents outside the sandbox.
Sonnet 5.5 ranks second on Artificial Analysis' Intelligence Index but uses about 193,000 output tokens per task at max effort. A comparison of AA Index results puts its cost at $7.60 per task.
Get the best stories deliveredto your inbox