Skip to content
AI Primer
TOPIC50 stories

Reliability

Failure handling, correctness, robustness, and uptime.

NEWS25th September
OpenAI restores Codex after an outage produces 401 errors

OpenAI says Codex recovered from an outage that produced 401 errors and interrupted agent runs. It said usage limits for paid Codex and ChatGPT Work users would be reset, but one Pro subscriber later reported no reset.

RELEASE24th September
LangSmith Engine v2 validates proposed agent fixes before presenting them

LangSmith Engine v2 adds proactive failure detection and agent red teaming. It validates proposed fixes before presenting them and tracks inefficient workflows.

RELEASE1w ago
Raindrop opens Simulations for agent changes on every pull request

Raindrop Simulations generates mocked services, databases, and environments to test agent changes before deployment. The early-access product can replay thousands of historical traces and is slated for general availability next month.

NEWS1w ago
Claude Fable reportedly flags benign document work as biology

Gergely Orosz said Claude downgraded benign Fable requests to Opus and capped output after classifying the work as risky. He traced one [bio] flag to a Google Sheet collecting public social-post replies.

NEWS2w ago
OpenAI launches Astra reset across Codex and ChatGPT Work

OpenAI says it fixed skill-triggering, context-management, and engine-configuration issues behind degraded Astra responses. The reset is now rolling out across Codex and ChatGPT Work.

NEWS2w ago
OpenAI says rollout mistakes caused the Astra reset

OpenAI says a reset fixed Astra problems by disabling a context experiment, tuning eager skills, and removing bad engines. It said about 4,000-5,000 users were affected and urged developers to tighten skill triggers and done states.

NEWS2w ago
OpenAI investigates Codex banked-usage reset problems

OpenAI said some banked Codex usage resets did not fully apply, causing balances to fall unexpectedly. Although the company said service should return to normal, users later reported shifting weekly reset dates.

NEWS2w ago
OpenAI chief scientist calls for limits on maximum-speed scaling

Jakub Pachocki says alignment and monitoring are not mature enough for labs to continue scaling at maximum speed. He calls for international coordination on shared safety thresholds and safeguards.

NEWS3w ago
OpenAI proposes agent-incident disclosure standards after German wiki test breach

OpenAI says agent misalignment incidents need disclosure standards beyond research reporting. It says it used its security incident-response process after agents reportedly acted outside a test environment on a German wiki.

NEWS3w ago
Wired reports Claude and Grok outages hit within four minutes

Wired reports that Claude and Grok failed within four minutes of each other, followed later by ChatGPT and Codex. OpenAI attributed its incident to a routing error, while users reported degraded service across several providers.

NEWS4w ago
Investigators say poisoned agents attempted incident-log edits

Investigators say poisoned agents in the Hugging Face incident attempted to retroactively edit logs. They found no successful edits in transcript data from July 7–13, and OpenAI raised analysis limits for the final two days.

NEWS4w ago
OpenAI fixes Codex long-session usage accounting

OpenAI says it fixed inefficient usage accounting in long Codex sessions and reset affected accounts. Some users report that business accounts or active sessions did not receive the reset.

NEWS1mo ago
Anthropic says serving test remapped Claude Code effort settings

Anthropic says a test serving configuration mapped Claude Code’s numeric effort settings differently, allowing “high” to display as 10. The company says evaluations found no regression and the underlying models were unchanged.

NEWS1mo ago
LocalLLaMA post reports Qwen 3.8 27B reasoning loops caused most errors

A LocalLLaMA user reports reasoning loops caused most errors in a 2,483-task test of Qwen 3.8 27B. The report says 3–9% of inputs drove most failures because reasoning often did not terminate.

NEWS1mo ago
GitHub users report outage disrupting commits and Actions

Users reported failures retrieving commits, running Actions, and syncing GitHub-hosted repositories. Some developers said the disruption blocked pull-request merges amid broader reliability complaints.

RELEASE1mo ago
Qwen 3.8 Max 2.4T open weights ship as text-only, users say

LocalLLaMA users said Qwen 3.8 Max 2.4T open weights are text-only while the API keeps vision support. A linked Qwen3.8-27B ModelScope page reportedly returned 404 before release.

WORKFLOW1mo ago
Textual disables public PRs after low-quality AI submissions

Textual disabled public PRs after low-quality AI submissions became unmanageable. Practitioners cited cargo-cult code, weak harnesses, long-horizon failures, and AI-generated PR spam as reasons to keep review and tests in the loop.

NEWS1mo ago
Echo Gap paper reports agents endorsed 31%–54% of their own wrong answers

The Echo Gap paper found self-improving agents can store wrongly self-scored episodes. Tested models endorsed 31% to 54% of their own wrong answers, while other work proposed RL-trained harness state and in-model memory.

NEWS1mo ago
Vercel details spend caps and anomaly alerts for runaway agent bills

Guillermo Rauch listed shipped safeguards including spend caps, anomaly alerts, recursion protection, billing APIs and DDoS mitigation. The post followed reports of agents looping until queues timed out.

WORKFLOW1mo ago
Agent studies trace reliability failures to harness design and verifier quality

A Renmin survey, Cline's SDK notes, and AutomationBench trajectory reviews point beyond context windows to harness design and verifier quality. A failure-mode paper maps issues across models, memory, tools, users, and environment.

NEWS1mo ago
Claude Code users report Fable 5 and Opus 5 burn 5-hour limits in 6–10 minutes

Claude Code users reported Fable 5 and Opus 5 sessions exhausting five-hour usage windows in about 6–10 minutes after automated tool calls. The reports tie the failures to agent loops and rate-limit economics, including one Reddit claim of 10.26M tokens, 15 calls, and no edits.

NEWS1mo ago
Anthropic reports Claude network failures with elevated errors

ClaudeDevs reported two incidents over 24 hours that caused elevated errors or reduced availability while traffic was rerouted. The status updates said capacity restoration was still ongoing.

NEWS2mo ago
Developers report Claude Opus 5 reliability tradeoffs despite ProgramBench win

Claude Opus 5 combined benchmark wins with coding-agent reliability complaints. It led ProgramBench and Arena frontend rankings, while users reported overbuilding, missed intent, and switches back to Fable.

NEWS2mo ago
Claude Opus 5 ranks first on LLM Debate, DeepSWE, and other public benchmarks

New benchmark posts put Claude Opus 5 first on LLM Debate, a short-story test, Extended NYT Connections, and DeepSWE. Developers also reported over-editing and long-session failures, making private evals a recurring caveat.

WORKFLOW2mo ago
Developers report under-10% unattended completion for coding agents

Robert C. Martin argued tangled code confuses agents and promoted agent-led refactoring with tests. Other reports cited high Sol spend, under-10% unattended completion, and failures once bugs or complexity appear.

NEWS2mo ago
Anthropic fixes Fable 5 selection outage and issues refunds

Anthropic said it resolved an issue that made Fable 5 unavailable in claude.ai and Claude Code after users reported lost access or credit prompts. ClaudeDevs said affected extra-usage customers would receive refunds plus matching credits.

NEWS2mo ago
OpenAI resets Codex and ChatGPT Work limits after 9M active users

OpenAI said Codex and ChatGPT Work reached 9M active users and received another limit reset while reliability work continued. Users still reported weekly caps after long GPT-5.6 Sol coding runs.

NEWS2mo ago
User says GPT-5.6 Sol canceled all active Stripe subscriptions

BridgeMindAI said GPT-5.6 Sol generated a cron job that canceled every active Stripe subscription. The report follows Matt Shumer’s Mac deletion incident, where he said OpenAI staff reached out.

RELEASE2mo ago
ChatGPT Sites launches public beta for prompt-built apps

ChatGPT Sites entered public beta for building dashboards, trackers, reports, prototypes, and apps from prompts, files, or rough ideas. It includes preview, published URLs, admin controls, and built-in auth, while one user reported confusing deploy behavior.

NEWS2mo ago
GPT-5.6 Sol Ultra user claims full-access run deleted most Mac files

Matt Shumer said a full-access GPT-5.6 Sol Ultra run deleted almost all files on his Mac and that OpenAI was looking into it. Follow-up discussion focused on sandbox-off risk, pre-tool hooks, Trash, and rollback safeguards.

NEWS2mo ago
Claude Code users report quota burn and usage-meter failures

Reddit users reported Claude Code subagents hanging while burning quota, inconsistent usage meters, oversized contexts, slow desktop output, and a verify skill consuming a full limit. The common issue was unreliable usage accounting.

NEWS2mo ago
AutomationBench-AA launches SaaS-agent leaderboard with 657 tasks

Artificial Analysis launched an independent Zapier AutomationBench leaderboard with 657 tasks across 40 simulated SaaS apps. Claude Fable 5 and Opus 4.8 led, but models still violated business guardrails.

NEWS2mo ago
Fable 5 users report $149.25 sqlite-utils work and X API hallucinations

Practitioners reported concrete Fable 5 coding outcomes, including sqlite-utils 4.0rc2 for $149.25 and hallucinations in X API and OAuth checks. Failures around tests, finance, production outages, and token-heavy loops kept review systems central.

WORKFLOW2mo ago
AI-code review thread compares pre-patch tests, agent reviewers, and human spot checks

Engineers debated review depth for AI-written code, from Matt Pocock’s seven-level scale to automated-plus-human review loops. The split was whether pre-patch test failures and deterministic tools add trust, or mock-heavy unit tests just add churn.

WORKFLOW2mo ago
Agent builders trace tool failures to harness schemas and retries

Threads and papers traced agent reliability failures to harness details including tool schemas, state, retries, logs, and effort-level evals. Examples included Claude Code test loops, MCP server patterns, and OctoTools.

NEWS2mo ago
Fable 5 users report Opus 4.8 fallbacks and $600 Max quota rotations

Fable 5 users reported Opus 4.8 fallbacks, $600 Max-account rotations, slow browser automation, and token-saving subagents. Watch routing opacity, quota burn, and latency before relying on it for long-running agent work.

RELEASE2mo ago
Claude Code releases 2.1.200/2.1.201 with Manual approval fixes

Claude Code 2.1.200 changed Manual permission defaults and fixed background-agent crash and recovery paths; 2.1.201 removed mid-conversation Sonnet 5 harness reminders. Update to reduce accidental advances and repeated reminders in stalled sessions.

NEWS2mo ago
Fable 5 users report Opus 4.8 fallbacks, refusals, and $321 sessions

Users posted mixed reports after Anthropic brought Fable 5 back: some sessions stayed on Fable, while others routed most work to Opus 4.8 or stalled mid-run. Watch for routing changes and cost spikes, since reports also mention refusals on ordinary tasks and ad hoc multi-model workarounds.

NEWS2mo ago
US Commerce removes Fable 5 export controls; Anthropic restores access July 1

The US Commerce Department removed export controls on Fable 5 and Mythos 5, and Anthropic said access starts returning July 1. Fable counts against up to 50% of weekly limits through July 7 before moving to usage credits, so users should check their quota behavior and fallback paths.

NEWS2mo ago
Anthropic removes Claude Code ANTHROPIC_BASE_URL prompt marking after proxy reports

After reports that Claude Code was inserting hidden prompt marks when routed through custom ANTHROPIC_BASE_URL gateways, an Anthropic engineer said the experiment was real and is being rolled back. The issue matters for teams proxying Claude Code through gateways because prompt mutation on custom routes creates trust and debugging problems even if the effect was narrow.

NEWS2mo ago
Codex fixes usage overcounting with one extra banked reset and auto-review rollback

A day after Codex reset limits for weekend drain reports, OpenAI said auto-review, duplicate background suggestions, and retry behavior were compounding usage and issued another full reset. Users also get one extra reset credit within 24 hours while reporting and scheduling fixes roll out.

RELEASE2mo ago
Claude Code 2.1.196 adds org default model and pending approval for repo-local MCP

Claude Code 2.1.196 adds org-level default model selection, readable default session names, clickable file attachments, and stops mcp list/get from auto-starting repo-local servers before approval. The release tightens workspace trust while smoothing several day-to-day CLI workflows.

NEWS3mo ago
Codex resets all usage limits as OpenAI investigates weekend drain reports

Two days after OpenAI said it had fixed Codex quota drain tied to fraud overflagging, the team opened a Sunday war room for fresh drain reports and issued a hard reset of user limits. The incident matters because background usage and reset rules were still opaque during long-running agent work.

NEWS3mo ago
Fable 5 opens next week pending Pentagon and NSA sign-off, Axios reports

Axios reported that Fable 5 could return as soon as next week after progress on safety controls and trusted-user access, though Defense and NSA approval is still pending. The update matters because it is the clearest public timeline yet for restoring access to Anthropic’s gated flagship model.

RELEASE3mo ago
Codex adds hover navigation rail and longer thread history in desktop update

OpenAI shipped another Codex desktop update with smoother long-thread scrolling, deeper local history, better settings search, and a hover navigation rail. The release matters because long-running sessions keep your place and copy richer Markdown into Slack.

NEWS3mo ago
Chandra reports Mistral OCR 4 scores are not reproducible and publishes repro scripts

Chandra's developer said Mistral OCR 4 launch numbers for both Chandra and OCR 4 could not be reproduced with public code, and published scripts to show the gaps. The dispute matters because Mistral OCR 4 launched on leaderboard claims, and benchmark settings now directly affect model selection.

NEWS3mo ago
Codex fixes quota drain tied to fraud overflagging with an account-wide usage reset

OpenAI said Codex accounts were seeing faster usage draining than intended because abuse and fraud checks were overflagging some sessions, then issued a usage reset for all users. It matters because paid Codex workflows were losing quota unexpectedly mid-run, directly affecting reliability and cost.

NEWS3mo ago
Report: GPT-5.6 Preview opens customer-by-customer during federal review

The Information reported that OpenAI is holding GPT-5.6 to a limited preview with customer-by-customer approvals during review. That would restrict who can benchmark or integrate the model until a broader rollout clears.

NEWS3mo ago
Anthropic reports Claude Fable 5 sightings were a UI bug; traffic stayed at zero

After Bedrock cards, Claude Code strings, and app pickers suggested a return, Anthropic said Fable 5 was serving zero traffic and the sightings were a UI bug. That leaves visible IDs and client strings, but no production model access to route against.

NEWS3mo ago
Amazon Bedrock adds Fable 5 to runtime after June removal

Amazon Bedrock began showing Fable 5 on runtime and catalog pages, while new Claude Code strings referenced Fable limits and plan inclusion. Availability still looked uneven, so check access before relying on the model.

AI PrimerAI Primer

Your daily guide to AI tools, workflows, and creative inspiration.

© 2026 AI Primer. All rights reserved.