All Stories
321 storiesSort:
Time:
20th July
19th July

Alibaba opens Qwen 3.8 Max Preview testing across Cloud, Qwen Chat, Qoder and web
Releaseπ§ Qwen19th July

Moonshot pauses new Kimi K3 subscriptions after GPU capacity crunch
π§ Kimi19th July

ChatGPT Work desktop adds cloud vs local run controls
ReleaseβοΈCodex19th July

Engineers replace broad agent loops with scoped workflows and SWE-bench harnesses
Workflowπ³Agent design patterns19th July

OpenBMB releases MiniCPM-Robot models and PhyAI runtime with 33-36 Hz throughput claim
Releaseπ§ Robotics19th July

Kimi K3 ranks No. 1 on Arena Frontend Code leaderboard
π§ Kimi19th July
18th July

Users claim DeepSeek V4 routes hard API prompts through Claude Fable 5
π§ Fable18th July

Kimi K3 benchmarks last at 53/67 in AlphaSignal repair harness
π§ Kimi18th July

Study reports Claude Code and Codex memory can store prompt-injection rules
π‘οΈClaude Code18th July

ChatGPT Work supports Plus, Pro, Business, and Enterprise on web and mobile
ReleaseπAgent product updates18th July

Code Arena users report Kaleb shows Qwen-like token quirks
π§ Qwen18th July

Developers report GPT-5.6 Sol and Fable overengineer small coding tasks
Workflowβ¨οΈCoding Agents18th July

Posts say Anthropic makes Fable 5 permanent on Claude Max and Team Premium at 50% limits
π§ Fable18th July

Slate, LangChain, and AI SDK add graph control for agent runtimes
WorkflowβοΈAgent Framework18th July
17th July

Anthropic adds Fable 5 to Max and Team Premium at 50% limits
π§ Claude17th July

Red-teamers claim Kimi K3 jailbreaks produced cyber and bio outputs
π§ Kimi17th July

Kimi K3 ranks #5 on Artificial Analysis as engineers dispute coding cost
π§ Kimi17th July

Posts claim GPT-5.6 Sol beats Mythos 5 on UK AISI and CyberGym tasks
π§ GPT17th July

Anthropic fixes Fable 5 selection outage and issues refunds
π§ Fable17th July

Anthropic releases Claude Managed Agents patterns for long-running agents
Releaseπ§ Claude17th July

Developers use Markdown files as long-term memory for agents
WorkflowπPersistent Storage17th July
16th July

OpenAI traces Codex file deletions to $HOME handling bug
β¨οΈCodex16th July

Baseten reports LLM fact writes can retain lift but vanish from answers
π§ Context Engineering16th July

Moonshot launches Kimi K3 with 2.8T parameters and 1M context
Releaseπ§ Kimi16th July

LocalLLaMA users report near-6x Qwen 3.6 27B speedups with MTP
Workflowπ§ llama.cpp16th July

Inkling adds early llama.cpp serving via 1-bit GGUF
π§ Open Models16th July

AI21 reports 80.8% on SWE-Bench Pro with model-team agent
WorkflowβοΈMulti-Agent Systems16th July

OpenAI updates ChatGPT desktop with history sync after Work feedback
Releaseβ¨οΈCodex16th July
15th July

Thinking Machines releases Inkling: 975B open-weight multimodal MoE
Releaseπ§ Open Models15th July

Users report Kivine on LMArena may be a Kimi K3 preview
π§ Kimi15th July

OpenAI introduces GPT-Red for prompt-injection red teaming
π‘οΈGPT15th July

Cursor users report agents switching to costly Claude Opus, Sonnet, Fable, or API calls
β¨οΈCursor15th July

OpenAI resets Codex and ChatGPT Work limits after 9M active users
π³Codex15th July

OpenAI and Work Louder release $230 Codex Micro control deck
Releaseβ¨οΈCodex15th July
14th July

Tencent releases 1-bit and 4-bit GGUF weights for 295B Hy3 single-GPU runs
Releaseπ§ llama.cpp14th July

Goodfire opens Silico private beta for automated interpretability and RL experiments
Releaseπ§ GLM14th July

OpenAI reports 8M Codex and ChatGPT Work users after 2.5x weekly usage jump
π§ Codex14th July

Developers report stale agents.md files derailing coding-agent runs
Workflowπ‘οΈCoding Agents14th July

Meta says its model scored 30/30 on Asian Physics Olympiad theory exam
π§ Muse14th July

Perplexity releases WANDR benchmark with 500 deep-research tasks
ReleaseπBenchmarks14th July

Posts claim Codex Desktop system prompt leaked with GPT-5.6 Sol tool list
π‘οΈCodex14th July