Claude 5 Family, LLM 0.32, Cursor Workspace Plugins
Wednesday, 5 August 2026 - AI News · (last 24h)
Anthropic’s Claude 5 family (Opus/Sonnet/Fable) lands via llm-anthropic 0.26, alongside a major LLM 0.32 release and Cursor Google Workspace plugins.
Must read
- LLM 0.32: reasoning traces, Responses API, resumable tool loops — Biggest LLM release since launch — structured messages, pausable tool loops and content-addressed logs matter for your LiteLLM gateway and eval tooling.
- llm-anthropic 0.26 ships claude-opus-5, sonnet-5, fable-5 — New Claude 5 tier lands via plugin — retest your Claude Code and agent-orchestration flows against Opus 5 before defaulting.
- Cursor adds Google Workspace plugins — Cursor agents can now read/write Docs, Sheets, Gmail — expands your one-person-team leverage beyond the repo.
- Claude Code v2.1.222 hardens worktree isolation — Fixes worktree-isolated sessions being able to run destructive git against main checkout — directly affects your overnight-agent-factory setup.
- Skill packs on skills.sh — Bundle agent-skills as a shareable pack via one URL — the discipline layer above vibe coding you’ve been writing about.
Tools & Frameworks
Microsoft Orchard: Kubernetes-native agentic modeling framework
Open-source thin environment service exposing generic primitives for trajectory distillation, on-policy RL rollouts and evals across harnesses.
Why this matters: Substrate for standardising agent training and eval across your team.
How Kiro built its one-agent harness
Kiro splits an IDE, CLI and Web agentic IDE across a single server-side harness with a fixed protocol boundary between agent and client.
Why this matters: Useful reference architecture if you’re consolidating your in-house MCP servers behind one harness.
DeepSeek V4 Flash 90% off via Novita on Vercel AI Gateway
Vercel Pro customers get 90% off DeepSeek V4 Flash through 11 August by routing via Novita on the AI Gateway.
Why this matters: Cheap fallback tier for your LiteLLM routing on non-critical agent paths.
GPT-5.6 Sol burns 2.25x the tokens of GPT-5.5 in Codex
GPT-5.6 Sol xhigh uses more than 2x tokens per Codex session vs 5.5, plus adds a cache-write charge 5.5 didn’t have.
Why this matters: Recalibrate your OpenAI vs Anthropic routing economics before defaulting Codex flows to 5.6.
MirrorCode: long-horizon reimplementation benchmark
25 target programs across Unix utilities, interpreters, cryptography and compression — models must reproduce end-to-end output without source access.
Why this matters: Better signal than SWE-bench for evaluating your coding agents on real long-horizon work.
CX agents in production at Lyft, Vodafone, LATAM
LangChain writeup of three production CX agent deployments with lessons on tooling, guardrails and human handoff.
Why this matters: Directly relevant if identity/fraud teams need agent patterns for regulated support flows.
Open Models & Local
Liquid AI ships LFM2.5-2.6B for local agents
2.6B parameter model tuned for on-device agents, deployable across mobile and edge alongside laptop targets.
Why this matters: Small enough for local fallback on Apple Silicon; watch as a router option.
MiniMax-H3 ported to MLX for Apple Silicon
Omni-modal model (text/image/audio/video, generates 15s clips with audio) now runs on Apple Silicon via MLX.
Why this matters: Watch — not for coding, but the MLX port matters for your local stack experiments.
Fast Gemma inference recipe on Gemma 4 E4B
VIDRAFT documents every optimisation behind its SOTA Fast Gemma tokens/sec run on a single A10G.
Why this matters: Concrete tuning recipe if you’re serving Gemma 4 in-house.
DeepSeek V4-Flash 105x cheaper than Claude Fable 5
Research firm finds DeepSeek V4-Flash the cheapest well-known model to run — 105x below Claude Fable 5 on cost.
Why this matters: Sharpens the cost-vs-capability call for your hybrid routing setup.
Industry & Trends
How OpenAI built GPT-Live
Full-duplex voice architecture combining stateful inference, async delegation, dynamic context management and low-latency media transport.
Why this matters: Reference architecture if voice agents ever intersect your identity/fraud flows.
OpenAI Astra solves ten open maths problems
Unreleased OpenAI model solved ten well-defined open mathematics problems with verifiable solutions.
Why this matters: Signal on verifiable-domain capability — read once, don’t act.
HBM4 forecast to double as hyperscalers hit capacity
HBM3e memory up 20%, HBM4 forecast to double; hyperscalers declared capacity-constrained on Q2 2026 calls.
Why this matters: Frontier API prices likely to firm — plan your local-plus-cloud routing accordingly.
Org & Leadership
Amodei worries Anthropic hires join for money, not mission
Why this matters: Retention signal at the top lab — compute, influence and autonomy are the levers, not pay.
Sources unavailable today: The Gradient, r/ChatGPTCoding top, r/ClaudeAI top, r/LocalLLaMA top, r/MachineLearning top
Auto-curated daily by Claude Opus 4.7 from Ben’s Bites, GitHub: anthropics/claude-code, GitHub: ggml-org/llama.cpp, GitHub: simonw/llm, Hugging Face blog, LangChain blog, Latent Space, Lenny’s Newsletter, NVIDIA developer blog, OpenAI blog, SaaStr (Jason Lemkin), Simon Willison, TLDR AI, Tomasz Tunguz, Understanding AI (Timothy B. Lee), Vercel blog, smol.ai news. Source list and editorial profile maintained by Daniel.