Claude Opus 5, Claude Code 2.1.219, Kimi K3 vs Fable 5
Samstag, 25. Juli 2026 - AI News · (letzte 24h)
Anthropic shipped Claude Opus 5 — frontier-adjacent coding at half of Fable 5’s price, 1M context, and the least prompt-injectable model yet.
Must read
- Introducing Claude Opus 5 — New default Opus: 1M context, $10/$50 per Mtok, pitched at long-horizon agentic coding — direct hit on your Claude Code workflow.
- Claude Code 2.1.219–2.1.220 — Opus 5 is now the Claude Code default, plus a strict network allowlist and DirectoryAdded hook — sandboxing and multi-repo dispatch just got better.
- Opus 5 is Anthropic’s least prompt-injectable model — Material for a RegTech CTO shipping agents against untrusted inputs; worth reading the system card section directly.
- Kimi K3 vs Claude Fable 5 on DeepSWE — Kimi K3 delivers 2.8x solves per dollar at pass@4 — a real routing candidate for your LiteLLM gateway on cost-sensitive agent runs.
- SaaStr: 30 agents down to 20 — consolidation in practice — Honest post-mortem on agent sprawl at an eight-figure business; useful counterweight to your overnight-agent-factory instincts.
Tools & Frameworks
Claude Opus 5 on Vercel AI Gateway
Opus 5 live on AI Gateway with low/medium effort settings tuned for token/latency savings on multi-file refactors.
Why this matters: Drop-in for anything you route via Vercel rather than LiteLLM.
Cline 4.0.11 adds Opus 5 and Kimi K3
Opus 5 (incl. 1M variants) and Kimi K3 with native tool calling added across Anthropic, Bedrock, Vertex, OpenRouter providers.
Why this matters: If your team uses Cline alongside Claude Code, day-one parity.
langchain-anthropic 1.5.2 with Opus 5
Opus 5 support landed in langchain-anthropic 1.5.2 same day as the model launch.
Why this matters: Relevant if any Python agents still sit on LangChain.
Codex + ChatGPT desktop get full-duplex voice
OpenAI wired GPT-Live’s full-duplex audio into Codex and the ChatGPT desktop app for hands-free agentic coding sessions.
Why this matters: Watch, don’t act — but the direction of travel for driving Claude Code by voice is now clear.
Claude voice mode adds Opus/Sonnet + app integrations
Claude voice now supports Opus and Sonnet with Gmail, Slack, Notion, and Calendar integrations in beta.
Why this matters: Consumer-side, but signals where Anthropic will land voice on the SDK next.
OpenWorker — local desktop AI coworker
Andrew Ng-adjacent OSS repo running a local coworker that acts across files and apps on your desktop.
Why this matters: Worth a look for your local-plus-cloud hybrid setup.
Open Models & Local
Kimi K3’s thinking-trace strategy
Kimi K3 spends 12x more reasoning tokens than Opus 4.8 and iterates like an internal agent inside its chain of thought.
Why this matters: Explains the pass@4 wins over Fable 5 — different cost/quality tradeoff to model in routing.
DeepSeek V4 trained on Huawei Ascend at 34.22% MFU
Huawei consortium published a technical report on full-parameter post-training of DeepSeek V4 on Ascend chips at 34.22% MFU.
Why this matters: Confirms a non-NVIDIA path for the models you may run locally next year.
SGLang 0.5.16 with DSpark speculative decoding
Confidence-driven speculative decoding hits 383.7 tok/s at accept length ~5 on DeepSeek-V4-Pro, TP8 on B300.
Why this matters: Relevant if your inference stack ever leaves the frontier APIs.
Industry & Trends
Fugu-Ultra v1.1
Multi-model orchestrator improves coding, agentic tasks and reasoning at the same price as v1.0, positioned as vendor-independent frontier access.
Why this matters: Another gateway-layer competitor to LiteLLM worth benchmarking.
Runway Media Router
Runway shipped an auto-routing API selecting image/video/audio models by quality, speed or cost, including third-party models.
Why this matters: Adjacent to your world; useful pattern reference for model-router UX.
Microsoft MAI-Image-2.5-Pro and MAI-Voice-2-Flash
Microsoft pushed its own high-fidelity image and low-cost voice models into public preview.
Why this matters: Skip unless media generation lands on your roadmap.
Cognition acquires The Interaction Company (Poke)
Devin-maker Cognition absorbed Poke, a personal agent that texts natively via Apple Messages.
Why this matters: Signals Cognition pushing beyond coding into consumer agent surfaces.
Sierra acquires TakeOff
Bret Taylor’s Sierra bought TakeOff to add a long-horizon multistep agent platform for enterprise deployments.
Why this matters: Enterprise agent M&A is heating up — competitive context.
NVIDIA ModelExpress for model artifact distribution
New NVIDIA tooling to move hundred-GB-to-TB model checkpoints faster across clusters.
Why this matters: Only relevant if you self-host large weights; useful pattern otherwise.
Sources unavailable today: r/ChatGPTCoding top, r/ClaudeAI top, r/LocalLLaMA top, r/MachineLearning top
Auto-curated daily by Claude Opus 4.7 from Apple ML research, Don’t Worry About the Vase (Zvi), GitHub: anthropics/claude-code, GitHub: cline/cline, GitHub: crewAIInc/crewAI, GitHub: ggml-org/llama.cpp, GitHub: langchain-ai/langchain, GitHub: sgl-project/sglang, Lenny’s Newsletter, NVIDIA developer blog, Not Boring (Packy McCormick), SaaStr (Jason Lemkin), Simon Willison, TLDR AI, The Algorithmic Bridge (Alberto Romero), Together AI blog, Vercel blog. Source list and editorial profile maintained by Daniel.