Skip to content

← AI Tracker

AI Briefing

Claude Opus 5, Claude Code 2.1.219, Kimi K3 vs Fable 5

Samstag, 25. Juli 2026 - AI News · (letzte 24h)

Anthropic shipped Claude Opus 5 — frontier-adjacent coding at half of Fable 5’s price, 1M context, and the least prompt-injectable model yet.

Must read

Tools & Frameworks

Claude Opus 5 on Vercel AI Gateway

Opus 5 live on AI Gateway with low/medium effort settings tuned for token/latency savings on multi-file refactors.

Why this matters: Drop-in for anything you route via Vercel rather than LiteLLM.

Cline 4.0.11 adds Opus 5 and Kimi K3

Opus 5 (incl. 1M variants) and Kimi K3 with native tool calling added across Anthropic, Bedrock, Vertex, OpenRouter providers.

Why this matters: If your team uses Cline alongside Claude Code, day-one parity.

langchain-anthropic 1.5.2 with Opus 5

Opus 5 support landed in langchain-anthropic 1.5.2 same day as the model launch.

Why this matters: Relevant if any Python agents still sit on LangChain.

Codex + ChatGPT desktop get full-duplex voice

OpenAI wired GPT-Live’s full-duplex audio into Codex and the ChatGPT desktop app for hands-free agentic coding sessions.

Why this matters: Watch, don’t act — but the direction of travel for driving Claude Code by voice is now clear.

Claude voice mode adds Opus/Sonnet + app integrations

Claude voice now supports Opus and Sonnet with Gmail, Slack, Notion, and Calendar integrations in beta.

Why this matters: Consumer-side, but signals where Anthropic will land voice on the SDK next.

OpenWorker — local desktop AI coworker

Andrew Ng-adjacent OSS repo running a local coworker that acts across files and apps on your desktop.

Why this matters: Worth a look for your local-plus-cloud hybrid setup.

Open Models & Local

Kimi K3’s thinking-trace strategy

Kimi K3 spends 12x more reasoning tokens than Opus 4.8 and iterates like an internal agent inside its chain of thought.

Why this matters: Explains the pass@4 wins over Fable 5 — different cost/quality tradeoff to model in routing.

DeepSeek V4 trained on Huawei Ascend at 34.22% MFU

Huawei consortium published a technical report on full-parameter post-training of DeepSeek V4 on Ascend chips at 34.22% MFU.

Why this matters: Confirms a non-NVIDIA path for the models you may run locally next year.

SGLang 0.5.16 with DSpark speculative decoding

Confidence-driven speculative decoding hits 383.7 tok/s at accept length ~5 on DeepSeek-V4-Pro, TP8 on B300.

Why this matters: Relevant if your inference stack ever leaves the frontier APIs.

Fugu-Ultra v1.1

Multi-model orchestrator improves coding, agentic tasks and reasoning at the same price as v1.0, positioned as vendor-independent frontier access.

Why this matters: Another gateway-layer competitor to LiteLLM worth benchmarking.

Runway Media Router

Runway shipped an auto-routing API selecting image/video/audio models by quality, speed or cost, including third-party models.

Why this matters: Adjacent to your world; useful pattern reference for model-router UX.

Microsoft MAI-Image-2.5-Pro and MAI-Voice-2-Flash

Microsoft pushed its own high-fidelity image and low-cost voice models into public preview.

Why this matters: Skip unless media generation lands on your roadmap.

Cognition acquires The Interaction Company (Poke)

Devin-maker Cognition absorbed Poke, a personal agent that texts natively via Apple Messages.

Why this matters: Signals Cognition pushing beyond coding into consumer agent surfaces.

Sierra acquires TakeOff

Bret Taylor’s Sierra bought TakeOff to add a long-horizon multistep agent platform for enterprise deployments.

Why this matters: Enterprise agent M&A is heating up — competitive context.

NVIDIA ModelExpress for model artifact distribution

New NVIDIA tooling to move hundred-GB-to-TB model checkpoints faster across clusters.

Why this matters: Only relevant if you self-host large weights; useful pattern otherwise.


Sources unavailable today: r/ChatGPTCoding top, r/ClaudeAI top, r/LocalLLaMA top, r/MachineLearning top

Auto-curated daily by Claude Opus 4.7 from Apple ML research, Don’t Worry About the Vase (Zvi), GitHub: anthropics/claude-code, GitHub: cline/cline, GitHub: crewAIInc/crewAI, GitHub: ggml-org/llama.cpp, GitHub: langchain-ai/langchain, GitHub: sgl-project/sglang, Lenny’s Newsletter, NVIDIA developer blog, Not Boring (Packy McCormick), SaaStr (Jason Lemkin), Simon Willison, TLDR AI, The Algorithmic Bridge (Alberto Romero), Together AI blog, Vercel blog. Source list and editorial profile maintained by Daniel.