Claude Opus 5, vLLM 0.26 / DeepSeek V4, Ollama 0.32.4
Sonntag, 26. Juli 2026 - AI News · (letzte 24h)
Anthropic shipped Claude Opus 5, claiming Fable-level performance at half the price — a direct upgrade path for Claude Code users.
Must read
- Claude Opus 5: Fable-level performance at Opus price — New default model for your Claude Code and overnight-agent-factory setup; halved price changes routing economics through your LiteLLM gateway.
- Claude Opus 5: The System Card — System card details matter before you swap Opus 5 into agentic workflows touching customer identity/fraud data.
- vLLM v0.26.0 — DeepSeek V4, Inkling, MTP speculative decoding — 411 commits, DeepSeek V4 kernels and MTP=1 speculative decoding — relevant if you’re serving open models behind LiteLLM.
- Ollama v0.32.4 — Laguna on Apple GPUs via MLX — MLX engine gains Laguna support and 4-9% faster Qwen3 MoE decoding on M5 Max; direct upgrade for your local coding stack.
Tools & Frameworks
Ruff v0.16.0 enables 413 default rules
Astral’s Ruff jumped from 59 to 413 default-enabled rules on 23 July, silently breaking CI on unpinned dev deps.
Why this matters: Pin ruff in your Python repos before GitHub Actions goes red overnight.
LangChain: Own Your Intelligence
LangChain argues durable AI advantage comes from owning agent systems, governance, context stores and feedback loops rather than model choice.
Why this matters: Aligns with your connected-data-model / embedded-governance thesis; watch, don’t act.
Open Models & Local
Ollama v0.32.4
MLX engine adds Laguna support on Apple GPUs, quantizes speculative-decoding draft heads, and fixes Qwen3 MoE decoding across mixed-quant experts.
Why this matters: Direct win for Qwen3-Coder on Apple Silicon in your local-plus-cloud routing.
vLLM v0.26.0
Adds DeepSeek-V4 routing kernel (2.94% E2E TPOT), Inkling family, Hopper FA4 attention, and MTP=1 speculative decoding with LoRA and NVFP4.
Why this matters: If you self-host open coders alongside frontier via LiteLLM, this is the version to test.
Industry & Trends
Azeem Azhar: the curious case of AI distillation
Azhar surveys how frontier labs are distilling large models into cheaper, capable ones — the pattern behind Opus 5’s price cut.
Why this matters: Context for why your token bill keeps dropping while capability climbs.
Build on the stack you have — Anthropic, Atlassian, Scale VP converge
Three SaaStr AI 2026 sessions land on the same playbook: layer AI onto existing product surfaces rather than rebuild from scratch.
Why this matters: Useful framing if you’re pitching board on identity/fraud AI roadmap rather than a rewrite.
Sources unavailable today: r/ChatGPTCoding top, r/ClaudeAI top, r/LocalLLaMA top, r/MachineLearning top
Auto-curated daily by Claude Opus 4.7 from Don’t Worry About the Vase (Zvi), Exponential View (Azeem Azhar), GitHub: ollama/ollama, GitHub: vllm-project/vllm, LangChain blog, Latent Space, Lenny’s Newsletter, Not Boring (Packy McCormick), SaaStr (Jason Lemkin), Simon Willison. Source list and editorial profile maintained by Daniel.