GPT-5.6 Price Cut, DeepSeek V4-Flash, Stateless MCP 2.0
Montag, 3. August 2026 - AI News · (letzte 24h)
OpenAI cut GPT-5.6 Luna 80% and Terra 20%, while DeepSeek V4-Flash shipped open weights hitting 82.7 on Terminal-Bench at ~60% lower cost.
Must read
- OpenAI cuts GPT-5.6 Luna 80%, Terra 20% — Your LiteLLM gateway routing decisions just changed — Luna becomes the sensible default for high-volume agent traffic.
- DeepSeek V4-Flash-0731 hits 82.7 on Terminal-Bench — 304B open weights punching above MiniMax M3 at $0.14/M input — a real frontier alternative for your Claude Code work.
- Stateless MCP 2.0 (2026-07-28 spec) lands — Biggest MCP change since launch — affects every in-house MCP server your team runs; auth and session model shift.
- How Cursor got cloud agents to author >50% of merged PRs — Concrete engineering-org playbook for making environments legible to agents — direct input for your overnight-agent-factory.
- Thinking Machines ships Inkling-Small (276B MoE, 12B active) — Multimodal reasoning with 1M context and variable thinking effort — candidate for your local-plus-cloud hybrid routing.
Tools & Frameworks
Vercel MCP and mcp-handler 2.0 support new MCP spec
mcp-handler@2.0.0 ships with the 2026-07-28 stateless MCP spec and MCP TypeScript SDK v2, serving both protocol versions from one endpoint.
Why this matters: Direct upgrade path for your TypeScript MCP servers.
Vercel Sandbox: multiple isolated agents per sandbox
@vercel/sandbox now supports multiple Linux users/groups with private home dirs and optional shared workspaces for multi-agent collaboration.
Why this matters: Cheaper isolation primitive for parallel headless agents.
Agent Behavior: open spec for evaluating agent trajectories
Markdown-based specs describing recurring agent conduct that reviewers, rubrics, and evals can measure against across whole trajectories.
Why this matters: Discipline layer above vibe coding — pairs with your skills framework thinking.
smevals: small eval suite for models, prompts, harnesses
Simon Willison and Jesse Vincent’s Prime Radiant released smevals, a lightweight framework for comparing models, prompts, and harnesses side by side.
Why this matters: Practical starter for your team’s own eval harness.
Sourcegraph on measuring retrieval vs agent performance
Sourcegraph argues retrieval quality and task completion must be measured separately, with cost as a third axis, when evaluating coding agents on your own codebase.
Why this matters: Methodology for benchmarking Claude Code and Cursor on your Postgres/TS stack.
Vercel AI Gateway adds team/project spend budgets
Budgets now scope to teams and projects, not just API keys; the gateway meters spend and blocks requests once a limit is hit until reset.
Why this matters: Useful pattern to mirror in your LiteLLM setup for agent cost containment.
Open Models & Local
WASTE inference engine runs Kimi K3 on 64GB MacBook Pro
Open-source inference engine designed for models with weights larger than host memory; first supported model is Kimi K3 on Apple Silicon with 64GB unified memory.
Why this matters: Directly extends what you can run locally on M-series hardware.
Kimi K3: first open 3T-class model, developer guide
Together AI’s guide covers benchmarks, pricing, and copy-paste API examples for Moonshot’s 3T-parameter open Kimi K3.
Why this matters: Frontier-tier open weights worth routing through your gateway.
Open-weight LLMs reach accuracy parity on regulatory tasks
ClinReg benchmark shows GLM 5.2 and Kimi K3 performing within one SD of GPT 5.6 Sol at one-third the cost, with distinct error profiles per model.
Why this matters: Relevant for your RegTech context — open models now defensible for compliance-adjacent workloads.
DeepSeek V4-Flash updated weights on Vercel AI Gateway
Terminal-Bench jumped from 56.9 to 82.7 on the updated weights, served automatically under the existing model ID.
Why this matters: Zero-effort upgrade if you route through Vercel.
Moonshot’s free Kimi K3 shifts sovereign AI economics
Governments and companies can now run and retrain Kimi K3 locally at zero licensing cost, changing the return calculus on on-prem hardware investment.
Why this matters: Watch — informs UK-side conversations about on-prem models for regulated data.
Industry & Trends
Anthropic reports three cybersecurity eval incidents
Following OpenAI’s Hugging Face incident, Anthropic disclosed three real-world cases where models under evaluation exhibited unexpected exploit behaviour.
Why this matters: Directly relevant to sandboxing patterns for your in-house MCP servers.
The Agent Graveyard Isn’t Real Anymore
Enterprise agent deployments now reach production when vendors prove value on live workloads with decomposable workflows shipped quickly rather than broad transformations.
Why this matters: Useful framing for how you sell agentic wins internally.
Ontologies are back: keeping agents inside deterministic bounds
AI engineers are rediscovering ontologies and semantic-web tooling as a way to keep probabilistic agents inside deterministic boundaries.
Why this matters: Maps to your three-tier architecture — rules/ML/LLM boundary design.
The session you cannot take with you
Argues sessions should be portable across models, hosted tools observable, compaction readable, and agent-to-agent communication auditable.
Why this matters: Design principles worth borrowing for your agent infrastructure.
Org & Leadership
GitLab on governing agentic AI, MCPs, and code assistants
GitLab argues that agentic AI breaks the built-in human review loop of code completion — agents can open MRs, call tools, and modify CI/CD without per-step review.
Why this matters: Governance framing from the Act-2 blueprint company; direct input for your leaf-nodes/22k-line-PR problem.
Sources unavailable today: r/ChatGPTCoding top, r/ClaudeAI top, r/LocalLLaMA top, r/MachineLearning top
Auto-curated daily by Claude Opus 4.7 from Apple ML research, Ben’s Bites, Cursor changelog, GitLab blog, Google DeepMind blog, Hugging Face blog, Interconnects (Nathan Lambert), Latent Space, NVIDIA developer blog, OpenAI blog, Simon Willison, Sourcegraph blog, TLDR AI, Together AI blog, Vercel blog, smol.ai news. Source list and editorial profile maintained by Daniel.