Claude Sonnet 5, Cursor iOS, Gemma 4 MTP
Dienstag, 30. Juni 2026 - AI News · (letzte 24h)
Anthropic ships Claude Sonnet 5 as default with 1M context and Opus-tier agentic capability at $2/$10 per Mtok promotional pricing.
Must read
- Claude Sonnet 5 default in Claude Code v2.1.197 — Sonnet 5 with 1M context is now the Claude Code default at $2/$10 per Mtok — directly changes your overnight-agent-factory economics.
- Cursor for iOS (public beta) — Dispatch and monitor cloud/local Cursor agents from your phone — the remote-control layer for parallel headless workflows you write about.
- Ollama 0.31.1: 90% faster Gemma 4 on Apple Silicon via MTP — Multi-token prediction nearly doubles Gemma 4 throughput on M-series — meaningful for your local-plus-cloud routing setup.
- Devin Fusion: multi-model harness cuts cost 35% — Dual-agent routing between frontier and cheap models mirrors your LiteLLM gateway thinking — a concrete pattern to steal.
- Impressions from visiting OpenAI, Anthropic & Cursor — Orosz on cloud-agent trends and coding harnesses spreading beyond craft — relevant priors before AI Engineer World’s Fair.
Tools & Frameworks
Cursor 3.10 changelog
Cursor 3.10 ships team marketplace updates on 30 June.
Why this matters: Your team runs on Cursor — check what changed.
Claude Sonnet 5 on Vercel AI Gateway
Sonnet 5 available via AI Gateway with new tokenizer; reaches Opus-tier outcomes at Sonnet pricing on coding and agentic work.
Why this matters: You route through LiteLLM — parallel option and useful pricing benchmark.
GitHub Copilot native agent in JetBrains IDEs
Copilot integrates as a first-class agent in the JetBrains agent picker, replacing the ACP Registry route.
Why this matters: Signal on how IDE vendors are converging on unified agent surfaces.
Sourcegraph Agentic Batch Changes (public beta)
Agent scopes, executes and ships large-scale migrations across hundreds of repos until every PR is mergeable.
Why this matters: The 22,000-line-PR problem, productised — worth stress-testing against your verification playbook.
Vercel Services: multiple frameworks in one project
Public beta lets a single Vercel project run multiple frontends and backends (e.g. Next.js + FastAPI) with atomic deploys.
Why this matters: Directly relevant to your TypeScript + Python + Vercel stack.
Vercel Agent: chat, investigations, approved actions
Vercel Agent now investigates prod issues and takes approved actions using deployment, log and metric context; new $0.25/Mtok pricing.
Why this matters: Platform-native ops agent for a Vercel-hosted stack.
Running untrusted agent code without a sandbox
LangChain Deep Agents use WASM + QuickJS in-process isolation with capability-based least-privilege and snapshot pauses instead of full sandboxes.
Why this matters: Concrete sandboxing pattern for MCP tool execution — worth comparing to your in-house MCP servers.
Mistral Workflows
Mistral ships a durable, fault-tolerant orchestration platform for multi-agent pipelines.
Why this matters: Another orchestration option alongside LangGraph/Temporal for your pipelines.
Open Models & Local
DeepSeek open-sources DSpark: up to 85% faster inference
DSpark is a speculative-decoding scout that runs ahead of the main model and lets it verify safe steps, without changing outputs.
Why this matters: Meaningful for local coding-agent latency on your Apple Silicon setup.
Sakana Fugu Ultra: 93.2 on LiveCodeBench
Fugu Ultra beats Fable on LiveCodeBench and starts at $5/Mtok input.
Why this matters: New coding-model contender to slot into your model gateway for A/B.
Industry & Trends
Sonnet 5 review: 64 blind generations
Lenny ran Sonnet 5 against four other frontier models across prototype, PRD and voice-agent tasks using Claude Code.
Why this matters: Independent multi-task eval to calibrate your Sonnet 5 rollout.
Claude Sonnet 5 on GitLab Duo
First model to complete every task in GitLab’s internal benchmark suite; live across all tiers via GitLab’s AI Gateway.
Why this matters: Enterprise validation of Sonnet 5 for multi-step dev workflows.
The CIO’s choices are clear in 2026
Across 87 public SaaS companies, only Infra/Dev Tools (+68.5%) and Security (+17.6%) are up 1Y; seat-priced app layer is being sold off.
Why this matters: Directly relevant to RegTech/identity positioning and buyer psychology.
How Rippling went AI-native in 6 months
Rippling deployed Deep Agents + LangSmith across HR, IT, finance, payroll and global ops in six months.
Why this matters: Adoption-story data point at the org scale you care about.
RoadmapBench: long-horizon agentic upgrades
115 tasks across 17 repos requiring median 3,700-line, 51-file modifications for real version upgrades.
Why this matters: Closer to your team’s actual migration workload than typical coding benchmarks.
Sources unavailable today: r/ChatGPTCoding top, r/ClaudeAI top, r/LocalLLaMA top, r/MachineLearning top
Auto-curated daily by Claude Opus 4.7 from Ben’s Bites, Cursor changelog, Don’t Worry About the Vase (Zvi), GitHub: anthropics/claude-code, GitHub: cline/cline, GitHub: langchain-ai/langgraph, GitHub: ollama/ollama, GitLab blog, Google DeepMind blog, Hugging Face blog, JetBrains AI blog, LangChain blog, Lenny’s Newsletter, NVIDIA developer blog, One Useful Thing (Ethan Mollick), OpenAI blog, Sourcegraph blog, TLDR AI, The Algorithmic Bridge (Alberto Romero), The Pragmatic Engineer (Gergely Orosz), Together AI blog, Tomasz Tunguz, Vercel blog, smol.ai news. Source list and editorial profile maintained by Daniel.