Claude Opus 5.5, GPT-6 Sol & Luna, MiMo V2.6 Pro
Wednesday, 23 September 2026 - AI News · (last 24h)
Anthropic ships Opus 5.5 and OpenAI counters with GPT-6 Sol and Luna within an hour — frontier price war is on.
Must read
- Claude Opus 5.5 — New default Opus in Claude Code v2.1.280: 1M context, $4/$20 per Mtok, targeting long-running agentic coding — retest your overnight-agent-factory routes.
- Introducing GPT-6 Sol and Luna — Two GPT-6 tiers, both cheaper than Astra, already live in Vercel AI Gateway — worth wiring into your LiteLLM gateway for A/B against Opus 5.5.
- Opus 5.5, GPT-6 Sol/Luna, and a new price war — Simon’s day-one comparison across Anthropic, OpenAI, xAI and MiMo — fastest way to calibrate before you commit routing changes.
- Claude Code v2.1.280 — Opus 5.5 is now the Claude Code default, plus a new env var to raise the 2,048-char MCP tool-description cap — matters for your in-house MCP servers.
- Xiaomi open-sources MiMo-V2.6 Pro and Flash — 1T-A42B omnimodal, trained for $3M, weights and RL code public — the strongest open-weights frontier candidate this year.
Tools & Frameworks
Cline Desktop 0.0.33: git worktrees for tasks
New ‘Work in: Worktree’ switch cuts a fresh cline/
Why this matters: Direct competitor pattern to Claude Code worktrees — worth stealing for your parallel-agent setup.
Drives for Vercel Sandbox in public beta
Persistent Drives mount into Vercel Sandboxes as directories, reusable across runs to preserve agent workspaces, on-disk memory, models and dependency trees.
Why this matters: Fills the persistent-state gap for sandboxed agents on your Vercel-hosted surfaces.
Devin Cloud in your terminal
Devin CLI can now hand any local task to a Devin Cloud VM and let you steer or resume from the terminal; free SWE-2 sessions until October 8.
Why this matters: Another data point on dispatch/remote-control patterns to compare with Claude Code headless.
AWS Strands Harness
Ready-to-run agent harness that ships with web search, shell, file edits, memory and sub-agent hand-off — bring-your-own model.
Why this matters: Reference architecture worth reading before extending your own in-house harness.
Better prompt caching for GPT-6
Higher cache hit rates, explicit breakpoints, new diagnostics and controls to reduce latency and cost on GPT-6 Sol/Luna.
Why this matters: Cache tuning knobs matter if you route through LiteLLM and pay per token at scale.
Cline 4.1.20: parallel sub-agent tool calls
Sub-agents spawned in the same step now execute independent tool calls concurrently; output budget defaults to 30% of the model’s advertised limit.
Why this matters: Concrete pattern for parallelising sub-agents you can port into your orchestration layer.
Open Models & Local
Transformers now runs llama.cpp quants
Hugging Face Transformers can load llama.cpp GGUF quantisations directly, unifying the Python and llama.cpp ecosystems around a single weights format.
Why this matters: Simplifies your MLX/Ollama/llama.cpp stack on Apple Silicon — one quant, many runtimes.
oMLX creator joins Hugging Face
Jun Kim, maintainer of oMLX, joins Hugging Face to support the MLX community and align tooling with the wider HF ecosystem.
Why this matters: Signals proper MLX first-class support in HF — good for your Apple Silicon local setup.
Grok 4.7
New larger base model at $2/$6 per Mtok with improved self-verification and constrained cybersecurity command execution.
Why this matters: Pricing undercuts Opus 5.5 by 5x — watch for routing, not urgent to adopt.
Qwen RecreationWorld
Five-platform training framework for hybrid computer-use agents that explore GUIs, write code, and visually verify their own work.
Why this matters: Useful reference if you’re evaluating open computer-use agents alongside frontier options.
Industry & Trends
How AI will change operating systems: Windows
Orosz’s deepdive into Microsoft’s push to make Windows agent-friendly via Linux-on-Windows, local models and GPU access to win developers back.
Why this matters: Context for where agent-native OS primitives are heading — relevant even on a Mac-first team.
The Great Unbundling of Intelligence
Argues agent economics push architectures toward capability-level routing — cheap specialised models for ranking/verification, frontier only for hard tasks.
Why this matters: Maps onto your three-tier deterministic/ML/LLM thesis — useful framing for your writing.
AI Comes for the If Statement
Small decision models like Jev2 and SemIf claim 99% cost reductions and lift accuracy from 47% to 80%+ for if-then primitives in production.
Why this matters: Directly relevant to your identity/fraud rules layer if the numbers hold up.
Org & Leadership
GitLab cut code-per-agentic-flow by 45%
GitLab Duo’s Flow Registry compiles reusable YAML components into LangGraph flows, cutting the code needed per agentic flow by 45%.
Why this matters: Concrete customer-zero metric from the Act 2 playbook — steal the declarative flow pattern.
Sources unavailable today: Last Week in AI, r/ChatGPTCoding top, r/ClaudeAI top, r/LocalLLaMA top, r/MachineLearning top
Auto-curated daily by Claude Opus 4.7 from Ben’s Bites, Don’t Worry About the Vase (Zvi), GitHub: All-Hands-AI/OpenHands, GitHub: BerriAI/litellm, GitHub: anthropics/claude-code, GitHub: cline/cline, GitHub: langchain-ai/langchain, GitHub: simonw/llm, GitLab blog, Hugging Face blog, Interconnects (Nathan Lambert), LangChain blog, Latent Space, Lenny’s Newsletter, NVIDIA developer blog, OpenAI blog, SaaStr (Jason Lemkin), Simon Willison, TLDR AI, The Pragmatic Engineer (Gergely Orosz), Together AI blog, Vercel blog. Source list and editorial profile maintained by Daniel.