GPT-5.6 Sol/Terra/Luna, Grok 4.5 + Cursor, SWE-1.7 Frontier Cheap
jeudi 9 juillet 2026 - AI News · (24 dernières heures)
OpenAI shipped the GPT-5.6 family (Sol/Terra/Luna) with programmatic tool calling and a new ChatGPT Work agent; SpaceXAI’s Grok 4.5 trained alongside Cursor lands the same day.
Must read
- GPT-5.6: Sol, Terra, Luna — Three-tier lineup at $1/$6, $2.50/$15, $5/$30 with 90% cache-read discount — retune your LiteLLM routing today.
- Grok 4.5, trained with Cursor — First Opus-class model post-Cursor acquisition; likely to reshape Cursor’s default coding model in your team’s workflow.
- SWE-1.7 hits frontier at fraction of cost — Devin’s model claims frontier coding at commodity pricing — directly relevant to overnight-agent-factory economics.
- ChatGPT Work: long-horizon agent — Merges Codex + ChatGPT into a desktop app with multi-hour project persistence — sets expectations for internal agent UX.
- 30% of SWE-Bench Pro tasks broken — OpenAI’s audit undermines the benchmark most vendors quote — reset how you evaluate coding models internally.
Tools & Frameworks
GPT-Live full-duplex voice
Full-duplex voice model listens and speaks simultaneously, delegating complex tasks to GPT-5.5 mid-conversation.
Why this matters: Voice-first agent patterns worth tracking for identity/fraud verification flows.
GPT-5.6 on Vercel AI Gateway
Sol, Terra, and Luna are live on AI Gateway in limited preview with agentic-optimised token efficiency.
Why this matters: Alternative routing path if you extend LiteLLM with Vercel-fronted endpoints.
Muse Spark 1.1 API launch
Meta’s 1M-context multimodal model gains an API with parallel tool calling, MCP support, and custom skills.
Why this matters: First Meta model that plausibly slots into an agentic dev pipeline; MCP support matters.
OpenHands 1.11.0
Adds Agent Profiles on cloud backend, Budgets dashboard, expanded usage monitoring, and repo metadata in observability traces.
Why this matters: Budget controls and profiles are the discipline layer your parallel-agent setup needs.
Cline v4.0.7
Adds ClinePass limit-reached error with one-click switch to usage-based billing and free-model tab.
Why this matters: Minor; watch as a Cursor/Claude Code alternative for team members.
Open Models & Local
Ollama Series B: 8.9M users
Ollama raises Series B led by Theory with 8.9M developers and 67,000 integrations across major model labs and hardware.
Why this matters: Ollama’s trajectory affects your local-plus-cloud hybrid setup on Apple Silicon.
Mistral Robostral Navigate 8B
8B model enables single-RGB-camera autonomous robot navigation from Mistral.
Why this matters: Watch but don’t act — off-stack for identity/RegTech.
Industry & Trends
Bun’s 11-day Rust rewrite with AI
Orosz breaks down a rewrite done in 11 days for $165K in tokens that would have taken a small team a year.
Why this matters: Concrete data point for your writing on leverage shift and the one-person-team thesis.
GPT-5.6 Sol vs Claude Fable benchmark
Head-to-head across prototypes, PRDs, and browser use; Sol wins on prototypes and PRDs.
Why this matters: Practical comparison for your Claude Code vs OpenAI routing decisions.
OpenAI acquires Northslope
OpenAI buys applied-AI firm Northslope, adding hundreds of forward-deployed engineers to enterprise ranks.
Why this matters: Signals OpenAI moving into Palantir-style FDE motion — competitive pressure on your vendor stack.
Ways to think about token pricing
Evans on whether model labs stay premium or become low-margin commodity infrastructure as supply crunch resolves.
Why this matters: Frames the cost curve underlying your build-vs-buy planning.
Anthropic GRAM: removable knowledge modules
Gradient-Routed Auxiliary Modules give models compartments for dual-use knowledge that can be deleted post-training.
Why this matters: Relevant safety pattern if you fine-tune models handling sensitive identity data.
Org & Leadership
Mosseri on Instagram’s 2026 product structure
Instagram’s head describes new product team structure, the rise of the ‘product staff’ role, and current hiring traits.
Why this matters: Named operating-model change from a large product org — compare against your Act-2-style thinking.
Sources unavailable today: r/ChatGPTCoding top, r/ClaudeAI top, r/LocalLLaMA top, r/MachineLearning top
Auto-curated daily by Claude Opus 4.7 from Apple ML research, Ben’s Bites, Benedict Evans, Don’t Worry About the Vase (Zvi), GitHub: All-Hands-AI/OpenHands, GitHub: cline/cline, GitHub: langchain-ai/langchain, GitHub: simonw/llm, GitLab blog, Latent Space, Lenny’s Newsletter, NVIDIA developer blog, OpenAI blog, Simon Willison, TLDR AI, The Algorithmic Bridge (Alberto Romero), The Pragmatic Engineer (Gergely Orosz), Tomasz Tunguz, Vercel blog, smol.ai news. Source list and editorial profile maintained by Daniel.