Gemini 3.8 Live, Claude Code 2.1.273, Cline Desktop
Wednesday, 16 September 2026 - AI News · (last 24h)
Google shipped Gemini 3.8 Live and Live Extended Thinking, two speech-to-speech models mirroring OpenAI’s GPT-Live family.
Must read
- Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking — New Google speech-to-speech models, already live on Vercel AI Gateway — routable via your LiteLLM setup today.
- Claude Code v2.1.273 — New gateway hint headers (request-class, agent-type, prev-tool-durations) and remote-control session forking — directly affects your LiteLLM routing and overnight-agent-factory setup.
- Inside OpenAI’s agentic software factory — Concrete look at how Codex is restructuring engineering inside OpenAI — reference material for your Act-2-style thinking.
- Cline Desktop: open-source app for open-weight models — Open-source parallel agent runner for local models — fits your Apple Silicon + hybrid cloud coding stack.
- Bad benchmarks and evals: Senior SWE-Bench, napkin math, and winter tires — Dan Luu on how to actually design evals — useful for the verification-of-agent-work problem you keep writing about.
Tools & Frameworks
Gemini 3.8 Live models on Vercel AI Gateway
Both gemini-3.8-live and gemini-3.8-live-extended-thinking are exposed on AI Gateway with real-time audio and visual grounding.
Why this matters: Drop-in via your model gateway; no direct Google integration needed.
Cline CLI v3.0.62 with Agent Plugins hub
CLI now surfaces Cline Desktop for open-weight models, adds scheduled runs, task import from Claude Code and Codex, and a plugin/MCP/skills marketplace.
Why this matters: Claude Code task import lowers switching cost for parallel local-model experiments.
Tau: terminal coding agent from Hugging Face
Minimal terminal-based coding agent with file editing, command execution, and durable sessions over OpenAI-compatible endpoints — designed as a teaching reference.
Why this matters: Readable reference implementation for anyone building in-house agent tooling.
Google ARTEMIS: Android agent for real devices
Open-source agent that drives real Android phones from natural-language instructions with element-index, coordinate, and visual locating fallbacks.
Why this matters: Useful reference if mobile QA touches your identity/fraud verification flows.
Your Agent Aced the Task. Will It Do It Again?
IBM Research on measuring agent consistency across repeated runs, not just single-shot success — with an evolve-consistency toolkit.
Why this matters: Direct answer to the verify-work-you-can’t-read problem in agentic pipelines.
Open Models & Local
Dense vs MoE: when to choose each (Nemotron 3.5 Lightning)
NVIDIA breaks down 30B-param MoE activating only 3B/token versus dense models on throughput, latency, and quality trade-offs.
Why this matters: Framework for local-vs-cloud routing calls on Apple Silicon.
StepAudio 3 Technical Report
Discrete autoregressive model over shared RVQ audio tokens produces speech, voices, sound effects, music and mixed audio in one model; SOTA TTS results.
Why this matters: Watch-don’t-act unless you have voice product ambitions.
Industry & Trends
Anthropic prepares Claude Money for personal finance
Why this matters: Anthropic pushing into a regulated finance vertical — relevant precedent for identity/RegTech vendor patterns.
Apple’s Siri can be swapped for Claude or ChatGPT
Why this matters: iOS 27 frameworks let Claude/GPT-5.6 be the Siri brain via Model Delegation — distribution shift for consumer AI.
Vercel: 90% inbound SDR automation, team cut 10→1.25
Why this matters: Concrete before/after numbers on agent-driven role right-sizing — Act-2 evidence in the wild.
OpenAI buys Glass Imaging for $300M+
Why this matters: Signals OpenAI’s Jony Ive device is real hardware — watch, not act.
Org & Leadership
Delphi: 10 engineers, 100+ prod deploys/day, no infra role
Why this matters: Small-team-large-leverage case study — everyone ships including product and design, matches your one-person-team framing.
Scaling agents in regulated industries: Madrigal, Abridge, Vizient
Why this matters: Constraint-heavy agent deployment lessons that map directly to identity/fraud/RegTech.
Sources unavailable today: Last Week in AI, r/ChatGPTCoding top, r/ClaudeAI top, r/LocalLLaMA top, r/MachineLearning top
Auto-curated daily by Claude Opus 4.7 from Ben’s Bites, Don’t Worry About the Vase (Zvi), Exponential View (Azeem Azhar), GitHub: anthropics/claude-code, GitHub: cline/cline, Google DeepMind blog, Hugging Face blog, LangChain blog, Latent Space, Lenny’s Newsletter, NVIDIA developer blog, SaaStr (Jason Lemkin), Simon Willison, TLDR AI, The Pragmatic Engineer (Gergely Orosz), Tomasz Tunguz, Vercel blog. Source list and editorial profile maintained by Daniel.