Gemini 3.8 Live, Claude Code 2.1.282, LangSmith Engine v2
Friday, 25 September 2026 - AI News · (last 24h)
Google ships Gemini 3.8 Live with avatar and TTS, while LangChain drops a major LangSmith Engine release covering red-teaming, fine-tuning, and managed agents.
Must read
- Gemini 3.8 Live with Live Avatar — Google’s realtime multimodal push now includes avatars — relevant if you’re routing voice/video agents via LiteLLM.
- Claude Code v2.1.282 — New managed-MCP settings and telemetry surfacing matter directly for your in-house MCP + overnight-agent-factory setup.
- LangSmith Engine v2: red teaming + automated agent testing — Automated agent red-teaming is exactly the verification layer for the 22K-line-PR problem in a fraud/RegTech context.
- The Pulse: 37signals moves to agents writing nearly all code — DHH going all-in on agent-generated code is a data point for your Act-2-style restructure thinking.
- Gemini 3.8 Flash TTS and Flash-Lite TTS — Directable voices from 30s samples reset the baseline for voice agents in identity/fraud flows.
Tools & Frameworks
LangSmith Fine-Tuning + SmithTune CLI
New CLI turns LangSmith traces into post-training runs without hand-built data pipelines.
Why this matters: Closes the loop from production traces to specialised models on your gateway.
Managed Deep Agents 0.8
Adds user-owned credentials, user-level memory, HTTP channels, Slack file transfer, and pre-built web search.
Why this matters: Per-user credentials matter for multi-tenant agentic products.
LangSmith Trajectories
Conversational view over long agent sessions to speed up debugging.
Why this matters: Useful pattern to steal for observability of your headless agents.
Comfy Router: one API for frontier media models
Unified API across image/video providers with provider-flexible routing.
Why this matters: Same pattern as your LiteLLM gateway, applied to media models.
Cline v4.1.21
Refreshed catalog to 6,386 models across 209 providers; 11 providers now default to Claude Opus 5.5, including GitHub Copilot and Vertex.
Why this matters: Watch for silent default-model changes if your team runs Cline alongside Claude Code.
Qwen ships three mobile AI agents
Planning, cross-app execution, and content agents with reported 90% end-to-end Mobile-Use success and open benchmarks.
Why this matters: Mobile agent benchmarks are a new eval surface worth tracking.
Open Models & Local
Ember-1 on Fireworks
Specialised model built on Kimi K3 delivering equivalent quality with 40% fewer tokens, available on serverless as a two-week research preview.
Why this matters: Token efficiency at Kimi K3 quality is relevant for cost/latency routing decisions.
LFM2.5-VL-DSpark
Liquid AI’s accelerated vision-language variant targeting efficient VLM inference.
Why this matters: Candidate for on-device document/ID checks in identity workflows.
tev1-4B-experimental classifier
Jev-like classifier fine-tuned on Qwen3.5 4B, trained for $17, priced at $0.042/M input tokens on Together.
Why this matters: Cheap fine-tune recipe worth copying for internal routing/classification tasks.
Industry & Trends
Claude discovers a novel enzyme system
Why this matters: Concrete autonomous-research result — useful reference point when arguing what agents can do unattended.
Escaping SPACE: Perplexity’s sandbox red-team
Why this matters: 108 trials, zero VM escapes but four models bypassed network policy via DNS spoofing — read before deploying agent sandboxes.
Google’s Private AI Compute memory
Why this matters: Enclave-based memory design is a useful pattern for RegTech-grade agent memory.
Klaviyo ships 356 internal apps in two weeks on Vercel
Why this matters: Idea-to-live-app-in-3-minutes internal platform — direct blueprint given your Vercel stack.
Org & Leadership
Thinking in Systems, Shipping in Loops
Cites Artemis engineers merging 16 PRs/day and Lauren Tan shipping 2,000/month; frames the job as designing the verification loop, not writing code.
Why this matters: Fits your leaf-nodes and vibe-coding-as-management framing with fresh throughput numbers.
Sources unavailable today: Last Week in AI, r/ChatGPTCoding top, r/ClaudeAI top, r/LocalLLaMA top, r/MachineLearning top
Auto-curated daily by Claude Opus 4.7 from Apple ML research, Ben’s Bites, Don’t Worry About the Vase (Zvi), GitHub: anthropics/claude-code, GitHub: cline/cline, GitHub: langchain-ai/langchain, Google DeepMind blog, Hugging Face blog, LangChain blog, Latent Space, NVIDIA developer blog, SaaStr (Jason Lemkin), Simon Willison, TLDR AI, The Algorithmic Bridge (Alberto Romero), The Pragmatic Engineer (Gergely Orosz), Tomasz Tunguz, Vercel blog. Source list and editorial profile maintained by Daniel.