Skip to content

← AI Tracker

AI Briefing

Gemini 3.8 Live, Claude Code 2.1.282, LangSmith Engine v2

Friday, 25 September 2026 - AI News · (last 24h)

Google ships Gemini 3.8 Live with avatar and TTS, while LangChain drops a major LangSmith Engine release covering red-teaming, fine-tuning, and managed agents.

Must read

Tools & Frameworks

LangSmith Fine-Tuning + SmithTune CLI

New CLI turns LangSmith traces into post-training runs without hand-built data pipelines.

Why this matters: Closes the loop from production traces to specialised models on your gateway.

Managed Deep Agents 0.8

Adds user-owned credentials, user-level memory, HTTP channels, Slack file transfer, and pre-built web search.

Why this matters: Per-user credentials matter for multi-tenant agentic products.

LangSmith Trajectories

Conversational view over long agent sessions to speed up debugging.

Why this matters: Useful pattern to steal for observability of your headless agents.

Comfy Router: one API for frontier media models

Unified API across image/video providers with provider-flexible routing.

Why this matters: Same pattern as your LiteLLM gateway, applied to media models.

Cline v4.1.21

Refreshed catalog to 6,386 models across 209 providers; 11 providers now default to Claude Opus 5.5, including GitHub Copilot and Vertex.

Why this matters: Watch for silent default-model changes if your team runs Cline alongside Claude Code.

Qwen ships three mobile AI agents

Planning, cross-app execution, and content agents with reported 90% end-to-end Mobile-Use success and open benchmarks.

Why this matters: Mobile agent benchmarks are a new eval surface worth tracking.

Open Models & Local

Ember-1 on Fireworks

Specialised model built on Kimi K3 delivering equivalent quality with 40% fewer tokens, available on serverless as a two-week research preview.

Why this matters: Token efficiency at Kimi K3 quality is relevant for cost/latency routing decisions.

LFM2.5-VL-DSpark

Liquid AI’s accelerated vision-language variant targeting efficient VLM inference.

Why this matters: Candidate for on-device document/ID checks in identity workflows.

tev1-4B-experimental classifier

Jev-like classifier fine-tuned on Qwen3.5 4B, trained for $17, priced at $0.042/M input tokens on Together.

Why this matters: Cheap fine-tune recipe worth copying for internal routing/classification tasks.

Claude discovers a novel enzyme system

Why this matters: Concrete autonomous-research result — useful reference point when arguing what agents can do unattended.

Escaping SPACE: Perplexity’s sandbox red-team

Why this matters: 108 trials, zero VM escapes but four models bypassed network policy via DNS spoofing — read before deploying agent sandboxes.

Google’s Private AI Compute memory

Why this matters: Enclave-based memory design is a useful pattern for RegTech-grade agent memory.

Klaviyo ships 356 internal apps in two weeks on Vercel

Why this matters: Idea-to-live-app-in-3-minutes internal platform — direct blueprint given your Vercel stack.

Org & Leadership

Thinking in Systems, Shipping in Loops

Cites Artemis engineers merging 16 PRs/day and Lauren Tan shipping 2,000/month; frames the job as designing the verification loop, not writing code.

Why this matters: Fits your leaf-nodes and vibe-coding-as-management framing with fresh throughput numbers.


Sources unavailable today: Last Week in AI, r/ChatGPTCoding top, r/ClaudeAI top, r/LocalLLaMA top, r/MachineLearning top

Auto-curated daily by Claude Opus 4.7 from Apple ML research, Ben’s Bites, Don’t Worry About the Vase (Zvi), GitHub: anthropics/claude-code, GitHub: cline/cline, GitHub: langchain-ai/langchain, Google DeepMind blog, Hugging Face blog, LangChain blog, Latent Space, NVIDIA developer blog, SaaStr (Jason Lemkin), Simon Willison, TLDR AI, The Algorithmic Bridge (Alberto Romero), The Pragmatic Engineer (Gergely Orosz), Tomasz Tunguz, Vercel blog. Source list and editorial profile maintained by Daniel.