Skip to content

← AI Tracker

AI Briefing

Kimi K3, Claude Code Migrations, LM Studio Bionic

Samstag, 18. Juli 2026 - AI News · (letzte 24h)

Moonshot released Kimi K3, a 2.8T-param multimodal model with 1M context and agentic coding focus; weights land July 27.

Must read

Tools & Frameworks

Claude Code v2.1.214

Fixes a permission-check bypass in PowerShell 5.1, a dir/** allow-rule over-match, and misjudged commands over 10,000 characters.

Why this matters: Security-relevant fixes for anyone running Claude Code with allow-lists in production.

GitHub Copilot SDK

New SDK embeds Copilot agents into custom applications and developer tools.

Why this matters: Another dispatch surface for agent orchestration alongside Claude Code SDK.

Cursor Slack improvements

Cursor’s Slack integration gets an update on July 17.

Why this matters: Relevant if you’re triggering Cursor background agents from Slack.

Vercel Chat SDK adds native Slack agent support

Slack adapter now supports Agent badge, Messages tab conversations, suggested prompts, rotating status, and native feedback buttons.

Why this matters: Cheap path to production Slack agents on your Vercel + TS stack.

Vercel Plugin now in Kimi Code CLI

Kimi Code CLI gets a Vercel Plugin with skills for Next.js, AI SDK, and Vercel Functions.

Why this matters: Kimi K3 in a coding CLI with real deploy context — worth a spike.

Vercel Sandbox downloads now free

Package installs, git clones, and dataset pulls no longer count toward Sandbox Data Transfer billing.

Why this matters: Meaningful cost cut for agent sandboxes on Vercel.

Harness Handbook for coding agents

Behaviour-level map connecting execution, permissions, and safety questions to the prompts, tools, state logic, config, and telemetry that implement them.

Why this matters: Discipline layer for the 22K-line-PR problem — good reference for your writing.

The best model routing is task-specific

Jerry Liu argues narrow workflow focus yields the most alpha in accuracy and cost when routing between models.

Why this matters: Directly applies to your LiteLLM gateway routing decisions.

Choosing GPT-5.6 Sol, Terra, or Luna in Codex

Field guide: Sol for ambiguous high-value problems, Terra for everyday implementation, Luna for fast bounded tasks; Sol Ultra adds multi-agent coordination.

Why this matters: Task-tier framing you can adopt for internal routing rules.

CrewAI 1.15.4 promotes Skills Repository

CrewAI moves its Skills Repository out of experimental status in 1.15.4.

Why this matters: Skills framework momentum beyond Anthropic — watch as a pattern, not urgent.

Open Models & Local

GLM 5.2 is 35% off via Novita on AI Gateway

GLM 5.2 discounted 35% through July 24 when routed via Novita on Vercel AI Gateway using zai/glm-5.2.

Why this matters: Cheap window to benchmark GLM 5.2 against your current coding tier.

llama.cpp adds Vulkan Q2_0 support

Vulkan backend gains Q2_0 with doubled rows per workgroup fixing early mat-vec-mul perf; multiple b100xx releases shipped in the window.

Why this matters: Marginal but relevant for local inference tuning on mixed hardware.

OpenAI’s AI scorecard from CFO Sarah Friar

Four metrics proposed: useful work, cost per successful task, dependability, and return on compute.

Why this matters: Useful vocabulary for board conversations about AI ROI.

Ramp expands AI token spend tracking

Ramp’s AI Token Spend Management now covers OpenAI, Anthropic, and Gemini with cross-provider visibility for finance teams.

Why this matters: Alternative to building your own token accounting on top of LiteLLM.

Fireworks hits $17.5B valuation on cheap-inference demand

Nvidia-backed Fireworks reaches $17.5B as buyers seek cheaper open-source model hosting.

Why this matters: Signal that open-weight hosting economics are hardening — watch for gateway pricing pressure.

NotebookLM becomes Gemini Notebook

Google folds NotebookLM into the Gemini app and Search under a new name.

Why this matters: Brand consolidation; watch but no action needed.

Org & Leadership

Replit’s self-driving company

Replit engineers report 3x code output using AI agents for PR review, incident investigation, and data analysis without quality loss.

Why this matters: Concrete adoption pattern from a named eng org — comparison point to GitLab Act 2.


Sources unavailable today: r/ChatGPTCoding top, r/ClaudeAI top, r/LocalLLaMA top, r/MachineLearning top

Auto-curated daily by Claude Opus 4.7 from Apple ML research, Cursor changelog, Don’t Worry About the Vase (Zvi), GitHub: anthropics/claude-code, GitHub: crewAIInc/crewAI, GitHub: ggml-org/llama.cpp, Hugging Face blog, LangChain blog, Latent Space, Not Boring (Packy McCormick), OpenAI blog, SaaStr (Jason Lemkin), Simon Willison, TLDR AI, Vercel blog, smol.ai news. Source list and editorial profile maintained by Daniel.