Kimi K3, Claude Code Migrations, LM Studio Bionic
Samstag, 18. Juli 2026 - AI News · (letzte 24h)
Moonshot released Kimi K3, a 2.8T-param multimodal model with 1M context and agentic coding focus; weights land July 27.
Must read
- Kimi K3 launches with 2.8T params and 1M context — Open weights on July 27 with coding/agentic focus — direct candidate for your local-plus-cloud routing setup.
- How Anthropic runs large-scale code migrations with Claude Code — Six-step rulebook with adversarial reviewers and mechanical verification — reference pattern for your overnight-agent-factory work.
- LM Studio Bionic: agent for open models — Local/cloud hybrid agent with sandbox and GLM 5.2 coding support — fits your Apple Silicon setup.
- Replit’s self-driving company — 3x code output — Concrete before/after from an eng org using agents for PR review, incidents, and analysis — feeds your public writing.
- Gemini 3.5 Pro delayed for coding improvements — Frontier coding model slip affects gateway routing plans; Alphabet dropped 4% on the news.
Tools & Frameworks
Claude Code v2.1.214
Fixes a permission-check bypass in PowerShell 5.1, a dir/** allow-rule over-match, and misjudged commands over 10,000 characters.
Why this matters: Security-relevant fixes for anyone running Claude Code with allow-lists in production.
GitHub Copilot SDK
New SDK embeds Copilot agents into custom applications and developer tools.
Why this matters: Another dispatch surface for agent orchestration alongside Claude Code SDK.
Cursor Slack improvements
Cursor’s Slack integration gets an update on July 17.
Why this matters: Relevant if you’re triggering Cursor background agents from Slack.
Vercel Chat SDK adds native Slack agent support
Slack adapter now supports Agent badge, Messages tab conversations, suggested prompts, rotating status, and native feedback buttons.
Why this matters: Cheap path to production Slack agents on your Vercel + TS stack.
Vercel Plugin now in Kimi Code CLI
Kimi Code CLI gets a Vercel Plugin with skills for Next.js, AI SDK, and Vercel Functions.
Why this matters: Kimi K3 in a coding CLI with real deploy context — worth a spike.
Vercel Sandbox downloads now free
Package installs, git clones, and dataset pulls no longer count toward Sandbox Data Transfer billing.
Why this matters: Meaningful cost cut for agent sandboxes on Vercel.
Harness Handbook for coding agents
Behaviour-level map connecting execution, permissions, and safety questions to the prompts, tools, state logic, config, and telemetry that implement them.
Why this matters: Discipline layer for the 22K-line-PR problem — good reference for your writing.
The best model routing is task-specific
Jerry Liu argues narrow workflow focus yields the most alpha in accuracy and cost when routing between models.
Why this matters: Directly applies to your LiteLLM gateway routing decisions.
Choosing GPT-5.6 Sol, Terra, or Luna in Codex
Field guide: Sol for ambiguous high-value problems, Terra for everyday implementation, Luna for fast bounded tasks; Sol Ultra adds multi-agent coordination.
Why this matters: Task-tier framing you can adopt for internal routing rules.
CrewAI 1.15.4 promotes Skills Repository
CrewAI moves its Skills Repository out of experimental status in 1.15.4.
Why this matters: Skills framework momentum beyond Anthropic — watch as a pattern, not urgent.
Open Models & Local
GLM 5.2 is 35% off via Novita on AI Gateway
GLM 5.2 discounted 35% through July 24 when routed via Novita on Vercel AI Gateway using zai/glm-5.2.
Why this matters: Cheap window to benchmark GLM 5.2 against your current coding tier.
llama.cpp adds Vulkan Q2_0 support
Vulkan backend gains Q2_0 with doubled rows per workgroup fixing early mat-vec-mul perf; multiple b100xx releases shipped in the window.
Why this matters: Marginal but relevant for local inference tuning on mixed hardware.
Industry & Trends
OpenAI’s AI scorecard from CFO Sarah Friar
Four metrics proposed: useful work, cost per successful task, dependability, and return on compute.
Why this matters: Useful vocabulary for board conversations about AI ROI.
Ramp expands AI token spend tracking
Ramp’s AI Token Spend Management now covers OpenAI, Anthropic, and Gemini with cross-provider visibility for finance teams.
Why this matters: Alternative to building your own token accounting on top of LiteLLM.
Fireworks hits $17.5B valuation on cheap-inference demand
Nvidia-backed Fireworks reaches $17.5B as buyers seek cheaper open-source model hosting.
Why this matters: Signal that open-weight hosting economics are hardening — watch for gateway pricing pressure.
NotebookLM becomes Gemini Notebook
Google folds NotebookLM into the Gemini app and Search under a new name.
Why this matters: Brand consolidation; watch but no action needed.
Org & Leadership
Replit’s self-driving company
Replit engineers report 3x code output using AI agents for PR review, incident investigation, and data analysis without quality loss.
Why this matters: Concrete adoption pattern from a named eng org — comparison point to GitLab Act 2.
Sources unavailable today: r/ChatGPTCoding top, r/ClaudeAI top, r/LocalLLaMA top, r/MachineLearning top
Auto-curated daily by Claude Opus 4.7 from Apple ML research, Cursor changelog, Don’t Worry About the Vase (Zvi), GitHub: anthropics/claude-code, GitHub: crewAIInc/crewAI, GitHub: ggml-org/llama.cpp, Hugging Face blog, LangChain blog, Latent Space, Not Boring (Packy McCormick), OpenAI blog, SaaStr (Jason Lemkin), Simon Willison, TLDR AI, Vercel blog, smol.ai news. Source list and editorial profile maintained by Daniel.