Inkling 975B open, Bonsai 27B on phone, Claude Code worktrees
Donnerstag, 16. Juli 2026 - AI News · (letzte 24h)
Thinking Machines Lab drops Inkling — a 975B/41B-active Apache 2.0 multimodal MoE with 1M context — as its first open-weights release.
Must read
- Thinking Machines releases Inkling: 975B-A41B Apache 2.0 multimodal — Best American open Apache 2.0 model to date, 1M context, day-0 on Together and Vercel AI Gateway — route it via your LiteLLM setup.
- Bonsai 27B: first 27B-class model to run on a phone — 1-bit/ternary quantisation at 3.9GB with tool use intact — the local-LLM story on Apple Silicon just got sharper.
- Claude Code v2.1.210: worktree isolation fix and elapsed-time counter — Fixes worktree subagents mutating the main checkout — critical if your overnight agent factory runs parallel Claude Code sessions.
- Ayush Paul finds exfiltration hole in Claude web_fetch — Direct hit on the tool you rely on; review your Claude Code fetch permissions and MCP sandboxing today.
- How Microsoft ships thousands of production AI agents — Rubric-based evals, agent identities, retrieval-as-sub-agent — concrete playbook mirrors your three-tier architecture thinking.
Tools & Frameworks
Claude Code v2.1.211 adds —forward-subagent-text
New flag and CLAUDE_CODE_FORWARD_SUBAGENT_TEXT env var expose subagent text and thinking in stream-json; also hardens permission-preview against unicode spoofing.
Why this matters: Better observability for your headless agent dispatch workflows.
Google’s Mantis Skills: security review harness for coding agents
Decoupled, sequential agent-skills toolkit for security review, designed to be tuned with internal docs and coding standards rather than used as-is.
Why this matters: Directly usable Skills pattern for RegTech code review — worth forking.
Long-Horizon Terminal-Bench: 46 stateful terminal tasks
Hidden verifiers rebuild and inspect final artifacts across hundreds of terminal interactions rather than trusting agent-reported progress.
Why this matters: Verification-not-trust matches your 22k-line-PR problem — useful eval methodology.
LangSmith adds tracing for Claude Code, Codex, Cursor, Copilot
LangSmith now traces tool calls, subagents, errors, costs, and retries across the major coding-agent CLIs from one place.
Why this matters: Cross-tool observability for teams running Claude Code and Cursor side by side.
LangChain on per-agent sandboxes booting in under a second
Argues each agent needs its own isolated computer; sandboxes boot fast and tear down cleanly, replacing shared VMs and containers.
Why this matters: Relevant infra pattern if you’re scaling parallel headless agents beyond git worktrees.
Inkling available on Vercel AI Gateway day-0
Set model to thinkingmachines/inkling in the AI SDK; supports controllable thinking effort across coding, reasoning, vision, and audio.
Why this matters: Route Inkling behind your LiteLLM gateway for A/B tests against Claude and GPT-5.6.
Prime Intellect ships verifiers v1 for agentic RL and evals
Decomposes environments into task set, harness, and runtime; supports complex agentic coding and computer-use tasks at scale.
Why this matters: Eval infra worth watching as harness quality becomes the differentiator.
Open Models & Local
Google announces Gemma 4 E2B optimised for Pixel 10 TPU
Runs natively on-device for offline conversation, image ID, audio transcription, and voice-driven phone control.
Why this matters: Signals where on-device model quality is heading; watch for MLX/Ollama ports.
The State of Open Source AI report
Majority of production tokens now route through open weights; five highest-volume OpenRouter models are all open.
Why this matters: Data for your local-plus-cloud routing arguments — hybrid is now the median case.
DeepSeek reportedly raising $1.5B ahead of IPO
Chinese open-weights lab in talks for pre-IPO round after DeepSeek V4 traction across llama.cpp and hosted providers.
Why this matters: Funding runway for the model family currently anchoring cheap open-coding stacks.
Industry & Trends
xAI open-sources grok-build after CLI uploads full home directories to GCS
xAI’s grok CLI was found uploading whole directories — including SSH keys and password managers — to xAI’s Google Cloud buckets; now open source under scrutiny.
Why this matters: Cautionary tale for vetting any coding CLI you install; audit MCP servers similarly.
Tomasz Tunguz: the harness is the new battleground
The software wrapping the model — deciding what data flows in, what gets logged, what trains the next model — is now the strategic asset, not the weights.
Why this matters: Frames why your in-house MCP servers and gateway are the moat.
Cognition: swapping Opus 4.8 for Fable 5 dropped Devin’s bill
Fable 5 costs 2x per token but architectural changes made it cheaper overall and score higher — a case study in agent economics.
Why this matters: Concrete evidence that harness design beats raw token price for agentic work.
Latent Space: 5 trends from AI Engineer World’s Fair 2026
Argues AI engineering has shifted to building systems around agents rather than building with agents.
Why this matters: Grounding read before AI Engineer World’s Fair 2026 in San Francisco.
Pragmatic Engineer: context engineering with Dex Horthy
Long-form interview on why context engineering — not prompt engineering — is the discipline that determines AI-assisted code quality.
Why this matters: Overlaps directly with your skills/spec framework writing.
Org & Leadership
Higgsfield: $500M ARR with 60 engineers, cash-flow positive
CEO Alex Mashrabov details how a 60-engineer team runs an AI video product at $500M ARR while cash-flow positive.
Why this matters: Concrete data point for the leverage-shift argument — small teams shipping at scale.
The Great Flattening
Argues coding orgs will flatten as multi-agent systems replace traditional engineering workflows, leaving customer insight as the human edge.
Why this matters: Adjacent to GitLab Act 2 thesis; watch for named companies making equivalent moves.
Sources unavailable today: r/ChatGPTCoding top, r/ClaudeAI top, r/LocalLLaMA top, r/MachineLearning top
Auto-curated daily by Claude Opus 4.7 from Apple ML research, Ben’s Bites, Don’t Worry About the Vase (Zvi), GitHub: anthropics/claude-code, GitHub: ggml-org/llama.cpp, Hugging Face blog, LangChain blog, Latent Space, NVIDIA developer blog, OpenAI blog, SaaStr (Jason Lemkin), Simon Willison, TLDR AI, The Pragmatic Engineer (Gergely Orosz), Together AI blog, Tomasz Tunguz, Vercel blog, smol.ai news. Source list and editorial profile maintained by Daniel.