Skip to content

← AI Tracker

AI Briefing

Inkling 975B open, Bonsai 27B on phone, Claude Code worktrees

Donnerstag, 16. Juli 2026 - AI News · (letzte 24h)

Thinking Machines Lab drops Inkling — a 975B/41B-active Apache 2.0 multimodal MoE with 1M context — as its first open-weights release.

Must read

Tools & Frameworks

Claude Code v2.1.211 adds —forward-subagent-text

New flag and CLAUDE_CODE_FORWARD_SUBAGENT_TEXT env var expose subagent text and thinking in stream-json; also hardens permission-preview against unicode spoofing.

Why this matters: Better observability for your headless agent dispatch workflows.

Google’s Mantis Skills: security review harness for coding agents

Decoupled, sequential agent-skills toolkit for security review, designed to be tuned with internal docs and coding standards rather than used as-is.

Why this matters: Directly usable Skills pattern for RegTech code review — worth forking.

Long-Horizon Terminal-Bench: 46 stateful terminal tasks

Hidden verifiers rebuild and inspect final artifacts across hundreds of terminal interactions rather than trusting agent-reported progress.

Why this matters: Verification-not-trust matches your 22k-line-PR problem — useful eval methodology.

LangSmith adds tracing for Claude Code, Codex, Cursor, Copilot

LangSmith now traces tool calls, subagents, errors, costs, and retries across the major coding-agent CLIs from one place.

Why this matters: Cross-tool observability for teams running Claude Code and Cursor side by side.

LangChain on per-agent sandboxes booting in under a second

Argues each agent needs its own isolated computer; sandboxes boot fast and tear down cleanly, replacing shared VMs and containers.

Why this matters: Relevant infra pattern if you’re scaling parallel headless agents beyond git worktrees.

Inkling available on Vercel AI Gateway day-0

Set model to thinkingmachines/inkling in the AI SDK; supports controllable thinking effort across coding, reasoning, vision, and audio.

Why this matters: Route Inkling behind your LiteLLM gateway for A/B tests against Claude and GPT-5.6.

Prime Intellect ships verifiers v1 for agentic RL and evals

Decomposes environments into task set, harness, and runtime; supports complex agentic coding and computer-use tasks at scale.

Why this matters: Eval infra worth watching as harness quality becomes the differentiator.

Open Models & Local

Google announces Gemma 4 E2B optimised for Pixel 10 TPU

Runs natively on-device for offline conversation, image ID, audio transcription, and voice-driven phone control.

Why this matters: Signals where on-device model quality is heading; watch for MLX/Ollama ports.

The State of Open Source AI report

Majority of production tokens now route through open weights; five highest-volume OpenRouter models are all open.

Why this matters: Data for your local-plus-cloud routing arguments — hybrid is now the median case.

DeepSeek reportedly raising $1.5B ahead of IPO

Chinese open-weights lab in talks for pre-IPO round after DeepSeek V4 traction across llama.cpp and hosted providers.

Why this matters: Funding runway for the model family currently anchoring cheap open-coding stacks.

xAI open-sources grok-build after CLI uploads full home directories to GCS

xAI’s grok CLI was found uploading whole directories — including SSH keys and password managers — to xAI’s Google Cloud buckets; now open source under scrutiny.

Why this matters: Cautionary tale for vetting any coding CLI you install; audit MCP servers similarly.

Tomasz Tunguz: the harness is the new battleground

The software wrapping the model — deciding what data flows in, what gets logged, what trains the next model — is now the strategic asset, not the weights.

Why this matters: Frames why your in-house MCP servers and gateway are the moat.

Cognition: swapping Opus 4.8 for Fable 5 dropped Devin’s bill

Fable 5 costs 2x per token but architectural changes made it cheaper overall and score higher — a case study in agent economics.

Why this matters: Concrete evidence that harness design beats raw token price for agentic work.

Argues AI engineering has shifted to building systems around agents rather than building with agents.

Why this matters: Grounding read before AI Engineer World’s Fair 2026 in San Francisco.

Pragmatic Engineer: context engineering with Dex Horthy

Long-form interview on why context engineering — not prompt engineering — is the discipline that determines AI-assisted code quality.

Why this matters: Overlaps directly with your skills/spec framework writing.

Org & Leadership

Higgsfield: $500M ARR with 60 engineers, cash-flow positive

CEO Alex Mashrabov details how a 60-engineer team runs an AI video product at $500M ARR while cash-flow positive.

Why this matters: Concrete data point for the leverage-shift argument — small teams shipping at scale.

The Great Flattening

Argues coding orgs will flatten as multi-agent systems replace traditional engineering workflows, leaving customer insight as the human edge.

Why this matters: Adjacent to GitLab Act 2 thesis; watch for named companies making equivalent moves.


Sources unavailable today: r/ChatGPTCoding top, r/ClaudeAI top, r/LocalLLaMA top, r/MachineLearning top

Auto-curated daily by Claude Opus 4.7 from Apple ML research, Ben’s Bites, Don’t Worry About the Vase (Zvi), GitHub: anthropics/claude-code, GitHub: ggml-org/llama.cpp, Hugging Face blog, LangChain blog, Latent Space, NVIDIA developer blog, OpenAI blog, SaaStr (Jason Lemkin), Simon Willison, TLDR AI, The Pragmatic Engineer (Gergely Orosz), Together AI blog, Tomasz Tunguz, Vercel blog, smol.ai news. Source list and editorial profile maintained by Daniel.