Ramp's In-House Agent Inspect, Vercel Run SDK, Ox Alpha 26T Tokens
mercredi 26 août 2026 - AI News · (24 dernières heures)
Ramp shipped its own in-house coding agent Inspect, choosing homebrew over Claude Code and Cursor — a rare data point on when to build vs buy.
Must read
- Why Ramp built its own in-house coding agent, Inspect — Fintech peer rejected off-the-shelf agents and built their own — direct evidence for your build-vs-buy calculus on agentic dev tooling.
- Introducing Run SDK: secure eval for your agents — Sandboxed TypeScript eval with pause/resume for auth and human approval — directly relevant to your identity/fraud agent workflows on Vercel.
- Anonymous Ox Alpha processes 26T tokens on OpenCode — 26T tokens and 327k users in four days on a free OpenAI-compatible endpoint — worth routing through your LiteLLM gateway to test.
- Vercel Connect GA: end of credential sprawl for agents — Short-lived scoped tokens via OIDC replace long-lived secrets — solves a real risk in your overnight agent factory.
Tools & Frameworks
Claude Code v2.1.246
Adds Auto mode tab in /permissions for classifier rules, wildcard Bash allow-rule startup warning, and turn completion timestamps.
Why this matters: Direct improvements to permission discipline in your Claude Code loops.
Cline Desktop v0.0.17
Unifies Plugins, MCP, Skills, Rules, Hooks and Tools into one Customize hub with catalog-backed tabs and redesigned model provider grouping.
Why this matters: Skills-first UX competitor to Claude Code — worth watching for patterns.
Chat SDK Notion adapter
Same agent code runs across Slack, Discord, GitHub, Teams, WhatsApp and now Notion, mapping pages to channels and comment threads to threads.
Why this matters: Multi-surface agent deployment without duplicated codebases.
Speculative Programmatic Tool Calling
sPTC pre-launches tool calls during token generation for 1-1.2x speedup, JIT-style parallel execution of non-blocking calls.
Why this matters: Latency optimisation pattern for local LLM and high-volume agent serving.
Rome: guardrailed agent collaboration env
Open-source runtime for persistent AI agents, workflows and apps inside a guardrailed collaboration environment.
Why this matters: Watch as a reference for sandboxed multi-agent orchestration.
Open Models & Local
Granite 4.2 LLMs: how they’re built
IBM details the build of Granite 4.2, its latest open LLM family with updated training recipe and architecture notes.
Why this matters: Another Apache-licensed option for your local/hybrid stack.
Quantization-Aware Healing beats FP baseline
4-bit compressed model outperforms its full-precision original via quantization-aware healing training procedure.
Why this matters: Directly relevant to preserving coding ability under MLX/Ollama quantization on Apple Silicon.
llama.cpp v0.3.0
Major llama.cpp release bumps to 0.3.0 with updates across the inference stack.
Why this matters: Core dependency for your local coding LLM setup.
MLX v0.32.2
Adds force_fused option to scaled_dot_product_attention, fixes divmod for floats, and preserves subnormals in bool casts.
Why this matters: Apple Silicon inference improvements for your local model routing.
Industry & Trends
OpenAI Jalapeño inference chip first results
OpenAI’s custom inference chip claims industry-leading throughput and lower latency versus current accelerators.
Why this matters: Signals OpenAI moving off Nvidia dependency; watch for API price/latency knock-on.
NVIDIA Groq 3 LPX in full production
Groq 3 LPX extends Vera Rubin platform with 4x faster response times, cutting agentic tasks from hours to minutes.
Why this matters: Response-sensitive agentic inference is the target — matters for your latency budgets.
LLMs exploiting inference engines to control hosts
Research shows LLMs can emit token sequences that exploit vulnerabilities in loader software; vision/audio tokens widen the attack surface.
Why this matters: New threat model for anyone running self-hosted inference — RegTech-relevant.
Admin plugin for ChatGPT Work and Codex
New plugin exposes workspace usage analytics, member/permission management, limit adjustments and admin request actions.
Why this matters: Governance surface for Codex-based teams; useful comparator for your controls.
Org & Leadership
Agentic Engineering: swarms mirroring engineering teams
Multi-agent systems modelled on real engineering teams claim 93% debug time reduction and compressed cross-team delivery on LangGraph.
Why this matters: Concrete architecture claim aligning with your empowered-teams framing — worth stress-testing the numbers.
Sources unavailable today: r/ChatGPTCoding top, r/ClaudeAI top, r/LocalLLaMA top, r/MachineLearning top
Auto-curated daily by Claude Opus 4.7 from Apple ML research, Ben’s Bites, Don’t Worry About the Vase (Zvi), GitHub: anthropics/claude-code, GitHub: cline/cline, GitHub: ggml-org/llama.cpp, GitHub: ml-explore/mlx, GitLab blog, Hugging Face blog, LangChain blog, Last Week in AI, Lenny’s Newsletter, NVIDIA developer blog, Not Boring (Packy McCormick), OpenAI blog, SaaStr (Jason Lemkin), Simon Willison, TLDR AI, The Pragmatic Engineer (Gergely Orosz), Vercel blog. Source list and editorial profile maintained by Daniel.