Cursor Projects, OpenAI Agents API, Claude Code Remote Control
Friday, 18 September 2026 - Weekly AI Briefing · (last 7 days)
The week the agent harness became the product. Cursor and Claude both shipped ‘Projects’ — persistent, long-horizon workspaces that delegate to swarms of agents and hold context across sessions — while OpenAI opened its managed agent runtime (the one behind Codex) as a public Agents API. For a London CTO already running an overnight-agent-factory on Claude Code, the harness is now where the margin lives: a Berkeley study this week showed identical models varying 71% in cost across harnesses with no quality delta. Meanwhile Claude Code v2.1.27x quietly added Remote Control, self-hosted runners with drain markers, and plugin evals — the dispatch primitives you’d otherwise have to build.
Launches & releases this week
Models
- Cognition SWE-2 — 50.0% on FrontierCode 1.1 at 64% lower cost than SWE-1.7; matches GPT-5.6 Sol and Fable 5.1 at a fraction of the price. (TLDR AI)
- DeepSeek v4.1-Flash — 763B-parameter causal encoder-decoder with 8B active and 16B distilled variant, adding vision — DeepSeek’s most ambitious architecture reset yet. (Latent Space)
- Gemini 3.8 Live — Two speech-to-speech models with real-time audio, visual grounding, and an extended-thinking variant; live on Vercel AI Gateway. (Google DeepMind blog)
- GPT-Live 1 — Full-duplex voice model at $0.05/min with 12 voices, client delegation, and 80% fewer interruptions than turn-based systems. (TLDR AI)
- Qwen3.8-Omni-Flash — Native omnimodal model with 1M-token context across text, image, audio, video; matches Gemini 3.8 Flash on audio-visual and beats it on audio overall. (TLDR AI)
- Bonsai 2 27B — Ternary {−1,0,+1} weights at 1.76 bits/weight, 5.9GB footprint, 262K context — runs on Apple Silicon via MLX custom low-bit kernels. (TLDR AI)
- TypeSafe Jev — System One Model returning typed Choice/Score/Boolean answers in parallel — >100x faster, >200x cheaper than small frontier LLMs, cannot hallucinate. (TLDR AI)
Features & Tools
- Cursor Projects — Persistent workspaces that delegate to thousands of agents, hold context across months, and run recurring work unprompted; internal users merge 30% more PRs. (TLDR AI)
- Claude Projects — Claude Code Projects manages parallel cloud sessions with shared memory and auto-delegation; in beta for select subscribers. (TLDR AI)
- Google Home MCP — Early-access MCP server lets Claude, ChatGPT, or any MCP client control every Google Home device via a Google Cloud project. (TLDR AI)
Products
- Astra for Law — GPT-6 Astra packaged with legal-grade privacy controls, connected legal data sources, and custom firm workflows. (OpenAI blog)
- Sourcegraph Agentic Batch Changes — One engineer drives codebase-wide changes across hundreds of repos; priced on outcomes — you only pay for merged changesets. (Sourcegraph blog)
- GitLab Duo CLI + MCP — Goal-to-done CLI plus new MCP tools for scoped pipeline/MR automation, hosted Kimi K3, GLM 5.3, and MiniMax M3 model choice. (GitLab blog)
Other Releases
- OpenAI Agents API — Public beta of the managed agent harness behind Codex — context, tools, subagents, persistent execution, files, and code environments as a hosted API. (TLDR AI)
- Claude Code v2.1.27x — Adds Remote Control sessions, self-hosted runners with drain markers, plugin eval CLI, gateway hint headers, and OTel spans for compaction and managed settings. (GitHub: anthropics/claude-code)
Stories to follow
The harness is the product
Every frontier lab shipped managed agent infrastructure this week: OpenAI opened Codex’s runtime as the Agents API, Cursor and Claude both launched persistent Projects, Google put Agent Substrate on GKE at 10x container density. Independent research made the case sharper — Arena AI’s HarnessTax study and a Berkeley paper both found harness choice barely affects quality but drives cost swings up to 71%. Josh Rosen’s ‘Managed Agent Architectures’ post frames the strategic choice: outsource generic loop plumbing, own product-specific logic.
- OpenAI Agents API — Codex’s managed harness — context, tools, subagents, persistent execution — now a public-beta API. (TLDR AI)
- Cursor Projects — Long-horizon workspaces with thousands of delegated agents; 30% more merged PRs internally. (TLDR AI)
- HarnessTax — 21 model-harness pairs: harness barely affects success rate but massively affects cost. (TLDR AI)
- The Harness Margin Opportunity — Same model, same output: 71% cost delta between harnesses drives 38% vs 75% gross margins. (Tomasz Tunguz)
- Managed Agent Architectures — Frontier labs are turning the agent loop into managed infrastructure; builders must pick what to own. (TLDR AI)
- Agent Substrate on GKE — Runtime for millions of sandboxes at 10x density with sub-500ms suspend/resume. (TLDR AI)
Ship-100-a-day org shapes
Real numbers from AI-native orgs landed this week. Delphi ships 100+ production deploys/day with 10 engineers and no dedicated infra role. Vercel cut its inbound SDR team from 10 to 1.25 with 90% automation and a single-digit-thousands infra bill. ICONIQ’s new pacesetter benchmark: 115% growth at $100M+, $655K revenue per employee. This is what the GitLab Act-2 blueprint looks like in numbers — smaller teams, end-to-end ownership, if-an-agent-can-do-it-automate-it as an operating principle rather than a slogan.
- Delphi ships 100/day — 10 engineers, no infra role, everyone including product and design ships behind feature flags. (Vercel blog)
- Vercel’s 10→1.25 SDR team — 90% inbound automation, 32x ROI on the SDR agent, single-digit-thousand-dollar total infra spend. (Tomasz Tunguz)
- ICONIQ Pacesetter Index — New bar: 115% growth at $100M ARR, 55% gross margins, $655K revenue per employee. (SaaStr (Jason Lemkin))
- OpenAI’s agentic software factory — How Codex ‘took over’ OpenAI internally and reshapes engineering for one billion users. (The Pragmatic Engineer (Gergely Orosz))
Local + open-weight closes the gap
Open-weight models hit 56% of Vercel AI Gateway token volume this month. On the local side: Bonsai 2 27B ships 27B parameters in 5.9GB with MLX kernels tuned for Apple Silicon, Ollama 0.34 makes MLX safetensors non-experimental, and Cline Desktop launches specifically for running open-weight models with parallel agents. Together AI added a five-stage playbook for closed→open migration. The hybrid routing question — which turn goes local, which goes to Astra — is now a first-class product decision.
- Open-weight hits 56% of tokens — AI Gateway Production Index: open-weight models now serve the majority of production token volume. (Vercel blog)
- Bonsai 2 27B — Ternary weights, 5.9GB, 262K context, custom MLX kernels for Apple Silicon. (TLDR AI)
- Ollama 0.34.2 — MLX safetensors ollama create now non-experimental; better Apple Silicon memory handling. (GitHub: ollama/ollama)
- Cline Desktop — Open-source app purpose-built for running open-weight models with parallel agents. (TLDR AI)
- Closed→open migration playbook — Together’s five-stage playbook: discover, evaluate, adapt, decide, production. (Together AI blog)
Skills as the discipline layer
The Notion Skills API lets teams edit agent instructions collaboratively and distribute them via GitHub sync or Vercel’s skills CLI — which now supports Notion-hosted skills natively. Cline’s SDK moved to hub-managed agent plugins under ~/.agents/plugins with explicit safeguards against repo-controlled MCP servers auto-starting. Matt Pocock’s Pragmatic Engineer interview on AI skills fills in the human side. This is the layer above vibe coding starting to consolidate on standards.
- Notion Skills API — Teams edit agent instructions together, distribute via GitHub sync or Vercel skills installer. (TLDR AI)
- Skills CLI supports Notion — skills@1.7.0 adds Notion databases as an install source — no Git repo required. (Vercel blog)
- Cline hub-managed plugins — Agent plugins under ~/.agents/plugins with skills and MCP server discovery; workspace scan deliberately disabled. (GitHub: cline/cline)
- AI Skills with Matt Pocock — How to plan and build software with agents; why engineering fundamentals matter more. (The Pragmatic Engineer (Gergely Orosz))
What I’m watching
- Frontier pacing debate goes political — Dario’s ‘pace the frontier’ essay triggered a preference cascade this week. Not directly actionable, but the regulatory shape it produces will constrain what capabilities your team can use.
- We Must Pace the Frontier (TLDR AI)
- Coxon warns of extinction (Don’t Worry About the Vase (Zvi))
- What does pacing mean? (TLDR AI)
- Evals you can trust — Multiple pieces this week confirmed labs are training against their guardrails and physics benchmarks were quietly broken. Your internal eval discipline matters more, not less.
- AI cheating is on the rise (TLDR AI)
- Bad benchmarks and evals (TLDR AI)
- Real-SWE benchmark (TLDR AI)
- Agents paying for content — x402 and pay-per-crawl are moving from spec to real transactions between Claude and websites. Relevant if your product publishes or consumes content agents want.
- Claude paid to read my page (TLDR AI)
- Is Agentic report categories (Vercel blog)
Top trending GitHub repos this week
browser-use/jev-ultrafast
4.7k★ · Python i. am. speed.
tamaratran/fast-jev-compaction
2.7k★ · TypeScript Claude Code plugin that replaces the compaction summary with Jev decisions: every tool call and result is scored in one fast request, stale ones are dropped or truncated, everything kept stays verbatim.
TheoLeeCJ/SemIf
1.5k★ · Python Semantic ifs from open models, on a 3090 at home. Independent; not affiliated with Jev or TypeSafe.
Chuloo/mural
1.4k★ · Kotlin · ios language-learning open-source swiftui
The language app you eventually delete. A native iPhone companion for learning through conversation.
yifanzhang-pro/recurrent-looped-tranformer
878★ · HTML Official Project Page for Recurrent Looped Transformer (RLT)
Read this weekend
Inside OpenAI’s agentic software factory
Gergely’s deepdive on how Codex ‘took over’ OpenAI internally is the most useful long-form of the week for someone running an agent-augmented team. Concrete engineering choices for one billion users, not vibes.
Quote of the week
Production code written by Claude should have a higher bar than if it was written by a human.
— Boris Cherny, Anthropic · link
Sources unavailable this week: Last Week in AI, r/ChatGPTCoding top, r/ClaudeAI top, r/LocalLLaMA top, r/MachineLearning top
Auto-curated weekly by Claude Opus 4.7 from Apple ML research, Ben’s Bites, Don’t Worry About the Vase (Zvi), Exponential View (Azeem Azhar), GitHub: All-Hands-AI/OpenHands, GitHub: BerriAI/litellm, GitHub: anthropics/claude-code, GitHub: cline/cline, GitHub: crewAIInc/crewAI, GitHub: langchain-ai/langchain, GitHub: ollama/ollama, GitLab blog, Google DeepMind blog, Hugging Face blog, Interconnects (Nathan Lambert), JetBrains AI blog, LangChain blog, Latent Space, Lenny’s Newsletter, NVIDIA developer blog, Not Boring (Packy McCormick), OpenAI blog, SaaStr (Jason Lemkin), Simon Willison, Sourcegraph blog, TLDR AI, The Algorithmic Bridge (Alberto Romero), The Pragmatic Engineer (Gergely Orosz), Together AI blog, Tomasz Tunguz, Understanding AI (Timothy B. Lee), Vercel blog. Source list and editorial profile maintained by Daniel.