Claude Code AGENTS.md, Claude Projects Redesigned, Qwen3.8-Omni-Flash
Saturday, 19 September 2026 - AI News · (last 24h)
Claude Code adopts the AGENTS.md standard in v2.1.277, while Anthropic redesigns Projects around parallel cloud sessions with shared memory.
Must read
- Claude Code adds AGENTS.md support (v2.1.277) — Your in-house Claude Code setup can now share project instructions with Cursor and Codex — one file, all harnesses.
- Claude Projects redesigned around parallel cloud sessions — Directly maps to your overnight-agent-factory pattern: threaded parallel work, shared memory, task delegation baked in.
- Qwen3.8-Omni-Flash: 1M-context omnimodal, beats Gemini 3.8 Flash on audio — Frontier-class omnimodal from the Qwen line you already track — worth routing tests via LiteLLM.
- Bonsai 2 27B: ternary weights, 5.9GB footprint, 262K context — 27B-class reasoning/coding model that fits comfortably on Apple Silicon via MLX — direct local-tier candidate.
- Anthropic: Claude now leads 26% of its research work — Concrete numbers on internal agent oversight at scale — useful reference when framing your own agentic adoption.
Tools & Frameworks
Claude Code v2.1.278: server-side auto-mode classifier
Auto mode now defaults to a server-side classifier that doesn’t bill classifier overhead on API/Enterprise/Bedrock/Vertex/gateways; opt-out flag provided.
Why this matters: Reduces token spend on your LiteLLM-routed Claude Code usage.
WebMCP support lands in Vercel’s mcp-handler
mcp-handler now exposes existing MCP tools to in-browser agents via a single script tag using the proposed WebMCP standard.
Why this matters: Path to reuse your in-house MCP servers inside browser-side agents on Vercel.
Notion Skills API with GitHub sync and Vercel skills installer
Teams edit agent instructions collaboratively in Notion and distribute them to other agents via GitHub sync or Vercel’s skills installer.
Why this matters: Skills-framework tooling for the discipline layer above vibe coding you write about.
GLM 5.3 FlashX on Vercel AI Gateway at ~200 tok/s
Z.ai’s multimodal coding model now served at ~200 tokens/sec via AI Gateway under zai/glm-5.3-flashx.
Why this matters: Fast cheap coding model to route through your LiteLLM gateway for tool loops.
Agora: Git as shared memory for AI research agents
Append-only Git DAG where agents commit hypotheses, results, verifications, and reports as reproducible artefacts.
Why this matters: Interesting primitive for your overnight-agent-factory audit trail problem.
SGLang v0.5.20: 713 PRs, adds GLM-5.3-Flash and Hy4-Preview
Major SGLang release with new model support including GLM-5.3-Flash across 237 contributors.
Why this matters: Serving-layer option if you self-host coding models alongside gateway calls.
Open Models & Local
Bonsai 2 27B ternary model, MLX-ready
27B ternary-weight model at 1.76 effective bits/weight, 5.9GB footprint, 262K context, multimodal, with custom MLX kernels for Apple Silicon.
Why this matters: Rare frontier-adjacent model that actually fits your M-series local tier.
Qwen3.8-Omni-Flash
Native omnimodal model with 1M-token context; audio-visual performance near Gemini 3.8 Flash and overall audio exceeding it.
Why this matters: Qwen line remains the strongest open bet for your local-plus-cloud routing.
Industry & Trends
Anthropic measures its own agent-led research
Claude leads 26% of Anthropic’s AI research work, overseeing tens of thousands of internal agents; new metrics track oversight and compute.
Why this matters: Benchmark data for the leaf-node verification problem you write about.
GLM built its own inference stack with an Infra Agent
Z.ai used a GLM-5.3-powered agent to build GLM-5.3-Flash’s serving stack on 100K+ accelerators in under two weeks, tripling throughput.
Why this matters: Concrete recursive-self-improvement case study with real numbers.
Gemini autonomously breached three companies in Irregular test
Google confirms Gemini completed unassisted breakout hacks against three companies in a May red-team run by Irregular.
Why this matters: Directly relevant to your identity/fraud domain and MCP sandboxing patterns.
Global bank scales coding agents on Together DMI
A global fintech moved coding agent traffic onto Together’s Dedicated Model Inference for self-serve scaling, model choice, and testing control.
Why this matters: Adoption pattern from a regulated-sector peer for hosted coding-agent inference.
Org & Leadership
GitLab CISO: securing the software factory at machine speed
New GitLab CISO frames agentic software delivery as a trust-scarcity problem, extending the Act 2 restructure into embedded security and governance.
Why this matters: The security half of the GitLab Act 2 blueprint you track — governance as the moat.
Sources unavailable today: Last Week in AI, The Gradient, r/ChatGPTCoding top, r/ClaudeAI top, r/LocalLLaMA top, r/MachineLearning top
Auto-curated daily by Claude Opus 4.7 from Apple ML research, Don’t Worry About the Vase (Zvi), GitHub: anthropics/claude-code, GitHub: langchain-ai/langchain, GitHub: sgl-project/sglang, GitLab blog, Latent Space, NVIDIA developer blog, Not Boring (Packy McCormick), One Useful Thing (Ethan Mollick), OpenAI blog, SaaStr (Jason Lemkin), Simon Willison, TLDR AI, Together AI blog, Vercel blog. Source list and editorial profile maintained by Daniel.