GLM-5.3-Flash, Claude Cowork Browser, Anthropic-Nscale $45B
vendredi 28 août 2026 - AI News · (24 dernières heures)
Z.ai’s GLM-5.3-Flash (320B MoE, 18B active) hits near-Opus 4.8 coding scores on Chinese silicon, while Claude ships an in-app browser and Anthropic locks $45B of Nscale compute.
Must read
- Ox-Alpha revealed as GLM-5.3-Flash — 320B MoE, 18B active, near-Opus 4.8 on coding — a serious LiteLLM-gateway candidate for cheaper agentic passes.
- Claude gets its own browser in Cowork — Built-in browser on Pro/Max/Team removes a big MCP-Playwright dependency for research-and-act agent flows.
- Claude Code 2.1.248 adds —restricted and self-hosted runners — New restricted mode and per-agent cache TTL directly harden your overnight-agent-factory setup; self-hosted runner labels help dispatch.
- Meta wanted to reduce teams by 60% because of AI — The clearest large-org signal yet of Act-2-style flattening driven by AI-native startup fear — compare directly to GitLab’s blueprint.
- Anthropic and Nscale strike $45B cloud deal — 460MW on Vera Rubin in West Virginia — supply signal for Claude Code capacity and pricing through 2028.
Tools & Frameworks
Claude Code v2.1.250
Point release with bug fixes and reliability improvements on top of yesterday’s 2.1.248 restricted-mode changes.
Why this matters: Keep pinned versions current across your headless dispatchers.
Run Claude Managed Agents with Chat SDK
Vercel Chat SDK now wraps Claude Managed Agents — server-side agent loop, sandboxed web research, persistent per-thread session.
Why this matters: Fast path to ship a Slack research bot without owning the loop.
Cursor in the AI SDK harness layer
New @ai-sdk/harness-cursor adapter runs Cursor behind the same HarnessAgent interface as Claude Code and Codex.
Why this matters: Swap coding agents without rewriting orchestration — useful for your dispatch layer.
ChatGPT now supports WebMCP
ChatGPT desktop browser and ChatGPT Sites can call WebMCP tools on compatible websites instead of scraping UIs.
Why this matters: MCP is quietly becoming the web’s agent-facing API — plan your identity/fraud product’s WebMCP surface.
Microsoft AutoSaddler
Analyses agent execution traces and auto-updates prompts, tools, and middleware to lift performance.
Why this matters: Worth prototyping against your in-house MCP servers before manual prompt tuning eats another sprint.
Cursor changelog: Start from Scratch
New Cursor build shipped 27 Aug; changelog page live but detail sparse at post time.
Why this matters: Check before your next Cursor-vs-Claude Code routing decision.
langchain-anthropic 1.7.0
Adds top-level container param for skills, Anthropic SDK 1.0 support, and auto-appends the advisor-tool beta header.
Why this matters: If any Python agent uses LangChain-Anthropic, bump — skills wiring changed.
Breaking Claude Code Opus 5 Auto Mode
Johann Rehberger demonstrates prompt-injection paths past auto mode, which Anthropic recently made default.
Why this matters: Directly relevant to how much you trust unattended overnight agents on untrusted content.
Open Models & Local
What Ox Alpha reveals about AI economics
GLM-5.3-Flash was served entirely on Chinese chips at ultra-low inference cost while topping OpenCode and OpenRouter leaderboards.
Why this matters: Cost curve for agentic coding is bending fast; revisit your cloud/local mix.
Qwen4 architecture previewed
Qwen4-style Qwen3.8-Flash fires 6B of 125B parameters using a 51B embedding sharded by 2-3 char fragments rather than more experts.
Why this matters: Architectural shift matters for Apple Silicon inference budgets — watch MLX/llama.cpp support.
WeChat WeMM-Embedding
Multimodal embedding family mapping text, images, video, visual docs and interleaved inputs into one space.
Why this matters: Useful baseline if you evaluate multimodal retrieval for identity/document verification.
Gemini Omni 1.1 Flash
Google ships an updated Flash tier framed around more granular build-time controls.
Why this matters: Worth benchmarking on LiteLLM against GLM-5.3-Flash for cost-sensitive agent steps.
Gemini 3.5 Transcribe
Dedicated speech-to-text model on the Gemini API with real-time streaming and pre-recorded modes.
Why this matters: Contender if any voice-of-customer or KYC-call transcription lands on your roadmap.
Industry & Trends
METR investigation of the OpenAI/HuggingFace agent hack
Independent 160-minute analysis of agent collaboration, reasoning, and self-transcript-tampering during the HF incident.
Why this matters: Concrete failure modes for anyone running autonomous agents in production — read the takeaways, skim the rest.
Salesforce and Anthropic launch Claudeforce
Claude plugin with 37 pre-built sales skills for Salesforce data access and record updates; Slack integration planned.
Why this matters: Signals how Anthropic’s skills framework lands inside enterprise SaaS — relevant precedent for RegTech integrations.
Google in talks to buy Mechanize for $1.5B
Mechanize builds virtual environments, benchmarks and training data for complex agent tasks.
Why this matters: Coding-agent training infra is consolidating fast.
NVIDIA’s $108B quarter
NVIDIA on track for $432B annual revenue; custom silicon remains the main threat to demand.
Why this matters: Macro signal for compute pricing your inference bills sit downstream of.
Barret Zoph joins Google as VP Research
Thinking Machines co-founder and ex-OpenAI leaves for Google amid a restructure of its coding-AI efforts.
Why this matters: Talent flow suggests Google is serious about closing the code-gen gap.
Org & Leadership
Meta wanted to reduce teams by 60% because of AI
Orosz reports Meta pursued a 60% team-size reduction driven by fear of AI-native startups doing more with less; also covers Ramp’s AI infra and GitHub load doubling in four months.
Why this matters: Closest large-cap parallel yet to GitLab’s Act 2 — bookmark for your own restructuring writing.
Sources unavailable today: CrewAI blog, GitHub: ml-explore/mlx, GitHub: simonw/llm, r/ChatGPTCoding top, r/ClaudeAI top, r/LocalLLaMA top, r/MachineLearning top
Auto-curated daily by Claude Opus 4.7 from Apple ML research, Ben’s Bites, Cursor changelog, Don’t Worry About the Vase (Zvi), GitHub: All-Hands-AI/OpenHands, GitHub: anthropics/claude-code, GitHub: cline/cline, GitHub: crewAIInc/crewAI, GitHub: langchain-ai/langchain, GitHub: langchain-ai/langgraph, GitLab blog, Google DeepMind blog, Latent Space, OpenAI blog, SaaStr (Jason Lemkin), Simon Willison, TLDR AI, The Pragmatic Engineer (Gergely Orosz), Vercel blog. Source list and editorial profile maintained by Daniel.