Hy3 295B MoE, Claude Cowork, Tencent Hy3 Open
Dienstag, 7. Juli 2026 - AI News · (letzte 24h)
Tencent drops Hy3, a 295B-parameter MoE (21B active) that rivals flagship open models 2-5× its size, free on OpenRouter until July 21.
Must read
- Hy3: Tencent’s 295B MoE open model — New open flagship with 21B active params; worth benchmarking against Qwen3-Coder for your local-plus-cloud coding routing.
- State of CLI Coding Agents, Mid-2026 — Head-to-head of Claude Code, Codex CLI, Omp on real repos — directly informs your Claude Code + Cursor stack decisions.
- Getting started with loops in Claude Code — Anthropic taxonomy of agent loops maps directly to your overnight-agent-factory patterns.
- Claude Cowork background agents, Gemini Managed Agents — Anthropic Cowork on mobile/web and Google’s background execution API — both reshape headless dispatch.
- JetBrains AI for Teams: shared context and org governance — Cross-tool governance layer over Claude Code/Codex — relevant if your LiteLLM gateway is becoming the control plane.
Tools & Frameworks
Claude Code v2.1.203
Adds login-expiry warnings for background sessions, manual-permission-mode badge, and MCP roots/list notifications for additional working directories.
Why this matters: MCP roots and background-session fixes matter for your headless overnight setup.
Claude Code loops: types and stop conditions
Anthropic categorises agent loops by trigger, stop condition, primitive, and task type, with token-cost management guidance.
Why this matters: Vocabulary for verifying and staffing loop-based agents on your team.
Vercel acquires Better Auth
Vercel buys Better Auth (4.7M weekly npm downloads, 850+ contributors) to focus on open-source auth and agent identity.
Why this matters: Agent identity is the emerging control problem for your MCP servers on Vercel.
GitHub Tools SDK for eve agents
New @github-tools/sdk/eve subpath registers full GitHub toolset in nine lines with write-tool approvals on by default.
Why this matters: Safe-by-default approval gates are the pattern your leaf-node PR review needs.
PyTorch Monarch on AMD ROCm
Monarch brings elastic fault-tolerant single-controller distributed training to AMD GPUs, recovering from node failures without halting jobs.
Why this matters: Watch-but-don’t-act: matters if AMD becomes a serious training alternative.
sqlite-utils 4.0 with schema migrations
First major bump since 2020: adds schema migrations (absorbing sqlite-migrate), breaking changes documented in an upgrade guide.
Why this matters: Useful for lightweight local evals and prototype datastores alongside Postgres.
Open Models & Local
Hy3 295B MoE from Tencent
295B params, 21B active, 3.8B MTP layer; beats similarly-sized models and rivals open flagships 2-5× larger. Free on OpenRouter through July 21.
Why this matters: Test it via LiteLLM against Qwen3-Coder for coding-quality-per-token.
MLX v0.32.0
CUDA qmm fixes, cmake-generated qmm implementations, gguflib validation kept in release builds, quantised-kernel improvements.
Why this matters: Direct impact on quantised model performance on your Apple Silicon dev boxes.
Decagon: 90% of workloads on open-source models
Decagon runs ~90% of production workloads on fine-tuned open-source models for latency and task-specific performance in customer service agents.
Why this matters: Concrete data point for your three-tier architecture: when specialised OSS beats frontier.
Industry & Trends
Anthropic signs $19B TeraWulf infrastructure lease
Why this matters: Confirms Anthropic capacity ramp through 2028 — reduces near-term rate-limit risk for Claude Code shops.
xAI rebrands to SpaceXAI
Why this matters: Watch-but-don’t-act: Musk consolidating AI narrative under SpaceX.
Replit on continual learning at the harness level
Why this matters: ViBench and Telescope show how to cluster production agent failures — directly applicable to your evals.
Anthropic: J-space, a global workspace in language models
Why this matters: Interpretability progress toward monitoring internal agent reasoning — relevant to your verification problem.
Org & Leadership
Pragmatic Engineer: hiring managers & job seekers in 2026
Based on 50+ hiring managers and job seekers: AI-adjacent roles are the hottest market; engineering leadership hiring is tight.
Why this matters: Data for your London hiring plans and role right-sizing decisions.
The $10B Forward-Deployed Engineer boom
AI companies have committed $9.75B over 12 months to FDE teams; Tunguz analyses three structural models and whether FDEs create a moat.
Why this matters: Relevant if you’re weighing customer-embedded engineering as a delivery model.
Sources unavailable today: r/ChatGPTCoding top, r/ClaudeAI top, r/LocalLLaMA top, r/MachineLearning top
Auto-curated daily by Claude Opus 4.7 from Apple ML research, Ben’s Bites, Don’t Worry About the Vase (Zvi), GitHub: All-Hands-AI/OpenHands, GitHub: anthropics/claude-code, GitHub: cline/cline, GitHub: ml-explore/mlx, Hugging Face blog, JetBrains AI blog, Latent Space, Lenny’s Newsletter, NVIDIA developer blog, OpenAI blog, Simon Willison, TLDR AI, The Pragmatic Engineer (Gergely Orosz), Tomasz Tunguz, Vercel blog, smol.ai news. Source list and editorial profile maintained by Daniel.