Tencent Hy3 295B, Claude Code 2.1.202, GPT-5.6 Preview
Montag, 6. Juli 2026 - AI News · (letzte 24h)
Tencent open-weights Hy3, a 295B MoE with 21B active params, 256K context and vLLM MTP support — a credible GLM-5.2 rival.
Must read
- Tencent releases Hy3 (295B MoE, Apache 2.0) — New open-weight frontier-class MoE with 256K context and MTP; worth benchmarking against Qwen3-Coder in your local-plus-cloud routing.
- Claude Code v2.1.202 — Dynamic workflow size control and OTel workflow.run_id/name attrs — directly useful for observing your overnight agent factory.
- OpenAI GPT-5.6 in narrow preview — Three tiers (Sol/Terra/Luna) plus reasoning-effort slider; plan LiteLLM routing before it lands in Codex.
- Closing the Verification Loop — Directly maps to your 22k-line PR problem: /ce-dogfood skill and persona strategy for verifying agent output at scale.
- jamesob’s guide to running SOTA LLMs locally — Concrete hardware/config recipes from $2k to $40k for Apple Silicon-adjacent local setups; feeds your published local-LLM writing.
Tools & Frameworks
Ollama v0.31.2
Flash attention on compute-capability 6.x NVIDIA GPUs, iGPU vision offload with padding, updated MLX and llama.cpp engines, ollama launch for Claude Code.
Why this matters: Closer Ollama/Claude Code integration matters for your local-hybrid workflow.
OpenHands cloud 1.41.0
Adds tree-sitter AST semantic file chunking, organization conversation admin dashboard, and production workspace state snapshots on start/archive.
Why this matters: Agent-harness competitor to Claude Code worth tracking for team-scale governance.
Own the Loop: A Field Guide to Agent Harnesses
Argues that as coding models commoditise, the durable edge shifts to model-agnostic harnesses managing tools, orchestration and routing.
Why this matters: Reinforces your case for owning the control loop over vendor lock-in.
Running autonomous coding agents from a phone (Symphony + Linear)
Alessio Fanelli demos a Symphony + Linear setup dispatching parallel Codex agents from mobile against a real Linear backlog.
Why this matters: Concrete dispatch pattern adjacent to your overnight-agent-factory thinking.
JetBrains tests the Caveman token-compression skill
Paired A/B on SkillsBench: Caveman’s advertised 65% token saving measures at 8.5% on real agentic tasks with the skill forcibly activated.
Why this matters: Sober empirical check on Claude Code skills before you standardise any across the team.
Open Models & Local
Mistral releases Leanstral
Open-source 119B-parameter theorem-proving and code-verification agent built on Mistral’s coding framework.
Why this matters: Verification-focused model relevant to your three-tier architecture for regulated identity/fraud work.
Open Source AI Gap Map
Interactive map of the open-source AI stack highlighting fragmentation, duplication and missing layers across model, tooling and infra tiers.
Why this matters: Useful reference when picking OSS components for the in-house stack.
Industry & Trends
Alibaba reportedly bans Claude Code
Alibaba plans to prohibit employee use of Claude Code from July 10 as high-risk software, redirecting engineers to its in-house Qoder tool.
Why this matters: Signal on geopolitical fragmentation of coding-agent access — watch, don’t act.
The end of compute scarcity? Not so fast
Meta and SpaceX reportedly reselling compute capacity, hinting at hyperscaler capex revisions though resold capacity clears immediately.
Why this matters: Watch item — could soften cloud inference pricing for your gateway.
Import AI 464: Fable writes GPU kernels
Jack Clark covers Fable-generated GPU kernels, automation trends, and analog computation directions in this week’s roundup.
Why this matters: Primary-source-adjacent context on where coding-model capability is heading.
Sources unavailable today: r/ChatGPTCoding top, r/ClaudeAI top, r/LocalLLaMA top, r/MachineLearning top
Auto-curated daily by Claude Opus 4.7 from Exponential View (Azeem Azhar), GitHub: All-Hands-AI/OpenHands, GitHub: anthropics/claude-code, GitHub: langchain-ai/langgraph, GitHub: ollama/ollama, GitLab blog, Hugging Face blog, Import AI (Jack Clark), JetBrains AI blog, Lenny’s Newsletter, NVIDIA developer blog, Simon Willison, TLDR AI, The Algorithmic Bridge (Alberto Romero), Tomasz Tunguz, smol.ai news. Source list and editorial profile maintained by Daniel.