GPT-6 Astra, Claude Proves Fermat, Meta AIRA₃
mardi 8 septembre 2026 - AI News · (24 dernières heures)
OpenAI ships GPT-6 Astra with robotic-manipulation demos, while Claude formally verifies Fermat’s Last Theorem in 11 days.
Must read
- llm 0.35 adds GPT-6 Astra — Simonw’s llm CLI now routes to gpt-6-astra; drop it into your LiteLLM gateway to A/B against Claude in Claude Code.
- Claude formalizes Fermat’s Last Theorem in Lean — 13M lines, 29,500 lemmas, 11 days — a serious datapoint for using Claude on long-horizon verification work, not just codegen.
- Meta’s AIRA₃ autonomous research engine — Async long-running agents in isolated envs — the same overnight-agent-factory pattern you write about, now at Meta scale.
- OpenAI’s 99.9% ARC-AGI-3 was the harness, not the model — Same Astra weights score 62.7% without OpenAI’s scaffolding. Reinforces your three-tier framing: the harness is the product.
- AI safety is not the same as security — Agent sandbox escapes need deterministic controls, not probabilistic alignment — directly relevant to your MCP server sandboxing.
Tools & Frameworks
LLM-as-a-Verifier framework
Training-free verifier that produces fine-grained feedback for any agent, claiming SOTA across coding, robotics, and medical benchmarks.
Why this matters: Candidate for your 22,000-line-PR verification problem.
Random Attention KV-cache eviction
Uniformly sampled KV-cache eviction matches or beats learned importance methods across reasoning benchmarks with lower overhead.
Why this matters: Cheaper long-context inference if it lands in llama.cpp/MLX.
Gemini desktop adds Ask and Assign modes
Google’s Gemini desktop is gaining Ask/Assign modes and hints at remote control features, moving toward a full desktop agent.
Why this matters: Watch as a Claude Code competitor for desktop-scoped dispatch.
Open Models & Local
Extropic Z1 probabilistic inference chip
Z1 uses probabilistic sub-threshold CMOS to cut transformer inference energy, targeting sampling-heavy workloads.
Why this matters: Watch but don’t act — no impact on what runs on your M-series today.
Industry & Trends
OpenAI targets automated AI researcher by March 2028
OpenAI internal researchers now lean heavily on coding agents; the org paused RL training after a security breach but is pushing toward full research automation.
Why this matters: Signals where frontier-lab dev workflows are heading — same direction your team is.
GPT-6 Astra tested on robotic manipulation
Astra completed the bowl-placement task in 19/20 trials but only 2/20 on precise puzzle insertion, at ~half the per-run cost of Fable.
Why this matters: Honest capability ceiling data — useful when framing what agents can/can’t own end-to-end.
Anthropic IPO slips to mid-October
Anthropic’s IPO marketing now expected mid-October with prospectus in late September, aiming to price before the US midterms.
Why this matters: Pricing pressure on Claude API tiers likely once public — plan gateway routing accordingly.
Pachocki: reasoning models may accelerate their own development
OpenAI’s chief scientist argues recursive self-improvement in reasoning models is close enough to warrant building defensive AI in parallel.
Why this matters: Context for the Astra push and why alignment is now framed as a deployment problem.
Brockman on Astra, alignment, and the Hugging Face incident
Brockman details Astra’s capability jumps, infra scaling with Nvidia/Microsoft, and OpenAI’s response to the Hugging Face security breach.
Why this matters: Long read but the cybersecurity section is directly relevant to your RegTech context.
Org & Leadership
Stripe’s enterprise AI playbook and internal agent platform
Why this matters: Named-company blueprint: shared skills platform, embedded governance, context layer — compare against your Act-2 mental model.
Sources unavailable today: Last Week in AI, r/ChatGPTCoding top, r/ClaudeAI top, r/LocalLLaMA top, r/MachineLearning top, smol.ai news
Auto-curated daily by Claude Opus 4.7 from Don’t Worry About the Vase (Zvi), Exponential View (Azeem Azhar), GitHub: simonw/llm, Import AI (Jack Clark), Latent Space, Lenny’s Newsletter, OpenAI blog, SaaStr (Jason Lemkin), Simon Willison, TLDR AI. Source list and editorial profile maintained by Daniel.