Skip to content

← AI Tracker

AI Briefing

Claude Opus 5.5, GPT-6 Sol & Luna, Cursor Security Reviewer

Thursday, 24 September 2026 - AI News · (last 24h)

Anthropic ships Opus 5.5 at 40% lower cost while OpenAI counters with cheaper GPT-6 Sol and Luna variants — the mid-market price war is on.

Must read

  • Claude Opus 5.5 — New default for your Claude Code workflows: matches Fable 5.1 quality at 40% lower Opus 5 cost.
  • GPT-6 Sol and Luna — Cheaper GPT-6 tier for coding and computer-use — worth routing tests via your LiteLLM gateway.
  • Cursor: Rollouts and Security Reviewer — Security review inside Cursor is directly relevant to identity/fraud-side code paths your team ships.
  • What a task costs on Opus 5.5 — Concrete cost-per-turn breakdown for cache-heavy sessions — informs your overnight-agent-factory budget.
  • SWE-Bench Pro V2 — Frontier models sit at ~23% on 642 realistic tasks — a sobering reference when you evaluate agents on your own repos.

Tools & Frameworks

Claude Code v2.1.281

Adds Bedrock assume-role via STS, per-developer sessions, Bedrock guardrail bindings, and desktop policy blocks for working-directory reads.

Why this matters: Cleaner enterprise auth path if you move Claude Code onto AWS Bedrock.

Cursor rollouts and Security Reviewer

Cursor ships staged feature rollouts and a Security Reviewer for agent-produced diffs.

Why this matters: Directly addresses the 22k-line-PR verification problem you keep writing about.

Better GPT-6 prompt caching

Higher default cache hit rates, discounted shared-prefix reuse within 30 minutes, plus new cache diagnostics.

Why this matters: Cuts token spend on repeated agent scaffolds routed through LiteLLM.

Vercel Connect supports TanStack AI

TanStack AI agents can now call OAuth-protected MCP servers via Vercel Connect with no stored credentials — token refreshed per request.

Why this matters: Removes MCP credential-rotation pain for your Vercel-hosted surfaces.

LiteLLM v1.101.2

Latest LiteLLM gateway release, cosign-signed Docker images with pinned commit hash verification.

Why this matters: Your model gateway — pick up Opus 5.5 and GPT-6 Sol/Luna routing here.

langchain-anthropic 1.7.4

Adds Opus 5.5 and GPT-6 profile augmentations plus mid-conversation tool changes on SystemMessage.

Why this matters: Mid-run tool swaps unblock dynamic agent skill loading.

Open Models & Local

Ollama v0.34.4

Single-pass structured outputs for thinking models, faster Qwen 3.8 prompt processing on Apple Silicon, per-image resolution selection for Gemma 4.

Why this matters: Direct upgrade for your local Apple Silicon coding stack.

llama.cpp v0.5.0

Adds HRM-Text, MiMo-V2.6 and HunyuanOCR conversion, Metal Apple Silicon backend improvements, and multi-address HTTP binding.

Why this matters: MiMo-V2.6 support widens the coding-model bench you can run locally.

Cline CLI v3.0.65

Compact-and-retry on output-token cutoffs for llama.cpp/Ollama/LM Studio so long local sessions don’t die mid-answer.

Why this matters: Fixes a real failure mode for overnight agent runs on local models.

Train your own Jev classifier for $17

Together fine-tunes a 4B Qwen3.5 classifier on serverless for roughly $17 with a reproducible recipe.

Why this matters: Cheap path to a bespoke fraud/identity classifier in the deterministic-plus-ML tier.

Hardware-agnostic models in vLLM

New abstraction layers hit 96.6% of native H100 performance while staying torch-compilable across accelerators.

Why this matters: Watch item if you’re planning any self-hosted inference beyond Apple Silicon.

Opus 5.5 becomes the new default; everyone cuts prices 40-50%

Frontier providers dropped prices 40-50% around the Opus 5.5 launch, with OpenAI’s efficient GPT-6 tier following.

Why this matters: Rebudget your per-agent token spend — the ceiling just moved.

The most important AI market is the middle

Fable 5.1 took only 3.7% of gateway spend in 12 days; open models now serve the majority of token volume at ~86% discount.

Why this matters: Reinforces the case for mid-tier routing via LiteLLM rather than defaulting to frontier.

Training AI from real-world tool use

Perplexity combines rejection-sampling fine-tuning with hint-guided self-distillation on successful sessions and user-corrected failures.

Why this matters: Concrete recipe for turning your agent’s own traces into training signal.

Google’s RRSI for self-improving agent harnesses

RRSI regularises recursive self-improvement of agent harnesses; improves out-of-distribution results across eight benchmarks with fewer policy tokens.

Why this matters: Relevant if you’re letting agents rewrite their own scaffolding overnight.


Sources unavailable today: Last Week in AI, The Gradient, r/ChatGPTCoding top, r/ClaudeAI top, r/LocalLLaMA top, r/MachineLearning top

Auto-curated daily by Claude Opus 4.7 from Apple ML research, Cursor changelog, Don’t Worry About the Vase (Zvi), GitHub: All-Hands-AI/OpenHands, GitHub: BerriAI/litellm, GitHub: anthropics/claude-code, GitHub: cline/cline, GitHub: ggml-org/llama.cpp, GitHub: langchain-ai/langchain, GitHub: langchain-ai/langgraph, GitHub: ollama/ollama, GitLab blog, Google DeepMind blog, Hugging Face blog, Latent Space, NVIDIA developer blog, Not Boring (Packy McCormick), OpenAI blog, SaaStr (Jason Lemkin), Simon Willison, TLDR AI, The Pragmatic Engineer (Gergely Orosz), Together AI blog, Tomasz Tunguz, Vercel blog. Source list and editorial profile maintained by Daniel.