Claude Opus 4.1 (claude-opus-4-1-20250805) model retired
Users and applications using Claude Opus 4.1 will receive errors and must migrate to Claude Opus 5. Researchers can apply for continued access through the External Researcher Access Program.
Issue #4 · Aug 9, 2026 · week of Aug 3 – Aug 9
The radar watched 63 sources across 32 tools this window and published 33 changes; these are the eleven worth your attention, ranked. One more than usual: a Supabase endpoint removal that fell between two issue windows has a September 23 deadline and should not wait another week. Every item carries the before/after evidence the pipeline captured; expand it before you take our word.
Anthropic has retired claude-opus-4-1-20250805; every request to it now returns an error, with Claude Opus 5 as the recommended replacement. Services still pinning this model ID will not degrade silently, they fail outright. Researchers who need continued access can apply through the External Researcher Access Program.
Users and applications using Claude Opus 4.1 will receive errors and must migrate to Claude Opus 5. Researchers can apply for continued access through the External Researcher Access Program.
Claude Code fixed a Bash permission bypass where a crafted command could hide parts of itself from permission checks. Teams that rely on permission modes to bound what an agent can run should upgrade now, especially for instances running unattended or in shared environments.
Resolved a security vulnerability where crafted Bash commands could evade permission checks
Supabase's logs.all Management API endpoint goes away on September 23; log querying moves to a new ClickHouse-backed endpoint that accepts ClickHouse SQL only. Scripts, integrations, and monitoring tools calling the endpoint directly need their queries converted before the cutoff. The announcement landed in late July and fell into the gap between two issue windows, so it takes its slot now, with six weeks left before the deadline.
Users querying the logs.all endpoint must migrate to the new logs endpoint and convert queries to ClickHouse SQL dialect before September 23, 2026. This affects scripts, integrations, and tooling calling this endpoint directly.
Kimi K3 is now generally available in GitHub Copilot, hosted by GitHub on Fireworks AI and billed by usage: $3 per million input tokens, $15 output, $0.30 cached input, matching the Together AI rates from last issue. It ships off by default for Business and Enterprise plans, so administrators have to enable the policy before anyone in the organization can select it.
Users will be charged usage-based pricing for Kimi K3 model, with distinct rates for input, output, and cached input tokens. Pricing documentation rollout temporarily paused due to GitHub Actions incident.
The groq.com/pricing page now carries the $650M fundraise announcement and positioning copy; the per-model token price tables are gone. Comparing costs now means digging through the console or docs, so teams with Groq in their cost model should keep their own copy of the last published rates.
Users visiting the pricing page for pricing information will no longer find the detailed pricing tables for models, tokens, and tools. Instead, they'll see marketing messaging about Groq's $650 million fundraise and company positioning. This significantly impacts developers trying to compare costs or integrate Groq's API.
OpenAI's Fast mode now takes long-context requests: prompts beyond 272K tokens on GPT-5.6 Sol, Terra, and Luna can run in the Fast tier, at up to 2.5x the speed of Standard. Workloads that fell back to Standard purely because of context length are worth re-benchmarking.
Developers using GPT-5.6 Sol, Terra, or Luna can now use Fast mode for long-context prompts exceeding 272K tokens, achieving significantly faster response times.
platform.openai.com · also seen at 1 more source
Zep added a 'Zep for Emerging Companies' tier: $13,000 for the first year, aimed at companies that have raised between $1M and $10M, with SOC 2 Type II, HIPAA BAA, DPA, guaranteed rate limits, and BYOK via AWS KMS included. Teams stuck between standard tiers that lack compliance features and an enterprise contract they cannot justify get a middle option.
Zep introduces a new mid-market pricing tier targeting companies that have raised $1M-$10M in funding, positioned between their standard tiers with enterprise features at a lower price point.
LangSmith's Managed Deep Agents entered public beta: deep agents can now run hosted on LangSmith rather than on infrastructure you operate yourself. Teams already building long-running agents in the LangChain ecosystem can weigh the managed option against self-hosting.
LangChain users can now use managed Deep Agents through LangSmith, expanding the platform's capabilities for agent orchestration.
blog.langchain.dev · also seen at 1 more source
Replicate added Nvidia H200 GPUs in single, 2x, 4x, and 8x configurations, with 144GB of memory per card at $0.001525/sec (about $5.49/hr) for a single GPU. One catch: H200 capacity requires a committed-spend contract, it is not on-demand.
Users now have access to H200 GPU instances in single, dual, quad, and 8x configurations. H200 GPUs with 144GB memory are available, supporting high-memory workloads. Capacity requires committed spend contracts.
The Gemini API released two embodied-reasoning endpoints in public preview, gemini-robotics-er-2-preview and a streaming variant, covering spatial reasoning, multi-robot coordination, and low-latency interaction. The same changelog carries an easy-to-miss deadline: gemini-robotics-er-1.6-preview shuts down on August 31, leaving projects on 1.6 the rest of this month to migrate.
Developers can now build real-time robotic agents with embodied reasoning capabilities, multi-robot coordination, and streaming support for low-latency interactions.
ai.google.dev · also seen at 1 more source
langchain-anthropic fixed caller-supplied tool_choice not being preserved on its way to the API. Before the fix, code forcing the model toward a specific tool this way was not necessarily taking effect; after upgrading, behavior returns to what callers intended, so paths that rely on it deserve a regression pass.
Tool selection behavior is now correctly preserved when passed by callers, ensuring expected tool invocation patterns.
That's it for this issue. Get the next one by email — meanwhile the directory updates daily and each issue lives here permanently.