AI Devtools Radar

Issue #4 · Aug 9, 2026 · week of Aug 3 – Aug 9

Anthropic retires Claude Opus 4.1, Claude Code fixes a Bash permission bypass, Groq pulls its public price list

The radar watched 63 sources across 32 tools this window and published 33 changes; these are the eleven worth your attention, ranked. One more than usual: a Supabase endpoint removal that fell between two issue windows has a September 23 deadline and should not wait another week. Every item carries the before/after evidence the pipeline captured; expand it before you take our word.

  1. Anthropic has retired claude-opus-4-1-20250805; every request to it now returns an error, with Claude Opus 5 as the recommended replacement. Services still pinning this model ID will not degrade silently, they fail outright. Researchers who need continued access can apply through the External Researcher Access Program.

    deprecationhigh

    Claude Opus 4.1 (claude-opus-4-1-20250805) model retired

    Users and applications using Claude Opus 4.1 will receive errors and must migrate to Claude Opus 5. Researchers can apply for continued access through the External Researcher Access Program.

    Evidence
    +We've retired the Claude Opus 4.1 model (claude-opus-4-1-20250805). All requests to this model will now return an error. We recommend upgrading to Claude Opus 5. Researchers can request ongoing access through the External Researcher Access Program.

    docs.claude.com

  2. Claude Code fixed a Bash permission bypass where a crafted command could hide parts of itself from permission checks. Teams that rely on permission modes to bound what an agent can run should upgrade now, especially for instances running unattended or in shared environments.

    otherhigh

    Fixed Bash permission bypass vulnerability via crafted commands

    Resolved a security vulnerability where crafted Bash commands could evade permission checks

    Evidence
    +Fixed a Bash permission bypass where a crafted command could hide parts of itself from permission checks

    github.com

  3. Supabase's logs.all Management API endpoint goes away on September 23; log querying moves to a new ClickHouse-backed endpoint that accepts ClickHouse SQL only. Scripts, integrations, and monitoring tools calling the endpoint directly need their queries converted before the cutoff. The announcement landed in late July and fell into the gap between two issue windows, so it takes its slot now, with six weeks left before the deadline.

    apihigh

    logs.all Management API endpoint being deprecated and replaced with new ClickHouse-backed logs endpoint

    Users querying the logs.all endpoint must migrate to the new logs endpoint and convert queries to ClickHouse SQL dialect before September 23, 2026. This affects scripts, integrations, and tooling calling this endpoint directly.

    Evidence
    +The logs.all Management API endpoint is being removed on 23rd September 2026 (Wednesday), two months from this announcement. Log querying moves to a new ClickHouse-backed logs endpoint. The new endpoint accepts ClickHouse SQL only.

    supabase.com

  4. Kimi K3 is now generally available in GitHub Copilot, hosted by GitHub on Fireworks AI and billed by usage: $3 per million input tokens, $15 output, $0.30 cached input, matching the Together AI rates from last issue. It ships off by default for Business and Enterprise plans, so administrators have to enable the policy before anyone in the organization can select it.

    pricinghigh

    Kimi K3 pricing announced: $3 per 1M input tokens, $15 per 1M output tokens, $0.30 per 1M cached input tokens

    Users will be charged usage-based pricing for Kimi K3 model, with distinct rates for input, output, and cached input tokens. Pricing documentation rollout temporarily paused due to GitHub Actions incident.

    Evidence
    +Kimi K3 pricing, which will be $3 per 1M input tokens, $15 per 1M output tokens, and $0.30 per 1M cached input tokens.

    github.blog

  5. The groq.com/pricing page now carries the $650M fundraise announcement and positioning copy; the per-model token price tables are gone. Comparing costs now means digging through the console or docs, so teams with Groq in their cost model should keep their own copy of the last published rates.

    otherhigh

    Groq pricing page completely redesigned to focus on company positioning and fundraise announcement rather than pricing details

    Users visiting the pricing page for pricing information will no longer find the detailed pricing tables for models, tokens, and tools. Instead, they'll see marketing messaging about Groq's $650 million fundraise and company positioning. This significantly impacts developers trying to compare costs or integrate Groq's API.

    Evidence
    Smart, Fast, and Affordable Unmatched Price Performance Fast responses, scalable performance, and costs you can plan for. Start Building Large Language Models *Approximate number of tokens per $ AI Model Current Speed(Tokens per Second) Input Token Price(Per Million Tokens) Output Token Price(Per Million Tokens) [Detailed pricing tables for GPT OSS 20B 128k, GPT OSS Safeguard 20B, GPT OSS 120B 128k, Llama 3.3 70B Versatile 128k, Llama 3.1 8B Instant 128k, Qwen 3.6 27B 131k, and other models with specific prices]
    +Announcing our $650 million fundraise to scale global inference Read more> Every customer served. Every product sold. Every commit merged. Every agent task completed. That's inference. Training creates the possibility. Inference creates the value. As AI does more, inference multiplies. And inference is becoming the bottleneck. Groq was built for this. We pioneered the LPU. Now, with LPX, it works alongside NVIDIA's next-generation GPUs to deliver unparalleled inference capability, reliably, affordably, at scale. Fast or affordable is no longer a tradeoff. AI keeps training. Now it needs a better way to work. Groq makes inference work at scale.

    groq.com

  6. OpenAI's Fast mode now takes long-context requests: prompts beyond 272K tokens on GPT-5.6 Sol, Terra, and Luna can run in the Fast tier, at up to 2.5x the speed of Standard. Workloads that fell back to Standard purely because of context length are worth re-benchmarking.

    featurehigh

    Fast mode now supports long-context requests (272K+ tokens) for GPT-5.6 models with 2.5× speed improvement

    Developers using GPT-5.6 Sol, Terra, or Luna can now use Fast mode for long-context prompts exceeding 272K tokens, achieving significantly faster response times.

    Evidence
    +Fast mode now supports long-context requests for GPT-5.6 Sol, GPT-5.6 Terra, and GPT-5.6 Luna. As of today, long-context prompts exceeding 272K tokens can run in Fast mode, delivering speeds up to 2.5× faster than the Standard tier.

    platform.openai.com · also seen at 1 more source

  7. Zep

    Zep added a 'Zep for Emerging Companies' tier: $13,000 for the first year, aimed at companies that have raised between $1M and $10M, with SOC 2 Type II, HIPAA BAA, DPA, guaranteed rate limits, and BYOK via AWS KMS included. Teams stuck between standard tiers that lack compliance features and an enterprise contract they cannot justify get a middle option.

    pricinghigh

    New 'Zep for Emerging Companies' pricing tier introduced at $13,000 per year

    Zep introduces a new mid-market pricing tier targeting companies that have raised $1M-$10M in funding, positioned between their standard tiers with enterprise features at a lower price point.

    Evidence
    +Zep for Emerging Companies $13,000for your first year Enterprise controls at a price built for companies that have raised between $1M and $10M.

    www.getzep.com

  8. LangSmith's Managed Deep Agents entered public beta: deep agents can now run hosted on LangSmith rather than on infrastructure you operate yourself. Teams already building long-running agents in the LangChain ecosystem can weigh the managed option against self-hosting.

    featurehigh

    LangSmith introduces Managed Deep Agents in Public Beta

    LangChain users can now use managed Deep Agents through LangSmith, expanding the platform's capabilities for agent orchestration.

    Evidence
    +LangSmith Managed Deep Agents is now in Public Beta Victor Moreira August 7, 2026 9 min

    blog.langchain.dev · also seen at 1 more source

  9. Replicate added Nvidia H200 GPUs in single, 2x, 4x, and 8x configurations, with 144GB of memory per card at $0.001525/sec (about $5.49/hr) for a single GPU. One catch: H200 capacity requires a committed-spend contract, it is not on-demand.

    featurehigh

    Added Nvidia H200 GPU options with multiple configurations and pricing

    Users now have access to H200 GPU instances in single, dual, quad, and 8x configurations. H200 GPUs with 144GB memory are available, supporting high-memory workloads. Capacity requires committed spend contracts.

    Evidence
    GPU1xCPU13xGPU RAM80GBRAM72GB 72GB
    +GPU1xCPU13xGPU RAM80GBRAM144GB Nvidia H200 GPU gpu-h200 $0.001525/sec $5.49/hr H200 capacity is available with committed spend contracts. 2x Nvidia H200 GPU gpu-h200-2x $0.003050/sec $10.98/hr Additional Multi-GPU H200 capacity is available with committed spend contracts. 4x Nvidia H200 GPU gpu-h200-4x $0.006100/sec $21.96/hr Additional Multi-GPU H200 capacity is available with committed spend contracts. 8x Nvidia H200 GPU gpu-h200-8x $0.012200/sec $43.92/hr Additional Multi-GPU H200 capacity is available with committed spend contracts. 144GB

    replicate.com

  10. The Gemini API released two embodied-reasoning endpoints in public preview, gemini-robotics-er-2-preview and a streaming variant, covering spatial reasoning, multi-robot coordination, and low-latency interaction. The same changelog carries an easy-to-miss deadline: gemini-robotics-er-1.6-preview shuts down on August 31, leaving projects on 1.6 the rest of this month to migrate.

    featurehigh

    Gemini Robotics ER 2 models released in public preview for advanced robotics applications

    Developers can now build real-time robotic agents with embodied reasoning capabilities, multi-robot coordination, and streaming support for low-latency interactions.

    Evidence
    +Gemini Robotics ER 2 in public preview: Released two new embodied reasoning model endpoints for robotics: gemini-robotics-er-2-preview: Advanced spatial reasoning, agentic code execution, multi-step tool orchestration, video moment finding, progress classification, and multi-robot coordination. gemini-robotics-er-2-streaming-preview: Optimized for real-time text streaming using the Live API, enabling low-latency robot agents with bidirectional audio and video input.

    ai.google.dev · also seen at 1 more source

  11. langchain-anthropic fixed caller-supplied tool_choice not being preserved on its way to the API. Before the fix, code forcing the model toward a specific tool this way was not necessarily taking effect; after upgrading, behavior returns to what callers intended, so paths that rely on it deserve a regression pass.

    api

    Fix tool_choice preservation in Anthropic integration

    Tool selection behavior is now correctly preserved when passed by callers, ensuring expected tool invocation patterns.

    Evidence
    +fix(anthropic): preserve caller `tool_choice`

    github.com

That's it for this issue. Get the next one by email — meanwhile the directory updates daily and each issue lives here permanently.