AI Devtools Radar

第 4 期 · 2026年8月9日 · 8 月 3 日至 8 月 9 日当周

Anthropic 退役 Claude Opus 4.1,Claude Code 修复 Bash 权限绕过,Groq 撤下公开价格表

本期雷达监控了 32 个工具的 63 个来源,窗口内发布 33 条变化,选出 11 条,按重要性排序。比常规多一条,因为补入了卡在前两期窗口交界处漏掉的 Supabase 端点移除,它的截止日期在 9 月 23 日,不该再压一周。每条都带流水线抓到的前后对照证据,别只听我们说,展开看原文。

  1. Anthropic 已退役 claude-opus-4-1-20250805,对它的所有请求现在直接返回错误,官方建议迁移到 Claude Opus 5。配置里还写着这个模型 ID 的服务不会静默降级,只会收到报错,需要尽快改。做研究需要继续访问的,可以申请 External Researcher Access Program。

    deprecationhigh

    Claude Opus 4.1(claude-opus-4-1-20250805)模型已停用

    使用 Claude Opus 4.1 的用户和应用将收到错误,必须迁移到 Claude Opus 5。研究人员可以通过外部研究人员访问计划申请继续使用。

    证据
    +We've retired the Claude Opus 4.1 model (claude-opus-4-1-20250805). All requests to this model will now return an error. We recommend upgrading to Claude Opus 5. Researchers can request ongoing access through the External Researcher Access Program.

    docs.claude.com

  2. Claude Code 修复了一个 Bash 权限绕过问题,构造过的命令可以把自身的一部分藏过权限检查。依赖权限模式来限制 agent 能执行哪些命令的团队值得马上升级,无人值守和共享环境里跑的实例尤其如此。

    otherhigh

    修复通过精心构造的命令进行的 Bash 权限绕过漏洞

    解决了精心构造的 Bash 命令可能绕过权限检查的安全漏洞

    证据
    +Fixed a Bash permission bypass where a crafted command could hide parts of itself from permission checks

    github.com

  3. Supabase 管理 API 的 logs.all 端点将于 9 月 23 日移除,日志查询迁到新的 ClickHouse 后端端点,而且新端点只接受 ClickHouse SQL。直接调这个端点的脚本、集成和监控工具,都要在截止日期前完成查询改写和迁移。这条 7 月下旬就公告了,正好卡在前两期窗口的交界处被漏掉,趁离截止还有六周补上。

    apihigh

    logs.all Management API 端点已弃用,将由新的 ClickHouse 支持的 logs 端点替代

    查询 logs.all 端点的用户必须在 2026 年 9 月 23 日之前迁移到新的 logs 端点,并将查询转换为 ClickHouse SQL 方言。这会影响直接调用此端点的脚本、集成和工具。

    证据
    +The logs.all Management API endpoint is being removed on 23rd September 2026 (Wednesday), two months from this announcement. Log querying moves to a new ClickHouse-backed logs endpoint. The new endpoint accepts ClickHouse SQL only.

    supabase.com

  4. Kimi K3 在 GitHub Copilot 正式可用,由 GitHub 托管在 Fireworks AI 上,按量计费:每百万输入 token 3 美元、输出 15 美元、缓存输入 0.3 美元,和上期 Together AI 的报价一致。对 Business 和 Enterprise 计划默认关闭,管理员在 Copilot 设置里启用对应策略后,组织成员才能选到它。

    pricinghigh

    Kimi K3 定价公布:输入 token 每 100 万个 $3、输出 token 每 100 万个 $15、缓存输入 token 每 100 万个 $0.30

    用户将按使用量计费购买 Kimi K3 模型,输入、输出和缓存输入 token 的费率不同。由于 GitHub Actions 事故,定价文档发布暂时暂停。

    证据
    +Kimi K3 pricing, which will be $3 per 1M input tokens, $15 per 1M output tokens, and $0.30 per 1M cached input tokens.

    github.blog

  5. groq.com/pricing 整页换成了 6.5 亿美元融资公告和公司定位文案,原来按模型列出的 token 价格表没有了。要比较成本现在得进控制台或文档里找数字;把 Groq 算在成本模型里的团队,建议自己留一份此前费率的存档。

    otherhigh

    Groq 定价页面完全重新设计,重点放在公司定位和融资公告而非定价详情

    访问定价页面的用户将无法找到模型、token 和工具的详细定价表。页面现在展示的是 Groq 6.5 亿美元融资和公司定位的营销信息。这大幅影响了试图比较成本或集成 Groq API 的开发者。

    证据
    Smart, Fast, and Affordable Unmatched Price Performance Fast responses, scalable performance, and costs you can plan for. Start Building Large Language Models *Approximate number of tokens per $ AI Model Current Speed(Tokens per Second) Input Token Price(Per Million Tokens) Output Token Price(Per Million Tokens) [Detailed pricing tables for GPT OSS 20B 128k, GPT OSS Safeguard 20B, GPT OSS 120B 128k, Llama 3.3 70B Versatile 128k, Llama 3.1 8B Instant 128k, Qwen 3.6 27B 131k, and other models with specific prices]
    +Announcing our $650 million fundraise to scale global inference Read more> Every customer served. Every product sold. Every commit merged. Every agent task completed. That's inference. Training creates the possibility. Inference creates the value. As AI does more, inference multiplies. And inference is becoming the bottleneck. Groq was built for this. We pioneered the LPU. Now, with LPX, it works alongside NVIDIA's next-generation GPUs to deliver unparalleled inference capability, reliably, affordably, at scale. Fast or affordable is no longer a tradeoff. AI keeps training. Now it needs a better way to work. Groq makes inference work at scale.

    groq.com

  6. OpenAI 的 Fast mode 开始支持长上下文请求,GPT-5.6 Sol、Terra、Luna 上超过 272K token 的 prompt 也能走 Fast 档,官方称速度最高是 Standard 档的 2.5 倍。之前因为上下文太长只能回落到 Standard 的负载,可以重新测一遍。

    featurehigh

    Fast 模式现在支持 GPT-5.6 模型的长上下文请求(272K+ token),速度提升 2.5 倍

    使用 GPT-5.6 Sol、Terra 或 Luna 的开发者现在可以对超过 272K token 的长上下文提示使用 Fast 模式,获得显著更快的响应时间。

    证据
    +Fast mode now supports long-context requests for GPT-5.6 Sol, GPT-5.6 Terra, and GPT-5.6 Luna. As of today, long-context prompts exceeding 272K tokens can run in Fast mode, delivering speeds up to 2.5× faster than the Standard tier.

    platform.openai.com · 另有 1 个来源报告了同一变化

  7. Zep

    Zep 新增「Zep for Emerging Companies」档位,面向融资额在 100 万到 1000 万美元之间的公司,首年 1.3 万美元,SOC 2 Type II、HIPAA BAA、DPA、保证速率上限和 AWS KMS BYOK 都包含在内。卡在标准档合规不够、企业档又谈不起之间的团队,多了一个中间选项。

    pricinghigh

    推出新的「Zep for Emerging Companies」定价层,年费 13,000 美元

    Zep 推出针对融资 100 万至 1000 万美元的公司的新中端市场定价层,定位于标准层之间,提供企业功能但价格更低。

    证据
    +Zep for Emerging Companies $13,000for your first year Enterprise controls at a price built for companies that have raised between $1M and $10M.

    www.getzep.com

  8. LangSmith 的 Managed Deep Agents 进入公开测试,深度 agent 可以托管在 LangSmith 上运行,不用自己维护执行环境。已经在 LangChain 生态里做长时任务 agent 的团队,可以拿托管版和自建部署比一比运维成本。

    featurehigh

    LangSmith 推出托管 Deep Agents 公开测试版

    LangChain 用户现在可以通过 LangSmith 使用托管 Deep Agents,扩展了平台的代理编排功能。

    证据
    +LangSmith Managed Deep Agents is now in Public Beta Victor Moreira August 7, 2026 9 min

    blog.langchain.dev · 另有 1 个来源报告了同一变化

  9. Replicate 上新了 Nvidia H200 GPU,单卡、双卡、四卡、八卡四种配置,单卡 144GB 显存,每秒 0.001525 美元(约每小时 5.49 美元)。注意 H200 容量要签承诺消费合同才能用,不是随开随用的按需档。

    featurehigh

    添加 Nvidia H200 GPU 选项,多种配置和定价方案

    用户现在可以访问单个、双 GPU、四 GPU 和 8x 配置的 H200 GPU 实例。配备 144GB 内存的 H200 GPU 可用,支持高内存工作负载。容量需要承诺花费合同。

    证据
    GPU1xCPU13xGPU RAM80GBRAM72GB 72GB
    +GPU1xCPU13xGPU RAM80GBRAM144GB Nvidia H200 GPU gpu-h200 $0.001525/sec $5.49/hr H200 capacity is available with committed spend contracts. 2x Nvidia H200 GPU gpu-h200-2x $0.003050/sec $10.98/hr Additional Multi-GPU H200 capacity is available with committed spend contracts. 4x Nvidia H200 GPU gpu-h200-4x $0.006100/sec $21.96/hr Additional Multi-GPU H200 capacity is available with committed spend contracts. 8x Nvidia H200 GPU gpu-h200-8x $0.012200/sec $43.92/hr Additional Multi-GPU H200 capacity is available with committed spend contracts. 144GB

    replicate.com

  10. Gemini API 发布了两个具身推理模型端点的公开预览版,gemini-robotics-er-2-preview 和它的流式变体,覆盖空间推理、多机器人协同和低延迟交互。同一份 changelog 里还有一条容易漏看的截止日期,gemini-robotics-er-1.6-preview 将于 8 月 31 日关停,还在用 1.6 的项目这个月内就得迁走。

    featurehigh

    Gemini Robotics ER 2 模型推出公开预览版,用于高级机器人应用

    开发者现在可以构建具有具身推理功能、多机器人协调和低延迟交互流媒体支持的实时机器人 agent。

    证据
    +Gemini Robotics ER 2 in public preview: Released two new embodied reasoning model endpoints for robotics: gemini-robotics-er-2-preview: Advanced spatial reasoning, agentic code execution, multi-step tool orchestration, video moment finding, progress classification, and multi-robot coordination. gemini-robotics-er-2-streaming-preview: Optimized for real-time text streaming using the Live API, enabling low-latency robot agents with bidirectional audio and video input.

    ai.google.dev · 另有 1 个来源报告了同一变化

  11. langchain-anthropic 修复了调用方传入的 tool_choice 没有被完整保留的问题。在修复之前,靠 tool_choice 强制模型调用特定工具的代码未必真的生效;升级后行为回到调用方期望的样子,依赖这条路径的代码值得回归测一遍。

    api

    修复 Anthropic 集成中的 tool_choice 保留问题

    工具选择行为现在在调用者传递时被正确保留,确保了预期的工具调用模式。

    证据
    +fix(anthropic): preserve caller `tool_choice`

    github.com

本期到此为止。下一期可以直接送进邮箱,同时目录每日更新,每期周报永久可读。