AI周刊 #102:OpenAI 发布 GPT-6 Sol/Luna,Anthropic 发布 Claude Opus 5.5
OpenAI 发布 GPT-6 Sol 和 Luna,Sol 定价 $2/$10 MTok、约为 GPT-5.6 的一半,DeepSWE 68.8%,Luna 定价 $0.10/$0.50 MTok。
Awaiting translation
OpenAI 发布 GPT-6 Sol 和 Luna,Sol 定价 $2/$10 MTok、约为 GPT-5.6 的一半,DeepSWE 68.8%,Luna 定价 $0.10/$0.50 MTok。
Awaiting translation
BB 是一个 MIT 许可的开源程序,让 Claude Code、Codex、Pi 和 ACP 智能体在同一个窗口里工作,7 个月内在仓库积累了约 4 千颗星。
Awaiting translation
NVIDIA 开源了 Agent 运行环境 OpenShell,用策略规定 Agent 能碰哪些文件、访问哪些服务,并在实际操作发生时执行这些规则。请求先经过沙箱识别发起程序,再由沙箱外的 Supervisor 核对策略后才建立外部连接;凭据由 Provider 管理,Agent 环境里拿到的是占位令牌,Proxy 转发前才替换成真实凭据。
Awaiting translation
作者统计自己项目五周的 Claude Code 日志(112 个会话、584 次运行),发现被入口文件(.claude/commands 和 .claude/agents)点名的文档平均被打开 12.3 次,而没有任何入链的 88 个 markdown 文件中 57 个一次都没被打开,平均仅 1.2 次。
Awaiting translation
A product designer with six years of SaaS experience used Claude Code to single-handedly build Котомка, a life-planning app. Nearly all the code was written by AI; his job was to define requirements, review the results, and make decisions.
Why it matters: Using a real repository, the author documented the pitfalls he hit while building a product on Claude Code alone, plus the rules, hooks, and testing guardrails he set up around the AI.
作者发布开源插件 claude-gym-coach,在 Claude Code 里完成训练计划、记录和周期复盘,采用 MIT 许可,只需已有的 Claude 订阅。
Awaiting translation
CopilotKit 团队在 GitHub 开源了 OpenMuse,一个可部署在自己服务器上的个人 AI 助理,iOS、Android 和网页三端代码全开源并采用 MIT 协议可商用。
Awaiting translation
OpenAI 将重新向新订阅者开放 200 美元的 Pro 订阅,但调整了额度计算方式。据 Тибо Соттио 说法,按 API 消耗折算,新方案提供的用量约为旧版 Pro 200 的一半。
Awaiting translation
作者根据 Claude 官方新出的提示词指南,整理出在 Claude Code 中使用 Opus 5.5 的几条做法。官方称 Opus 5.5 每次回复前都会自行决定思考多少,删掉提示词里的 think carefully 后回复更早且质量没有下降,作者用 grep 命令清理了 CLAUDE.md、rules 和 Skill 中的这类指令。
Awaiting translation
Martin Alderson 认为前沿实验室的推理价格战正在加速:OpenAI 在约两个月内把 GPT-5.6 Luna 降价 90%、GPT-6 Sol 降价 60%,Anthropic 的 Opus 5.5 输入输出降价 20%、缓存读取降价 60%。
Awaiting translation
Claude Opus 5.5 被称为截止今天的世界最强模型,同时也是最便宜的模型。作者用 Claude Max Plan 20x 实测:月费 $200 美金,5 天消耗 164 亿 Token,折合每月可消耗 600 亿 Token,考虑缓存后实际价值约 $15000~$20000。
Awaiting translation
skillmem 作者发布 0.12 版本,先把十六个不变量写进 docs/INVARIANTS.md,再让 Codex 和 Claude Opus 5.5 分别寻找反例,规则是发现必须附带可复现的失败测试,由 scripts/release-gate.sh 校验。
Awaiting translation
作者导出自己 49 个 Claude Code 会话、138 小时活跃工作的日志,统计出每小时平均消耗 1530 万缓存读取、86 万缓存写入、5.1 万输出和 560 普通输入 token,再按这套固定用量给 TeamoRouter、LiteAI、RouterAI、Polza AI、ProxyAPI 五家 API 网关算每小时花费。
Awaiting translation
营销人 Anton Budon 用 Claude Code 配合三星 S23 上的 Android 声级计 App,测出阁楼听音位在 125 赫兹处有低频隆起,并发现功放 loudness 按钮会整体抬高 7 分贝低频、参考音变化会让曲线失真。随后他让 Claude Code 接入 Яндекс 音乐 API,做了一个局域网遥控器,并在其中加入针对该频段削减分贝的均衡器。
Awaiting translation
有用户反映公司稳定使用一年多的 Claude Code Team 套餐被 ban,发帖求助稳定使用 cc 的方法,并表示愿意多付费继续使用。回帖者称这波是大范围封禁 Team 套餐,建议改用 Bedrock 和 Vertex 开账户,或使用老 Google 账号注册。
Awaiting translation
文章拆解了 JEV 这类专注快速判断的决策模型在 Agent 系统中的 8 类应用场景,包括 ReAct 行动循环、Computer Use、工具与技能路由、工具安全守卫、任务评估与观测、模型与 Agent 路由、RAG 与上下文管理、实时动态决策。
Awaiting translation
作者在实验用 BugSink 在异常发生时触发云端 Claude Code,让它读取生产环境只读数据和仓库、定位原因并提交修复,项目为实验性质、容错成本低。作者强调这套流程依赖透明的日志、追踪和遥测,而代码智能体默认做不好这些。
Awaiting translation
Drawgent 发布 0.2.1,把用户自己安装的 Claude Code、Codex 或 opencode 接到实时 Excalidraw 画布上,通过 MCP 提供画布工具,智能体可读取场景和截图、实时改图并检查自己的布局。
Awaiting translation
Anthropic 为 Claude 开发者开放插件提交门户,可提交扩展审核并发布到 Claude Directory。付费套餐开发者可提交带远程服务器的 MCP 连接器,或整合 MCP 服务器与 Agent Skills 并托管在 GitHub 的 bundle。
Awaiting translation
Claude Code 2.1.283 于 9 月 25 日发布,管理员可通过 availableModelsMatch: "exact" 只允许指定模型版本,并用 deniedModels 单独禁用模型,避免新版本未经批准进入受管环境。
Awaiting translation
Awaiting translation
Claude Tag in Slack can now use your personal connectors! You can now securely access that Drive doc, Salesforce account, or Warehouse table that you have personal access to right where the work happens. Avail. on Teams today and Enterprise next week https://claude.com/blog/claude-tag-now-supports-personal-connectors-in-channels
作者用 AI 智能体写出 1700 行 JS,在单个 HTML 文件里用 Canvas 绘制一分钟 1080p60 带音效的影片,再由 headless Chrome 逐帧渲染成 MP4,全程没有图片、视频和音频采样文件。
Awaiting translation
Anthropic 发布 Claude Code 2.1.282,修复了使用 --continue 和 --resume 启动时重发旧消息、extended thinking 块丢失,以及恢复后的会话可能无法处理请求的问题。
Awaiting translation
作者更新了自己在 Claude Code 中按角色分工的路由配置,适配 Opus 5.5:协调者用 Opus 5.5 拆解任务并验收,explorer 用 Haiku 4.5/low 检索代码。
Awaiting translation
作者主张在别人讲解方案时平均每三十秒问一个确认性问题,因为早期的小误解会层层放大,等到讲完再一起问就来不及了。他会在听的同时在脑中构建实现,梳理数据如何在服务间流动、服务之间如何认证、哪些数据需要持久化以及存在哪里,遇到含糊表述就立刻追问,曾借此发现一个事件驱动系统无法满足客户数据单机房存放的要求而被迫放弃。
Awaiting translation
METR 2025 年 7 月的实验让 16 名开源开发者随机在允许和禁止使用 AI 的条件下完成任务,结果用 AI 时反而慢 19%,而参与者自认快了 20%。
Awaiting translation
The author tested 10 open-source prompt injection detectors against 629 AgentDojo injection attacks—each buried in real tool output—plus 97 benign samples.
Why it matters: The author tested 10 open-source detectors against 629 real injection attacks, with full comparison data at both default thresholds and after calibration.
作者依据三家厂商 2026 年 9 月 24 日的文档,对比 GitHub Copilot code review、Anthropic 的 Claude Code Review 和 Gemini Code Assist on GitHub 在触发方式、读取的规则文件、能否拦截合并和单次成本上的差异。
Awaiting translation
作者用 LangChain 的 MongoDBGraphStore 做对比测试,仅输入 5 份文档就自动生成 17 种实体类型、34 种关系类型,part_of、Part Of、part of 被识别为三类独立关联,导致检索时关联片段丢失。
Awaiting translation
Claude Code 2.1.277 宣布支持 AGENTS.md,但作者实测发现关闭遥测后该文件从不加载:内置插件 agents-md 的 isAvailable 依赖远程开关 tengu_agents_md_mod。
Awaiting translation
On September 22, Anthropic released its flagship model Claude Opus 5.5, aimed at developers and teams who want agents to handle multi-step tasks like coding and data analysis. The company says it delivers better performance and lower cost than Opus 5.
Why it matters: Anthropic's published pricing and the default workload cost reduction help developers estimate the migration cost for long-running agent tasks.
Boris Cherny 用 Opus 5.5 配合 Lean 对 Claude Agent SDK 做形式化验证,几段简短提示词换来 16 个 PR,修复了多个 bug 和竞态条件。
Awaiting translation
Lovable has shipped Opus 5.5, which the company says matches Opus 5 in results while cutting the number of steps by one-third to one-half. On Lovable's internal benchmarks, Opus 5.5 ties Opus 5 on 0-to-1 builds and iterative code changes, and comes out 4% to 6% ahead on validation discipline; across all reasoning effort levels, steps per task drop by 26% to 57% and input tokens fall by 21% to 59%, with the differences significant at the 95% confidence level.
Why it matters: Lovable shares official comparison data between Opus 5.5 and Opus 5 on step counts and tokens, so readers can judge the real change in build efficiency.
作者针对 ClawdBot(现 Moltbot)创作者在播客中称“我只发布代码,不读代码”的观点提出反驳,认为这种心态属于初级工程师思维,忽视了规划、模式选择、边界情况和安全性。
Awaiting translation
Drawing on his own Codex setup, the author built a role-based model routing system for Claude Code: the main thread acts as coordinator, explorer uses Haiku for code search only, worker uses Opus for TDD implementation, verifier uses Sonnet to run checks independently, senior uses high-tier Opus for money, data, and concurrency, and reviewer switches to a different model for semantic review.
Why it matters: Based on a week of hands-on testing, the author shares the configuration, Hook enforcement, and cost trade-offs of multi-model division of labor in Claude Code, which you can adapt to your own multi-agent workflows.
TypeSafe 的 JEV 是一款专做 Agent 内部判断任务的模型,只回答 Choice、Score、Noul 三类问题,用自然语言定义问题后毫秒级返回结构化答案与概率分布,训练方法称为 RLCD,强调校准概率而非生成文本。
Awaiting translation
Paper Compute 工程师用 TypeSafe 的 Jev 做智能体会话标签实验,在 1,781 条工程师对话轮次上,Jev 1.13.0 中位响应 188 ms,比 GPT-4o mini 的 1,023 ms 和 Claude Haiku 4.5 的 1,050 ms 快约 5.4 倍和 5.6 倍,本地 Qwen3 8B 为 2,311 ms。
Awaiting translation
The author shows how to connect Claude Code to any model with OpenRouter: put OPENROUTER_API_KEY in ~/.config/openrouter.env.
Why it matters: With about twenty lines of shell, the author decouples Claude Code's harness from the model and shares the routing and pitfalls for four model slots.
作者用智能体从零开发了 pi 的移动端前端 Pi Pocket,六周内达到 300 次提交,如今几乎不再逐行审查代码。他认为智能体适合边界清晰的小任务、重构和规划,但产出的代码普遍过度设计,规模比自己写大约 10-20%,样式和文档也偏模板化,且模型在主观判断上总顺着用户。作者目前只把这种方式用于个人项目,工作代码仍以 Claude 生成为主,自己写的不到 1%。
Awaiting translation
Claude Code 从 2.1.277 版本开始支持 AGENTS.md:当文件夹中没有 CLAUDE.md 时,Claude 会检查并使用 AGENTS.md。该支持基于 Claude Code mods 构建,这是其即将推出的定制 Claude Code harness 的方式,属于内置 mod,用户之后也可以自行构建自定义版本的项目指令。
Awaiting translation