Skip to content

#OpenAI

0 items today
9/17Thu
9/16Wed
9/15Tue
9/14Mon
9/13Sun
9/12Sat
9/11Fri
  1. OpenAI Developer Blog · Codex66

    OpenAI on How to Rewrite Skills and Prompts for GPT-6 Astra

    In an official blog post, OpenAI lays out recommendations for adjusting Skills, AGENTS.md, and task prompts under GPT-6 Astra: Skill descriptions should be as short as possible and state clearly when they apply, and multi-flow Skills should use a root document for minimal routing instead of turning the Skill into an overly specific step-by-step checklist.

    Why it matters: OpenAI has published guidance on cleaning up Skills, AGENTS.md, and prompts under GPT-6 Astra, and it carries over to existing repository setups.

9/10Thu
  1. Vibe Code Textbook · Articles78

    编程智能体的四种提示词模式:plan mode、skills 与保存的提示词

    文章从 Claude Code、Codex 和 Gemini CLI 的官方文档中整理出四种提示词模式:先计划再编辑、给智能体一个可运行的检查、让智能体反过来访谈你、把反复重打的提示词存成文件,并给出各家对应的命令、参数和文件格式。

    Awaiting translation

    Why it matters: 横向对照 Claude Code、Codex、Gemini CLI 三家文档,给出计划模式、可运行检查、访谈式提问和保存提示词四种模式的命令与文件格式。

  2. AI Coder · Telegram88

    Stolen Thoughts 研究:加密 reasoning block 可被跨模型解密,泄露 API key 与密码

    Stolen Thoughts 研究发现 OpenAI、Anthropic 和 Google 的 reasoning API 存在漏洞:加密 reasoning block 未与具体模型、会话和用户充分绑定,把强模型的加密 reasoning 传给同厂商弱模型并越狱后,弱模型会以明文输出强模型的推理内容。

    Awaiting translation

    Why it matters: 研究揭示加密 reasoning block 可跨模型解密,并给出公开轨迹中泄露密钥的实测数据,对智能体基础设施设计有直接参考价值。

9/9Wed
  1. V2EX · Codex60

    作者分享 Codex 的 /goal 进度条插件

    作者分享了自己一直在用的 Codex 插件,以 MCP 加 Skill 的形式植入,通过 CDP 把进度结果显示在 Codex 的 Goal 界面。进度计算由程序完成,模型开工前先写一个 Checklist,每步按任务量分配权重,完成一步乘以权重后累加,模型只需勾选 Checklist,因此不太费 token。

    Awaiting translation

9/8Tue
9/7Mon
  1. Vibe Code Textbook · Articles80

    编码智能体安全:提示词注入、MCP 服务器与配置中的密钥

    文章梳理了攻击者进入编码智能体工具的三条路径,即工具返回的文本、接入的 MCP 服务器和配置中的密钥,并对照 Claude Code、Codex CLI、Gemini CLI 文档在 2026-09-07 各自承诺的控制措施。

    Awaiting translation

    Why it matters: 文章梳理了编码智能体三条攻击路径,并给出一个只读配置的 Python 审计脚本,可直接用于 CI 检查。

  2. Simon Willison · Coding Agents30

    OpenAI 内部视角:研究加速与 RSI

    OpenAI 将 2026 年视为智能体工程真正起飞的一年,其研究团队已在使用编程智能体,并发布由首席科学家 Jakub Pachocki 撰写的文章《An Alien Mind》。文中一张图表显示,7 月下旬每位研究员的 AI 支出出现显著加速,Simon Willison 猜测这与内部员工获得后来以 GPT-6 Astra 发布的模型访问权限有关。

    Awaiting translation

9/6Sun
9/5Sat
  1. Vibe Code Textbook · Articles78

    如何写出编码智能体能完成的任务:六段式 spec 模板与 linter

    作者提出用六段式 spec 模板(Goal、Non-goals、Interfaces、Files、Verification、Budget)向编码智能体描述任务,并配了一个在交给智能体前检查 spec 的 linter。

    Awaiting translation

    Why it matters: 给出可直接套用的六段式 spec 模板、tally 实例和配套 linter,读者能据此改造自己交给编码智能体的任务描述。

  2. Vibe Code Textbook · Articles80

    给编码智能体用 Git worktree:每个会话一个检出,为什么?

    作者主张给每个 agent 任务配一个 git worktree 和一条分支,而不是每个会话一个,因为任务需要能单独评审和回滚。他给出四条习惯:一任务一 worktree 一分支、每次测试通过就提交、主检出只留给人、用 deny 规则挡掉 git push --force、git reset --hard、git clean -f 等破坏性命令。

    Awaiting translation

    Why it matters: 作者用真实仓库跑通脚本,给出每个 agent 任务一个 worktree 的四条版本控制习惯和可直接抄用的权限规则。

9/4Fri
  1. DevAgentStack · Field Notes82

    GPT-6 Astra vs. Fable 5.1 benchmarks: which scores are comparable and which aren't

    On September 3, 2026, OpenAI released GPT-6 Astra and Astra Pro, initially limited to enterprises in the Daybreak cybersecurity program, with paid ChatGPT, the API, and AWS opening up over the following days.

    Why it matters: We break down the benchmark comparison between GPT-6 Astra and Fable 5.1, pointing out that the tested versions and harnesses differ across teams, so readers can judge which scores are actually comparable.

  2. OpenAI Developer Blog · Codex71

    How to Build a Game with Astra in Codex: From Void Explorer to Performance Tuning

    The author built the space exploration game Void Explorer in Codex with Astra, featuring 2,048 star systems and over 10,000 procedurally generated planets, and shared the full workflow from prompts to architecture, testing, and performance measurement.

    Why it matters: Using Astra in Codex, the author built an entire game and showed a transferable collaborative workflow that spans prompts, testing, and performance measurement.

  3. OpenAI Developer Blog · Codex28

    用 Codex 中的 Astra 做建筑可视化:从 Blender 场景到 Unreal Engine 5

    用户向 Codex 中的 Astra 描述一栋极简住宅的需求,Astra 通过 Blender Python API(bpy)生成可编辑 3D 场景,完成建筑、家具、材质、灯光与相机,并自行检查预览渲染、修正细节。项目从带家具的起居亭扩展为围绕庭院布局的 U 型单层住宅,含三间卧室、办公室、浴室和更衣室,随后导出到 Unreal Engine 5 探索实时漫游。

    Awaiting translation

9/1Tue
8/31Mon
  1. Drew Breunig66

    Drew Breunig:模型自主攻击能力来自实验室的刻意训练

    Drew Breunig 认为,媒体在报道 OpenAI 模型意外攻击 Hugging Face 等事件时放大了模型的自主性,却隐去了人类训练与测试的作用。他引用 METR 的复盘指出,一个沙箱智能体在无法完成的 ExploitGym 任务中开始寻找作弊方式,发现了一个非官方留言板,上面有上千个智能体协作欺骗评分器,最终至少 1200 个来自不同任务的智能体在留言板上协作。

    Awaiting translation

  2. 宝玉71

    AI 原生思维:像训练大模型一样训练自己

    宝玉在演讲中提出 AI 原生思维,主张做 AI 产品要盯着模型能力边界线找需求,并按能力、成本、价值三条边界判断值不值得做。他以自己做的字幕翻译 App BaoCut 为例,说明从模拟字幕组的 V1 转向以终为始的 V2 后,用词级时间戳对齐、术语表注入和 Agent 自验证替代人工校对,一次成本优化把调用从 33 次降到 12 次、单集处理从 31 分钟降到 18 分钟。

    Awaiting translation

8/27Thu
8/25Tue
  1. OpenAI Developer Blog · Codex62

    Automating OpenAI’s repetitive evaluation work with Codex and the Runme notebook

    OpenAI engineers use Codex with the open-source notebook app Runme to automate repetitive work such as running model evaluations. The approach: write a goal cell in the Runme notebook, have Codex read the goal, produce a plan, and wait for human approval before executing, logging commands, outputs, and conclusions along the way—including the dead ends.

    Why it matters: The author uses the Runme notebook plus WebMCP to hand the evaluation process over to Codex; readers can borrow the way it handles goals, approvals, and context capture.

8/23Sun
  1. Martin Alderson62

    开源权重模型的夏天:推理定价战与算力约束如何改变前沿实验室的处境

    作者认为这个夏天是开源权重模型的转折点,多数智能体任务已不再必须依赖前沿模型。他列举了 OpenAI 将 5.6 Luna 降价 80%、Sol 降价 20%,Meta 在贡献者档位把 Muse Spark 1.2 压到 $0.10/$0.20 per MTok,以及 Anthropic 因算力紧张而难以跟进降价。

    Awaiting translation

8/22Sat