Cloudflare 开源决策模型 Clef:基于 Qwen3.8-27B 与 Qwen3.5-9B
Cloudflare 在 Birthday Week 期间开源两个决策模型 Clef(27B)与 Clef-flash(9B),Apache 2.0 许可,权重挂在 HuggingFace,分别基于 Qwen3.8-27B 和 Qwen3.5-9B,原生带视觉编码器,上下文窗口扩展到 64k。
Awaiting translation
Cloudflare 在 Birthday Week 期间开源两个决策模型 Clef(27B)与 Clef-flash(9B),Apache 2.0 许可,权重挂在 HuggingFace,分别基于 Qwen3.8-27B 和 Qwen3.5-9B,原生带视觉编码器,上下文窗口扩展到 64k。
Awaiting translation
OpenAI 上线 GPT-6.1 Sol,作者实测后认为它在软件工程测试 DeepSWE v1.1 上已追到 Astra 水平,价格比 Astra 便宜 80%,每百万 token 输入 2 美元、输出 10 美元,命中缓存输入 0.1 美元。
Awaiting translation
Google DeepMind 发布 Gemini 4 Argon,单次输出上限从 64K 提升到 1M tokens,主打长任务与多步骤推理。
Awaiting translation
Why it matters: 汇总了 Gemini 4 Argon 的官方评测数据与 Google 内部落地案例,可对照各家旗舰模型的能力差异。
谷歌发布旗舰模型 Gemini 4 Argon,官方公布的 18 项测试中独占第一 12 项,DeepSWE v1.1 以 77.9% 位列全球第一,AutomationBench-AA 以 77.5% 领先 Claude Sonnet 5.5 的 71.3%。
Awaiting translation
OpenAI has released GPT-6.1 Sol for coding, document processing, and task automation. The company says it comes close to GPT-6 Astra on some tests. On DeepSWE 1.1, a benchmark of real-world codebase tasks, the model matches Astra while costing about one-fifth as much to run. On OSWorld 2.0, which tests app control, it beats GPT-6 Sol by 7 percentage points at the highest reasoning tier.
Why it matters: GPT-6.1 Sol matches Astra on DeepSWE 1.1 at roughly one-fifth the cost, which gives you a sense of how the price-performance tradeoff for coding tasks has shifted.
OpenAI 在 9 月 29 日旧金山 DevDay 上发布 GPT-6.1 Sol,API 名为 gpt-6.1-sol,定价为每百万输入 token 2 美元、输出 10 美元,缓存输入 0.10 美元,标准价格是 GPT-6 Astra 的五分之一。
Awaiting translation
Why it matters: OpenAI DevDay 发布 GPT-6.1 Sol,价格降至 Astra 的五分之一,并同步更新 Codex、Agents API 与插件体系,可据此判断成本与工具链变化。
OpenAI 于 9 月 22 日在 API 发布 gpt-6-sol 和 gpt-6-luna,两者接受文本和图像输入、只生成文本。标准处理下,GPT-6 Sol 每百万输入 token 收费 2 美元、缓存 token 0.20 美元、输出 10 美元;GPT-6 Luna 分别为 0.10、0.01 和 0.50 美元。
Awaiting translation
On September 22, Anthropic released its flagship model Claude Opus 5.5, aimed at developers and teams who want agents to handle multi-step tasks like coding and data analysis. The company says it delivers better performance and lower cost than Opus 5.
Why it matters: Anthropic's published pricing and the default workload cost reduction help developers estimate the migration cost for long-running agent tasks.
Awaiting translation
MiMo-V2.6-Pro debuts as the top open weights model on the Artificial Analysis Intelligence Index (46). At $0.13 per Intelligence Index task, it lands on the Intelligence vs. Cost per Task Pareto frontier @Xiaomi has just released MiMo-V2.6-Pro, an open weights model with major advances in intelligence over its predecessor, MiMo-V2.5-Pro (Intelligence Index: 26). Despite the improvement, it retains the same attractive pricing at $0.435 per 1M input tokens (with a 99% cache-hit discount) and $0.87 per 1M output tokens. This makes MiMo-V2.6-Pro one of the most cost-efficient models to deploy. MiMo-V2.6-Pro is an MoE model with 1.02T total parameters and 42B active parameters. Stay tuned for additional analysis of the model. Check out MiMo-V2.6-Pro full benchmarking breakdown here: https://artificialanalysis.ai
NVIDIA has released Nemotron 3.5 Lightning, a customizable open-source model built for persistent agents, and it's now available for free in Cline.
Why it matters: NVIDIA's new open-source model is free to use on Cline, so you can decide whether it's worth switching for high-frequency agent workloads.
Andrej Karpathy 评价 Claude Fable 5 发布,指出它与 Mythos 是同一底层模型,只是增加了安全防护,在几乎所有基准上以明显优势达到 SOTA。
Awaiting translation
Fable 5 is state-of-the-art on nearly all tested benchmarks, with exceptional performance in software engineering, knowledge work, scientific research, and vision. The longer and more complex the task, the larger Fable 5’s lead over our other models.