跳到正文

#模型发布

今日 0 条
10/1周四
9/30周三
  1. Habr · Codex76

    OpenAI 发布 GPT-6.1 Sol,API 价格约为 Astra 的五分之一

    OpenAI 发布面向编程、文档处理和任务自动化的 GPT-6.1 Sol,公司称其在部分测试中接近 GPT-6 Astra。在真实代码库任务测试 DeepSWE 1.1 中,该模型追平 Astra,而执行成本约为其五分之一;在应用操控测试 OSWorld 2.0 中,最高推理档位下比 GPT-6 Sol 提升 7 个百分点。

    推荐理由:GPT-6.1 Sol 在 DeepSWE 1.1 上追平 Astra 而成本约为其五分之一,读者可据此判断编程任务的性价比变化。

9/29周二
  1. Tproger · Программирование88

    OpenAI 在 DevDay 发布 GPT-6.1 Sol,价格仅为 Astra 的五分之一

    OpenAI 在 9 月 29 日旧金山 DevDay 上发布 GPT-6.1 Sol,API 名为 gpt-6.1-sol,定价为每百万输入 token 2 美元、输出 10 美元,缓存输入 0.10 美元,标准价格是 GPT-6 Astra 的五分之一。

    推荐理由:OpenAI DevDay 发布 GPT-6.1 Sol,价格降至 Astra 的五分之一,并同步更新 Codex、Agents API 与插件体系,可据此判断成本与工具链变化。

9/23周三
  1. Tproger · Программирование80

    Anthropic 发布旗舰模型 Claude Opus 5.5

    Anthropic 于 9 月 22 日发布旗舰模型 Claude Opus 5.5,面向让智能体执行编程、数据分析等多步任务的开发者与团队,官方称相比 Opus 5 性能提升且成本下降。

    推荐理由:Anthropic 官方公布的定价与默认负载成本降幅,可帮开发者判断长时智能体任务的迁移成本。

9/22周二
  1. Sebastian Raschka72

    小米发布开源权重模型 MiMo-V2.6-Pro,在 Artificial Analysis 智能指数上以 46 分成为开源权重模型第一,每任务成本 0.13 美元,输入 0.435 美元/1M tokens、输出 0.87 美元/1M tokens,采用 1.02T 总参数、42B 激活参数的 MoE 架构。

    引用Artificial Analysis@ArtificialAnlys

    MiMo-V2.6-Pro debuts as the top open weights model on the Artificial Analysis Intelligence Index (46). At $0.13 per Intelligence Index task, it lands on the Intelligence vs. Cost per Task Pareto frontier @Xiaomi has just released MiMo-V2.6-Pro, an open weights model with major advances in intelligence over its predecessor, MiMo-V2.5-Pro (Intelligence Index: 26). Despite the improvement, it retains the same attractive pricing at $0.435 per 1M input tokens (with a 99% cache-hit discount) and $0.87 per 1M output tokens. This makes MiMo-V2.6-Pro one of the most cost-efficient models to deploy. MiMo-V2.6-Pro is an MoE model with 1.02T total parameters and 42B active parameters. Stay tuned for additional analysis of the model. Check out MiMo-V2.6-Pro full benchmarking breakdown here: https://artificialanalysis.ai

8/11周二
6/10周三
  1. Andrej Karpathy75

    Andrej Karpathy 评价 Claude Fable 5 发布,指出它与 Mythos 是同一底层模型,只是增加了安全防护,在几乎所有基准上以明显优势达到 SOTA。

    引用Claude@claudeai

    Fable 5 is state-of-the-art on nearly all tested benchmarks, with exceptional performance in software engineering, knowledge work, scientific research, and vision. The longer and more complex the task, the larger Fable 5’s lead over our other models.