Skip to content

Costs and usage limits

Where tokens go, how to make usage limits last longer, and how to diagnose runaway costs.

Latest curated items

Items 1–20 · 33 total
10/6Tue
  1. DEV Community · MCP78

    FP8 pitfall: GPU bill dropped 47%, but the model outputs “!!!!!!”

    The author ran Qwen2.5 7B/32B/72B on a single AMD MI300X with vLLM ROCm, priced at $2.99/GPU-hr. The BF16 baseline was 7B at $0.227/M, 32B at $0.77/M, and 72B at $1.67/M output tokens.

    Why it matters: The author benchmarked FP8 quantization on the MI300X and found that per-token billing can hide the model's output degrading into gibberish, then gave a reusable way to verify it.

  2. Reddit · ClaudeCode / Codex / VibeCoding76

    A local proxy spreads Claude Code requests across multiple Max accounts and switches before the quota runs out.

    The author open-sourced claudemanager, a local daemon that Claude Code points to via ANTHROPIC_BASE_URL. It only changes the request's Authorization header to route sessions to the Max account with the most remaining capacity in its 5-hour, weekly, and per-model windows, switching at custom thresholds before those windows fill up.

    Why it matters: The author also open-sourced a local proxy that automatically distributes Claude Code traffic across multiple Max accounts based on remaining quota, and logs requests along the way.

  3. Hacker News · MCP78

    Flash-Agents: an MCP plugin that hands Claude Code's coding tasks to a DeepSeek Flash worker

    Flash-Agents is a Claude Code plugin that delegates bounded coding work—implementing slices, porting tests, reviewing diffs, mapping out a codebase—to a DeepSeek V4.1 Flash worker, while Claude keeps architecture, acceptance criteria, and final review.

    Why it matters: The author outsources Claude Code's coding tasks to a DeepSeek Flash worker and shares the sandbox, patches, and measured data, so you can judge the cost and safety boundaries for yourself.

10/5Mon
  1. DEV Community · Codex78

    How I Stopped Codex from Burning Through My Usage Quota

    The author found that two Codex browser automation tasks consumed 170,123 and 110,180 tokens respectively, so they set out to control usage through model selection, configuration files, and task splitting.

    Why it matters: Drawing on real measurements where two browser tasks burned through hundreds of thousands of tokens, the author shares a quota-saving approach: switch models and configurations based on task difficulty.

10/4Sun
  1. 佬刘AI78

    Planning with GPT, Execution with DeepSeek: A Cost-Saving Two-Model Workflow

    The author has GPT (gpt-6.1-sol, reasoning tier high) handle planning, key decisions, and acceptance in Codex, and calls DeepSeek-V4.1-Flash through the official DeepSeek Harness to write code, run experiments, and fix bugs—together producing a local PDF toolkit with five working features.

    Why it matters: The author splits the work between GPT for planning and DeepSeek for execution to get a PDF toolkit running, and shares three prompts plus cache usage data that can carry over to cutting costs on long tasks.

9/30Wed
  1. Habr · Вайбкодинг80

    How Flawwow Used a Product Sandbox to Let 150 Non-Engineers Ship 262 AI-Generated Projects

    Artem Gambitsky, co-founder of the Russian e-commerce platform Flawwow, walks through the company's internal product sandbox: it lets colleagues with no engineering background push apps written by AI agents straight to production. In four months, 150 people submitted 262 projects and ran about 5000 deployments—none of that code was ever read by a developer.

    Why it matters: The author lays out the four layers of protection that let non-engineers write code with AI agents and ship it safely, plus the resource pitfalls hit along the way. All of it can be adapted to your own in-house sandbox.

  2. Habr · Codex76

    OpenAI releases GPT-6.1 Sol, with API pricing at roughly one-fifth of Astra's

    OpenAI has released GPT-6.1 Sol for coding, document processing, and task automation. The company says it comes close to GPT-6 Astra on some tests. On DeepSWE 1.1, a benchmark of real-world codebase tasks, the model matches Astra while costing about one-fifth as much to run. On OSWorld 2.0, which tests app control, it beats GPT-6 Sol by 7 percentage points at the highest reasoning tier.

    Why it matters: GPT-6.1 Sol matches Astra on DeepSWE 1.1 at roughly one-fifth the cost, which gives you a sense of how the price-performance tradeoff for coding tasks has shifted.

9/29Tue
  1. Tproger · Программирование88

    OpenAI 在 DevDay 发布 GPT-6.1 Sol,价格仅为 Astra 的五分之一

    OpenAI 在 9 月 29 日旧金山 DevDay 上发布 GPT-6.1 Sol,API 名为 gpt-6.1-sol,定价为每百万输入 token 2 美元、输出 10 美元,缓存输入 0.10 美元,标准价格是 GPT-6 Astra 的五分之一。

    Awaiting translation

    Why it matters: OpenAI DevDay 发布 GPT-6.1 Sol,价格降至 Astra 的五分之一,并同步更新 Codex、Agents API 与插件体系,可据此判断成本与工具链变化。

9/23Wed
  1. Tproger · Программирование80

    Anthropic Releases Flagship Model Claude Opus 5.5

    On September 22, Anthropic released its flagship model Claude Opus 5.5, aimed at developers and teams who want agents to handle multi-step tasks like coding and data analysis. The company says it delivers better performance and lower cost than Opus 5.

    Why it matters: Anthropic's published pricing and the default workload cost reduction help developers estimate the migration cost for long-running agent tasks.

9/21Mon
9/7Mon
9/5Sat
  1. GitHub Blog · Copilot71

    GitHub Copilot launches Project HydraFusion, using multi-model runtime orchestration to improve coding quality

    GitHub has launched Project HydraFusion as a research preview in the Copilot CLI. It uses runtime orchestration to pick an execution plan across models from multiple providers. Users select it just like any other model, and billing follows each model's standard rates.

    Why it matters: GitHub lays out three orchestration modes for HydraFusion and compares cost versus quality across three benchmarks, so you can judge the trade-offs of multi-model orchestration on real coding tasks.

9/3Thu
9/2Wed
  1. 陈与小金 · AI Coding 博客78

    Claude Code subagents: dispatch 13 AI workers at once, and the main conversation only gets 3 conclusions

    Drawing on a hands-on session where he dispatched 13 subagents to build a storyboard, the author walks through the subagents feature that both Claude Code and Codex have: subagents work in their own separate windows and hand only their conclusions back to the main conversation.

    Why it matters: Using a hands-on session where he dispatched 13 subagents to build a storyboard, the author shows how subagents keep their work outside the main conversation and send back only the conclusions.

  2. 宝玉78

    Anthropic's E-commerce AI Agent Engineering Guide: Architecture, Latency and Cost Optimization, and Production Practices

    Anthropic has published a guide dissecting e-commerce AI agents. Drawing on deployment experience with retailers, e-commerce platforms, and teams in travel, entertainment, and telecom, it proposes a single-agent architecture that puts Claude in a standard agent loop, uses skills to cover long-tail needs, and calls tools to work with existing systems. The guide says that in comparative testing, this architecture beats both sub-agent designs and the approach of cramming everything into the prompt.

    Why it matters: Drawing on enterprise e-commerce agent deployment experience, Anthropic lays out a complete engineering approach covering a single-agent-plus-skills architecture, latency and cost optimization, and memory and security evaluation.

  3. Cursor · Changelog66

    Cursor launches self-hosted machines, keeping tool execution within your own network

    Cursor supports self-hosted machines: code repositories, build artifacts, and secrets all stay on internal machines within your own infrastructure, and the agent handles tool calls locally. My Machines connects a single laptop or VM to a personal workflow, while Team Pools are named worker queues for teams or enterprises—scaling capacity up with requests and down when workers disconnect. Pools aren't tied to code repositories, and idle machines can sleep and then resume within a reconnection window.

    Why it matters: The official docs lay out pooled scheduling and sandbox integration for self-hosted machines, so readers can judge whether tool execution can stay within their own network.

  4. Paper Compute · Engineering Blog76

    别只量代码,量工程决策:用会话记录算出一次架构决策的 65 倍放大

    作者提出用 agent 会话记录衡量一次工程决策的下游影响,即 blast radius(影响范围),并用自家 tapes 项目的一次架构决策做验证:设计文档会话花费 57.05 美元,后续引发 4704.74 美元工作量,分布在 70 个人工会话、3 名工程师、8 个仓库和 27 天中,另有 404 个自动化评测会话花费 431.54 美元。

    Awaiting translation

    Why it matters: 作者用自家一次架构决策的 474 条会话记录,展示如何把决策的下游成本量化成可复用的指标。

8/27Thu
  1. Addy Osmani · Blog71

    如何审计你的 Agent 配置文件:CLAUDE.md、Skills 与 Hooks 的定期清理

    Addy Osmani 建议每隔几周运行一次 Claude Code 的 /doctor,单独用 /memory 检查记忆,并让每条指令重新证明自己的价值,因为模型、harness 和代码库都在变,旧配置会留下。

    Awaiting translation

    Why it matters: 作者结合自身配置审计经验与近期研究,说明 Agent 配置文件为何会腐化,以及如何按节奏清理。

  2. Cline · Blog71

    Cline 实测八个模型做 IMO 2026:DeepSeek V4 Flash 以 0.12 美元拿到金牌线

    Cline 让八个模型在自家 harness 里做 IMO 2026 六道题,证明由 GPT-5.5 和 Claude Opus 5 双盲按 0–7 分制评分、Gemini 3.1 Pro 仲裁,金牌线为 29 分。

    Awaiting translation

    Why it matters: Cline 用同一套 harness 盲评八个模型做 IMO 2026,给出分数与单次成本对照,可看开源权重模型的实际性价比。