Skip to content
Original
Reddit · ClaudeCode / Codex / VibeCoding· /u/Longjumping-Play6541·· 1 day agoAI score61

rashomon 更新:检测测试被填充或翻转,并估算 Claude Code 用量成本

Original title: We shipped what this sub asked for: catching tests that get padded or flipped instead of fixed

The title and summary in the selected language are awaiting translation.

AI overview

rashomon 发布新功能,--timeline 会标记两类模式:只有测试文件被改动后原本失败的测试立刻变绿,以及同一条测试命令在中间无改动的情况下既通过又失败。

Full text

A week ago I posted here asking how people know if Claude Code did something unexpected. A few of you had already written your own hooks or were grepping the transcript to check, and in the replies, three separate people independently described the same failure mode: an agent that can't fix a broken test leaves it alone and adds new, trivial passing tests next to it, so the suite exits green and the summary says "all tests pass."

We just shipped the thing that catches that.

rashomon --timeline now flags two specific patterns: a failing test that goes green right after only test files got edited, and the identical test command passing and failing with nothing in between. Neither requires storing test names or output, it's built entirely from command-line matching and file categorization that was already part of how we capture things, no new privacy tradeoff.

We also shipped rashomon spend, which estimates what your Claude Code usage would cost at API list prices, broken down by agent, model, and session, plus the one line we think is actually new: spend in turns that had a failed call the closing summary never mentioned.

Both of these came directly out of this thread and the comments on a follow-up post. Happy to walk through exactly how if anyone wants the details.

Repo: https://github.com/altrace-dev-role/rashomon

Do these kinds of checks and the failed-call spend metric seem useful for catching things you wouldn't notice from the agent's final summary alone?

Source: Reddit · ClaudeCode / Codex / VibeCoding · reddit.com