Skip to content
Original
Habr · Вайбкодинг· serjfedik·· 1 day agoSelectedAI score76

When Automated Checks Lie: Five Cases from a Project Where an AI Agent Writes the Code

Original title: Когда автоматические проверки говорят неправду: пять случаев из проекта, который пишут ИИ‑агенты

AI overview

In a product project where an AI agent writes the code and the author doesn't read it, automated checks repeatedly reached the wrong conclusion. The author found 86 checks that no workflow had ever triggered, a secret scan that missed 438 of 1413 files because Git escapes Russian filenames by default, a new check that mistook WHERE for a table alias and let an injection slip through, and three false alarms from the test dashboard and the agent's replica.

Why it matters

The author walks through five real cases to show why automated checks produce false greens or false reds, and lays out validation rules that carry over to other projects.

Full text · AI translation

When Automated Checks Lie: Five Cases from a Project Written by AI Agents

Simple

5 min

1.2K

Case Study

I’m not a programmer. At least, despite having graduated in “Applied Informatics in Management” (DGTU) back in 2010 and learning C++, I don’t consider myself one today.

However, since June 2026, I’ve been managing a large product—a proprietary analytics platform, essentially an in-house ERP—whose code is written by AI agents. I issue tasks via voice (I can’t even remember the last time I typed out a task description on the keyboard); the agents write the code, run tests, and deploy the new version to our website. The product is built on Next.js 16, TypeScript, and PostgreSQL. As of October 1st, the project’s history already includes over 8000 commits and nearly 1,8 million lines of code.

I don’t read the code myself. So between the agent saying “done” and the new version appearing on the site, there are automated checks—called “watchdogs,” now numbering around two hundred. If even one watchdog turns red, the new version doesn’t make it to the site, and I get a Telegram message saying “deployment halted,” along with a list of the offending checks.

That’s how I find out whether the product is ready. And several times, these checks have given false results. Below are five examples: what happened, how we noticed it, and what we changed.


Eighty-Six Checks That No One Ever Ran

On September 1st, it turned out that the project had 86 checks, yet no automated process was ever running them. The new version would go straight to the server after the code was pushed. The checks only triggered when someone manually entered a command before deployment.

I can’t say exactly how long this went on—the project’s history doesn’t show it. All we can see is that the checks were written and sat in a folder, but they weren’t actually protecting the site.

Now, before each deployment, all checks are run, and the rollout waits for their results. If even one check turns red, the previous version stays live. The list of checks is automatically generated from the project’s command list, so you can’t forget to enable a new one. There’s a rule: you can’t disable a check just because it turns red and gets in the way. Everyone who knows about a disabled check still counts on it.


The Secrets Check Didn’t Scan Every Third File

(Here follow the remaining cases where the checks gave false results, how we discovered them, and what fixes we made.) On that same day, an audit was underway before deployment: five independent reviews, with no changes made to the code. One of those reviews reached a check that searches the history for leaked keys and passwords. This check runs before every deployment and always reports “no secrets found.”

My project contains many files with Russian names—guides, reports, analyses. By default, Git displays such filenames in encoded form. For example, a file named “guide.md” appears like this in the output:

"\320\263\320\260\320\271\320\264.md"

The check took that name as-is, reported “file not found,” suppressed the error, and moved on. In the end, 438 out of 1413 files—including 286 guides—never got checked, even though the check honestly reported “no secrets found.”

We ran the missed files through the fixed check—no real keys turned up. But the whole time, the third batch of unread files still had a green tick next to it.

The fix was two lines. The code was written by agents; here it is as-is:

git -c core.quotePath=false … -z
while IFS= read -r -d ''

The first line makes Git output file names without encoding them. The second reads a list of names separated by null bytes, so spaces and Cyrillic no longer break the process.


On September 2, we found a class of bugs in database queries where the counter quietly shows the wrong number. No message, no blank screen—just a different figure. We didn't find any such spots in our code, but we added a check for this class ahead of time.

To make sure the check actually worked, we deliberately put the same bug back into a working file. The check stayed silent. It took the keyword WHERE on the next line for a table alias—the table's second short name—and decided everything was fine.

If we hadn't injected the bug on purpose, the check would have stayed on the list as protection against a problem it can't see. After that, we made a rule: every new check gets tested with an injected bug right away. If a check never turns red, it proves nothing. Since September 10, a separate check enforces this—without that test, a release won't go out.

Three false alarms in one week

On the night of August 21, a message came in about two unresolved errors on the site, one of which was said to have been hanging for 1072 hours. At that moment there were no open errors in the production database, and a record that old was impossible—the database was younger.

It turned out someone on the team had spun up a copy of the panel with test errors, and the copy messaged me as if it were the live site. After a similar incident on August 18, we'd already put a safeguard in place, but it looked for signs of a copy in the message text, and there were none.

Now the panel only sends alerts after confirming it's the same panel that receives requests at the production address.

On August 22, it happened a third time. An agent was testing notification delivery on its own copy and sent me a real Telegram message saying "Background process has stalled." Everything on the live site was working at the time.

What stuck with the project from that week is a rule: a false alarm is worse than silence. After a few false red alerts, people stop looking at red, and a real failure goes by under the same color.

All checks are green, but the screen fell apart

I saw this one myself.

In September, one page of the site came to me with broken layout three deploys in a row. Card icons sat at different heights, titles at three different levels, text stuck to the cards, and the footer links were out of place. Every time, I spotted it in a second.

The build passed, code checks found no errors, and every run was green. But the checks read the text of files, while the layout only shows up on the rendered page in the browser. Not a single check ever looked at the rendered page. After the third time, on September 13, I added a rule: the agent opens the screen in the browser and checks it before proposing a deploy. This applies to any work you can see with your eyes. Now a screen counts as done only after the browser has opened it and measured three widths, including the phone width.

What changed
By date it looks like this:

  • 21 August — only the working dashboard sends failure alerts;

  • 1 September — a new version doesn't reach the site without checks;

  • 2 September — every new check is verified with a planted error;

  • 13 September — the screen is opened in the browser and measured.

On top of that came a rule from my second project. There the counter showed “1” when the database had 23 records, and all the checks were green. Since then, a section that displays numbers on screen doesn't count as done until real data has gone through it.

These rules boil down to three questions with the word “done” in them. Has this check ever gone red? Did you open the screen yourself? Has real data gone through that section?

There's also an unsolved case. A check that compares a value with itself is always green, and a machine can't tell it apart from the text. Only a planted error catches it, and the person who plants it has to be the one who remembers this rule.

Source: Habr · Вайбкодинг · habr.com