What a green light can hide
Ten ways an automation fakes success, and the check that catches each
An automated system rarely fails loudly. It fails green: the run finishes, no error shows, and nothing happened. Every pattern below was caught in my own lab, on my own systems, and each one now has a check that fires without a person watching. Use the list on any system anyone offers you.
The ten quiet failures
A quiet failure is a failed run that reports success. No error, no alarm, no result. They rot a system because nothing rings.
- 01
The login worked.
A service refused the request. The refusal sits deep in the output, and the run ends green.
The check: any refusal from any service ends the run red, by rule, before anyone reads a line.
- 02
The data came in.
Every source came back empty, the fallbacks ran out, and the report still says gathered.
The check: an empty result is a failure until a person says otherwise.
- 03
The query ran.
Zero records, clean exit. The shape is right and the content is nothing.
The check: zero rows with a clean exit is red by default, never green by default.
- 04
The file was written.
It was. It is empty.
The check: a written file must have size and the expected content, read back at its exact path before anyone claims it.
- 05
It succeeded.
The body of the reply says failed. The wrapper around it says success. The wrapper is what got read.
The check: the content is inspected, never the wrapper.
- 06
The call went through.
It timed out, was caught quietly, and was retried in silence, again and again.
The check: a repeated timeout pages a person on the third try.
- 07
The parse worked.
Nothing matched, and an empty answer went forward as if it were the answer.
The check: nothing matched is a result, and it is reported as one.
- 08
Every route was tried.
Three quiet fallbacks in a row, and no error at the end of them.
The check: a run that falls back and never ends in an error is red.
- 09
Nothing else was touching that file.
Two processes raced for the same file. One skipped its write and said nothing.
The check: a lock, a count of skips, and a page at the fifth.
- 10
It was saved in the right place.
Written to a different folder, or never checked at all. Everything downstream reads the old copy.
The check: every write is read back at the exact path before it is claimed. The read-back is the claim.
Four ways a report lies about a failure
The failures above hide. A review of a failure can hide too. It fails in four named shapes.
The rule that catches all four: any review whose root cause lands entirely outside our own decisions gets a second pass. There is always a decision we owned: the service we picked, the gate we skipped, the assumption we never tested, the monitoring we never built.
Case: three lies at once
The strongest entry in my own record, because it was not one failure. It was three independent mechanisms manufacturing one false success, found together.
From my own lab · August 17, 2026
A setup job reported success. It had been refused.
The job needed elevated rights and did not have them. The system refused it. The transcript still read like success, for three reasons at once. First, the error was swallowed: a setting meant to stop the script on any error does not reach the scheduler's own commands, so the refusal printed as a note and the script carried on. Second, the success banner was unconditional: it printed whatever happened. Third, the self-check was hollow: it accepted a flag it never used, so it had never once looked at the thing it claimed to check, and it printed OK over an empty scan.
The rule it produced, now in the shared template every setup script inherits: success is a read-back, never a print. A pass names what it scanned, and a pass over nothing fails. Errors stop the run or surface; they never fall through. A flag that does nothing is a lie at rest. The scripts not yet refit are listed by name until they are.
Case: the capability that was never there
From my own lab · July 8, 2026
Two agents imported a module from the day they were written. The module did not exist.
No error, no alarm. The feature it enabled reported zero on every single run, and zero looked like a slow week instead of a dead limb. The fix was not a better prompt. It was a rule: a capability counts as present only when it has run once and been read back.
Case: eight agents listening, one clock
From my own lab · September 14, 2026
The board showed eight agents polling for work. All eight stamps carried the same second.
One script took one timestamp and wrote it eight times. A poll stamp proves only that a timer fired, and here it did not even prove that. The check built to notice a silent board had no state file, so it had never run. The first report counted 17 events in the board's whole life; an older copy on another drive held 146, so the first count was corrected in writing too.
The rule it produced: count outcomes, never heartbeats. A system is alive when it has done something a person can point to, not when it says it is listening.
Case: healthy, and two weeks behind
From my own lab · October 1, 2026
Every health light on the knowledge index was green. It knew nothing newer than two weeks.
The store was up and the model service was up. The retrieval test returned its hits, because it matched old entries and would have matched them forever. Three documents written after the last rebuild returned zero matches; one older control returned 141. The age probe alone told the truth: red at fourteen days. The trend probe read green on too little history, which is green by ignorance.
The rule it produced: a liveness check never stands in for a freshness check. Ask the system about something it should have learned yesterday.
One word for a lie
Status words drift. Done gets gamed, so a new word appears, and each new word buys one more state of plausible deniability. My systems are allowed five status words, computed by a machine from evidence. Prose cannot change them.
- DeliveredA reference was fetched again, independently. The only green.
- StagedThe work exists and waits on a named human gate.
- BlockedCannot proceed. Names the blocker and the owner.
- DarkThe default. Everything unproven lands here, never upward.
- FraudA reference was given and failed the re-check. The row lied. An incident, not a miss.
Teardown: what a cheap chatbot costs a twenty person company
A chatbot costs a few hundred dollars and answers in seconds. Its real bill arrives later, in places nobody puts on an invoice. Here is where the cost lands, so you can price it yourself before you buy one.
- It does not know your businessIt learned from the open web, not from your service list, your prices or your policies. Asked something specific, it produces a confident sentence with nothing behind it. You find out when a customer does.
- It keeps no record you can useWhen an answer was wrong, you cannot reconstruct what it said, when, or to whom. There is nothing to correct and nothing to show. The first time this matters is the first time a customer disputes what they were told.
- It is connected to nothingEvery answer still gets copied into your inbox, your calendar or your books by a person. The typing moved. The work did not.
- Someone on your team becomes its unpaid maintainerAnything it cannot answer escalates to whoever set it up. That person already has a job. This quietly becomes a second one, and it appears on no schedule.
- You end up checking every answerThe point of buying one is to stop reading every inquiry. Instead you read every reply before it goes out, because you cannot tell which ones are safe. That is the original job, with a tool attached.
- The person who set it up is goneIt was configured in an afternoon by someone who has since moved on. Nobody owns it now. The day it breaks, you are looking for a login.
When a chatbot is the right purchase. One question, one answer, no records, no money, no customer relationship. A widget that answers your opening hours and hands the person to a human is a good buy. It stops being a good buy the moment the answer has to be right about your business, or the conversation has to be remembered.
Three questions that price any vendor. Where does the answer come from, and can you see it. What gets written down when it answers, and can you read it later. Who is responsible when it is wrong, and what do they do about it. A vendor with straight answers to those three is worth a conversation, whatever the tool costs.
Three cheaper paths, and what each one actually is
Cheaper than a custom build is easy to find. Knowing what you are buying is the part that decides whether it was cheap. Here are the three paths people take instead, what each one really is, and the case where it is the right call.
- A chatbot on the websiteIt answers from general knowledge, and it carries no record. Right when the questions are general and the answers carry no consequence for anyone. The teardown above prices this one.
- A low cost outfit building it for youYou get a working system and nobody who owns it. The build gets priced. The changes do not. Right when your process is finished, written down, and will not move for a year, and when you have somebody in house who can hold the code themselves.
- Doing it yourself with general toolsThe most expensive of the three if your hours are billable, and the best of the three if they are not. Right when you enjoy this work, you have the time, and nobody downstream is waiting on the result.
The question that separates all three. Not what it costs to build. What it costs to change. Every system inside a working company gets changed, and the change is where the money actually goes. Ask any path you are considering what a change costs, who does it, and how long you wait for it. A build with no answer to that question is a build you will be replacing.
What this page is not saying. None of the three is a mistake. Each is the right purchase for somebody, and two of them are the right purchase for most companies. A custom build is the right purchase in one case: the work crosses your inbox, your books and your calendar, it repeats every day, and it has to answer to written rules. Outside that case, one of the three above will cost you less.
Teardown: where the document goes when someone pastes it into a free tool
Somebody on the team has ten minutes and a lease to read. They paste it into a free AI tool, get a clean summary, and move on. Nothing was charged. Something left the building, and the client was not told. Here is where that document can go, and how to find out which of these applies to the tool you already use.
- It sits in a conversation that is keptMost tools store what you send so the thread still works tomorrow. The question is for how long, and who inside that company can open it. The answer is written down in the tool's own terms, in a section about retention.
- It becomes training materialSome free tiers take what you give them as the price of the free tier. Many paid tiers let you switch that off, and some do not offer the switch at all. Look for the setting. If there is no setting, that is your answer.
- A person reads itSeveral services review a sample of conversations for quality and safety. That is a support employee, in another country, reading your client's lease. It is disclosed. It is rarely read.
- It leaves with no record on your sideYou cannot show what was sent, when, or by whom, because the tool was used on a personal login that your company does not own. The first time this matters is the first time a client asks.
What it costs, and why it is not the ten minutes. Nothing on that list is a fine. The cost is that a client's information left your building without the client knowing, and in several trades you owe that client a duty about exactly this. A privacy notice on your website does not cover a tool you did not tell them about.
Three questions you can answer today. Does the free tier let you turn training off, and where is the switch. Who can read a conversation, and is it disclosed. What happens to the data when the account is closed. The answers are in the terms, in plain enough language, and they take about twenty minutes per tool to find.
When a free tool is fine. Public information, generic drafting, and anything that names nobody. It stops being fine the moment the text identifies a client, a tenant, an address, an employee, a claim, or a figure that is not yours to publish. That line is the whole rule, and it is easy to teach.
Teardown: what a green dashboard actually proves
A vendor opens a dashboard. Everything on it is green. The room relaxes. Green is the least informative colour on that screen, and here is how to find out what it is really reporting.
- It can mean the check never lookedA test that trusts a flag nobody reads, or a check pointed at a demonstration set instead of your live one, prints a pass every single time. It has never once inspected the thing you believe it inspects.
- It can mean the check looked at nothingAn empty result and a healthy result are the same colour. No errors across no records is a perfect score for a system that did no work at all.
- It can mean the number is real and oldThe light was accurate when it was last written. If the writing stopped a fortnight ago, the light stayed green and nobody noticed, because a missing update is not an alarm unless somebody made it one.
- It can mean the check cannot failAsk what happens when it cannot reach its target. Does it go red, does it go quiet, or does it go green. A check that cannot fail is decoration. It exists to be photographed.
The one question that separates those four. Ask what that light would have to see in order to turn red. An answer naming a specific condition, on your live data, describes a real check. An answer describing the light itself describes a picture.
Three questions for any dashboard you are shown. What did this check read, by name. When did it last go red, and what was done about it. What does it do when it cannot run at all. A screen that has never been red has either an excellent system behind it or a very decorative one.
When a green light is enough. When you can name the condition that turns it red, and when you have seen it red at least once. Until both are true, that screen is a statement of confidence. Whether your system is healthy remains a separate question.