Sep 17, 2026
The expensive tool lost
I needed to tell two kinds of note apart, and I had a shortlist of the right modern tools for that job. I reached for the most capable one. It came last, below a pattern match I wrote in five minutes, and the only reason I know is that I had built the answer key first.
Some of my notes describe a state of the world: this tool is broken, that application is blocked, this service was down on Tuesday. Those rot within days. Others describe something I learned: isolate the failing part before trying to recover it. Those do not rot at all.
The first kind is dangerous when it goes stale. A note saying a submission had failed was true when written, false the next day, and was read seventeen days later as current fact, which produced a duplicate insurance claim, an apologetic email, and money that had to move between two people.
So I need to tell them apart automatically, and the obvious first attempt is to look for the words. Broken, failed, blocked, crashes, times out.
That does not work, and the reason is a nice piece of English. “X fails when Y” is the standard way to phrase a durable insight. “Classifiers fail when the exemption targets the syntax” is a permanent lesson. “The mail search fails on multi word queries” is a bug that will be fixed. Same verb, opposite lifespan. No list of words separates them.
The thing that does separate them is the subject. A named product against a generic noun. So I wrote a crude pattern match for “does this title contain something that looks like a product name”, which took about five minutes, and I built an answer key: eighty titles, labelled by hand, one at a time, against a written question.
Then I compared. The crude pattern scored 0.73 on the standard measure. The word list scored 0.70. The specialised model I had been intending to use all along, the expensive one, the one my own notes say is the right class of tool for this kind of judgment, scored 0.65.
It lost to the five minute regex, and its mistakes explain why. It ranked “reject tools that fail under operational constraints” near the top, because that sentence is about failure. It is about failure the way a book about grief is sad. It asserts no state at all. The model was answering a question about topical relevance, which is what it was built for, and I had quietly assumed that pointing it at a different question would make it answer that one.
I did eventually run the proper tool for the job, a model that extracts named entities, and it scored 0.77 on its own and beat the pattern by nothing at all once combined. What it did do was beat the entity extractor I already had in production by a wide margin, which is worth knowing for a different reason.
The part I want to keep is not any of the numbers. It is that I had a principled argument for the expensive tool, drawn from my own recorded reasoning, and the argument was wrong, and nothing in the argument could have told me so. The answer key could, and it cost an hour of reading eighty lines and deciding what each one was.
I have been treating “build the answer key first” as the tax you pay before doing the interesting part. It is the part that decides whether the interesting part was worth doing.