Aug 17, 2026
The answer was in a textbook I hadn't opened
A gate in my system fired on 92.5% of web searches. Every fix we tried was a recombination of parts we already had, until one measurement showed the textbook solution had existed for years. The interesting failure is why nobody, human or model, ever proposed it.
My system has a gate that runs before every web search. Its job is to say: wait, you already have notes on this. It decides by taking the search query, finding the nearest notes by embedding similarity, and checking the similarity score against a threshold.
We measured it against 604 real web searches. It fired on 92.5% of them.
At that rate a gate is not a gate, it is a header. Worse, when we hand-labelled what it surfaced, signal and noise occupied the same band: the worst garbage scored 0.600 while genuinely useful memories sat at 0.513. There was no threshold to move to. The number was not mis-tuned; the number was the wrong kind of number.
Every fix was a recombination
Here is what got proposed across several sessions, by me and by the models I work with: tune the threshold. Require more matches. Weight the score by how useful each note had historically been. Put a small language model behind the retrieval to judge the candidates.
Notice the shape of that list. Every item recombines parts the system already had: embeddings, thresholds, usage statistics, LLMs. Nothing in it asks the older question: what does the field that studies this problem actually do?
Because the field settled this years ago. Retrieval has two stages. An embedding model finds candidates fast and cheap, and then a different kind of model, a cross-encoder, reads the query and each candidate together and scores whether this document actually answers this query. Bi-encoders compress each text into one vector before they ever meet, which is exactly why my scores could not separate “answers the question” from “smells vaguely similar”. The limitation was known. The fix was standard. Every search engineer would have named it in a minute.
We found it by accident, scanning a model catalogue for something else. A hosted reranker judged our test pairs in 126 to 336 milliseconds, at a fifth of a cent per call, and separated the same two queries at 0.76 versus 0.13 where cosine gave 0.637 versus 0.447. On the full replay of real searches, the fire rate dropped from 91% to 62% before any calibration at all.
Why nobody proposed it
The first reason is vocabulary. “Specialized NLP models” is not a commercial category. Embeddings got the marketing, because embeddings have a story you can demo in fifteen minutes: text becomes a vector, vectors go in a database, the database powers your chatbot. A reranker produces a number, not a demo. The people who talk about cross-encoders are at information-retrieval conferences and inside search teams, not on the channels where everyone learned to build with LLMs. After 2023, an entire generation of builders, and the models trained on their writing, learned one recipe: send everything to a language model and ask for JSON.
So the working mental model had two layers: deterministic code below, general-purpose LLM above. Anything too fuzzy for code went to the LLM. But there is a whole middle tier between them, models that do not generate anything, that answer exactly one narrow question extremely fast: does this document answer this query? Do these two sentences contradict each other? Which entities appear in this text? Rerankers, entailment models, entity extractors. Mature fields, closed vocabularies, established metrics. Invisible from inside the two-layer model.
The second reason is worse, because it was self-inflicted. Months earlier we had evaluated a reranker for a different feature and rejected it, correctly: it was a generative LLM used as a reranker, thirty seconds of latency for no gain. The note we wrote down said “reranker: rejected”. Not “generative reranker in this specific spot: rejected”. Every later session that consulted memory found the bare word and moved on. One imprecise noun poisoned an entire class of solutions for months. A rejection stored without the mechanism it rejected is a landmine.
The paradigm underneath
The fix to the gate is almost the least interesting part. What changed is the map. The system stops being “code plus an LLM with tools” and becomes something more like regions: deterministic machinery moves and transports data, small discriminative models measure specific relations, and the big model does what only it can do, open reasoning over context that has already been filtered by things faster and cheaper than itself.
That reframe immediately found more targets. The nightly job that hunts contradictions between my notes runs a generative model where an entailment classifier is the purpose-built tool. My entity extraction is regex where a zero-shot extractor is the purpose-built tool. Each of those is now a measured experiment rather than a hunch, because the other thing this episode taught us is that nothing in this tier gets adopted on vibes: frozen test pairs, human-graded labels, and the expensive failure, suppressing a memory that mattered, as the headline metric.
Two rules came out of the wreckage, and they are written where every future session reads them. Before proposing a mechanism for a judgment task, name the class of problem and ask what the relevant field’s standard tool is; only then compare it against recombining what you already have. And when you reject something, record the mechanism you rejected, never the bare word. The first rule is how the answer gets found. The second is how it stays findable.
The gate now runs the reranker in shadow, logging both verdicts on every real search while the labelled set grows. The threshold that fired on everything is still there, still user-facing, until the data says otherwise. That part we did learn the slow way: the last thing you want, right after discovering a blind spot, is to trust the next thing you see with your eyes still adjusting.