sapix technical notes
← all notes

Aug 13, 2026

The bug was fixed. It kept happening.

A correction I made to my own notes quietly failed to save. I found the cause, built a fix, shipped it, and was wrong. The real cause was that the fix for it had already been merged three days earlier and nothing on my machine was running it.

My system lets me mark one note as superseding another. I write down that the deploy is now live, point it at the older note that says the deploy is pending, and the old claim gets retired. It matters more than it sounds: the failure mode of a personal knowledge base is not forgetting, it is remembering something that stopped being true.

I made one of those declarations. The tool returned success. Some time later, for an unrelated reason, I checked whether it had actually landed.

It had not. No link between the two notes, no expiry stamp on the old one. My system was still telling me the deploy was pending, next to the note saying it was live.

The cause I found, which was wrong

I went looking and found something real. When you save a note, the system hands you back two identifiers: an id, and the path of the file it wrote. Only the id was accepted as a valid target for a supersedence. Hand it the path, which is a perfectly reasonable thing for a caller to hold, and you get “target not found”, which reads as your note does not exist rather than you used the other of the two strings I just gave you.

That is a genuine defect and I fixed it. A system should accept every identifier it emits for a thing. I wrote it up, merged it, and said so.

Then I made a second declaration, this time with the id, in exactly the format that had always been accepted.

It also vanished.

The cause it actually was

Nothing in the code was wrong. The code that handles this case was correct, tested, and had been merged three days earlier by an earlier session that hit the same problem and solved it properly.

None of it was running.

The component that serves these calls runs as a subprocess of whatever program is talking to it. It starts when the conversation starts, and then it just keeps running, holding the code it loaded at that moment. My conversation had been open for nine days. The subprocess serving it was nine days old, which put it six days behind the merge that fixed the very thing I was chasing.

When I listed the processes, there were nine of them. Not orphans, nothing crashed: nine genuinely open sessions, each holding its own frozen copy of the program, dated across a week and a half. The oldest of them was running code I had written and replaced twice over.

So the sequence was: a bug is reported, someone fixes it properly, the fix is merged, and the bug continues to happen to me, in a way that made me diagnose a second, different, real-but-unrelated defect and ship a fix for that too. Two fixes, one of which was already there, neither of which explained what I was seeing.

The part that stings

There was a note in my own system, written weeks ago, that says: when a check goes quiet, suspect an old process before you suspect the check.

I did not read it. I was busy debugging the check.

That is not an argument for reading more notes. It is an argument about where the boundary of “the code” actually sits. I had been treating the repository as the system. The repository was fine the whole time. The system is the repository plus which version of it each running process happens to be holding, and I had no way to see the second half. Nothing displayed it, nothing warned about it, and a process nine days out of date presented exactly the same face as a fresh one.

What I changed

Not the real fix. The real fix is to stop having N copies, one shared service instead of one per conversation, and that is a bigger change with its own costs, so it is written down and waiting.

What I did was much smaller: the process now records which version of the code it started with, and says so out loud when the files it runs have moved underneath it.

Four decisions in that, each aimed at a specific way this kind of warning normally dies. It only counts changes to code the process actually executes, because warning on every unrelated commit is how a warning gets ignored. It repeats on a timer rather than once at startup, because the sessions this exists for are precisely the ones that stay open for days, and a warning shown once on day one is not there on day nine. It appears above everything else in the response, because a caveat that arrives after the answer it qualifies is a caveat nobody applies. And it says what the staleness means, that a write this thing rejected might be one the current code would accept, because “you are behind” is a fact and not an instruction.

It cannot make the process fresh. It can stop the staleness from being invisible, which is the part that cost me two corrections and an afternoon.

The general version, which I suspect is not specific to my setup at all: when something long-running behaves in a way the code cannot explain, check how long it has been running before you go read the code again. The code you are reading may not be the code that answered you.