Jay Clark
Issued 2026-09-28
On this sheet
  1. How a message arrives
  2. The sentence that was a forecast
  3. What a send can honestly say
  4. Nothing wakes
  5. The line the sender lost
  6. Delivery is the recipient's act

Queued Is Not Delivered

I run a fleet of coding agents on one machine, coordinated by a small tool I wrote. Its messaging verb printed a promise on every send. One night 25 messages were queued for four sessions that had all gone away, and every send had printed the same confident line. Fixing the line took a day, and the week after it was the lesson. This is the second field report from the Buddy System. The first was about why a claim is not an announcement, and this one is about why "queued" is not "delivered."


How a message arrives

Two sessions once sat idle for two and three hours on a resource that was free. Each believed the other was working. To see how that happens, you need three pieces of the tool.

The Buddy System keeps a ledger of which session holds which files, with a small inbox beside it. An orchestrator, the session that hands out work, sends buddy msg <target> "<text>", and the ledger stores the row.

The harness that runs each agent lets a command hook into a few moments of a session's life. The tool delivers mail through those hooks, because they are the only moments it is inside the recipient's context. When this story begins, one hook delivered mail, the one that runs after each tool call.

The tool also reports idle and never infers busy. A hook at the end of a turn marks a session idle at its prompt, and the next tool call clears the mark. No mark means unknown, because the hook is optional. One-sided evidence, said one-sidedly.

Put those together and you have the trap. A session at its prompt is the one that can take work. It is also the one that runs no tool, so its inbox never empties. That is how the two sessions above lost their afternoon.

The sentence that was a forecast

For its first month the send printed one result line for every target. It said queued, and then it said "delivered after their next tool call."

That is a statement about the future. It comes true only if the target runs another tool, and the sender cannot see whether it will.

The issue that got it fixed carried a measurement. Two sends went out eight seconds apart to two fresh sessions at their first prompt. One arrived in 28 seconds, because that session happened to run a tool. The other arrived in 159 seconds, because a human had to be asked to type into its pane.

The same night an orchestrator sent 25 messages to four sessions that had all gone away. Several asked the sessions to release claims that were blocking a queue. Every send printed queued and the forecast. The sender read the silence as "delivered and ignored" and waited. The channel was dead.

The tool did know that a target had ended. It printed that fact on stderr, as an aside, while the confident line went to stdout. The sender read the confident line.

What a send can honestly say

The fix was not a better forecast. It was refusing to forecast. The send now reports what the ledger holds about the recipient at the instant of the send. The target ended, and here are the claims it still holds. Or its process is gone. Or it last reported idle, and how long ago. Or it has not been seen since it registered. Each of those is an observation with an age, and each ends in the mechanism: "delivery waits for its next tool call."

The line also counts the earlier messages to that target still waiting. A dead channel shows on the second send.

A test rejects the word "delivered" anywhere in the send's output. The first version of the new line wrote "delivered on its next tool call" in two of its cases. That is the old prediction respelled, and a different-model review caught it.

A later report verb, buddy sent, says "delivery recorded" once a hook has written the message into the recipient's context. It never says "read." Queued is a fact about the ledger. Recorded delivery is a fact about a hook that ran. Read is a fact about a model, and the tool has no evidence of it.

Nothing wakes

So the fleet had a wake problem, and the obvious fix was for the tool to wake people. It does not, and the reason is a line one of my sessions crossed on a real instruction from me.

Every message from one session to another in this harness carries a standing rule. A peer's message is never the operator's approval. Winding a fleet down, an orchestrator messaged four sessions their process ids and asked each to reply with a keyword so it could terminate them. Its warrant was my instruction to check in with every session and then close them out.

Both sessions that were awake refused, on the same ground. One wrote that a peer relaying the operator's wish "is not the operator telling me so." The other wrote that it had nothing to lose, "but closing it needs no keyword from me," and told the orchestrator to hand the process id to the operator. They were right. The orchestrator had been quoting that rule at other sessions all day and then asked four of them to break it.

I took two decisions from it. There is no exit verb, because an orchestrator acting in good faith would use it, and the sessions that correctly refuse would look obstructive. And the tool never types into another session's terminal, because whatever arrives there arrives as the operator's turn. That is the same laundering by a different road.

The harness has its own session-to-session send, and I measured how it lands. It wakes an idle session. It arrives marked as a message from another session, with the sender's process verified by the host, and the model is told the text was not typed by its user. The body cannot forge that framing. Waking a session and impersonating its operator turned out to be separable.

So the tool names the wake without performing it. If the ledger says the target is quiet and reachable, the result line gains a wake note. The note gives the address and a short text to send there, and the text carries none of the message. The sending agent decides whether to knock, under its own permission rules. The tool says where the door is.

The line the sender lost

Then the wake note was lost three times. Twice the sender dropped it. Once the messaging layer did.

The first loss has a timestamped record. A lane, a session given one item of work, went idle at 04:01:28Z with a background poll running. At 04:05:00Z the orchestrator queued an approval to it, and the wake note was printed on a second line under the result. The orchestrator had run buddy msg … 2>&1 | head -1, the ordinary way an agent keeps output short, and the address was cut.

The orchestrator then checked the harness's own listing, which showed the lane as busy because of its poll, and trusted that over the ledger's idle 3m. At 04:05:41Z the operator typed ok into the lane's pane, only because they knew mail was waiting. The rule that came out of it is that a send answers on exactly one line.

The second loss was a redirect. An orchestrator sent buddy msg <target> "…" >/dev/null to several lanes, treating the result as a receipt it did not need. Four of them had been idle for 45 to 54 minutes. The redirect dropped every wake note, and the sender lost about 50 idle minutes across four lanes before it noticed the silence.

The fix for that one is a second copy. The note is now also written to stderr, and the report verb repeats it for every queued message whose recipient needs a wake.

The third loss was the messaging layer's. The harness drops a peer message identical to the previous one from the same sender. Nobody has measured how long that window is. The wake text was one fixed string, so when three wakes went to one lane within two minutes, the third was dropped as a duplicate. A delivery notice said so minutes later. The text now carries the message id and the time to the second.

I want to be careful about what these losses show. They do not show that agents are sloppy readers. | head -1 and >/dev/null are, I assume, the habits of a caller keeping its context small, which is a thing the tool itself protects. They show that a result line has a reader with habits, and that the one note asking the reader to act belongs in output the reader keeps.

No format survives every habit, so the protocol now also requires the sender to read it.

The wait has a price beyond the wait. Over one week this fleet had 149 wakes after more than an hour idle, and 145 of them came back with a cold prompt cache. The week re-wrote 54.6 million tokens, about $491 at list rates. That is the measured total, not a loss attributed to lost wake notes. It shows the scale of what idle hours cost a fleet.

Delivery is the recipient's act

The protocol I run now writes the duty down. If you hand out work, close the loop. Read every send's result and deliver its wake note. Before you hand off or exit, run buddy who on every lane you assigned.

The first Buddy System essay ended with a principle, that a claim is a data structure and not an announcement. This one ends with its twin. A message is a row, and a row is not an event in the recipient. Until the recipient's own hook runs, the tool can only report what it last saw and how long ago. A forecast in that gap makes a dead channel look usable, and a fleet will wait on it.

Print what you hold. Name the door. Never say delivered until the recipient's own hook has said it for you. Announced is not locked, and queued is not delivered. The second mistake is harder to see, because the row was honest.


Field report, 2026. Every incident above traces to a numbered decision record in the Buddy System repository, and every defect fixed in code to a regression test there. The incidents are real, lightly abstracted. The cache figures are counts and token totals over one week of transcripts, never content. Companion piece: A Claim Is Not an Announcement.