← Writing

The agent pauses. The human approves. The approval lands on a different machine.

12 min readAI AgentsDistributed SystemsArchitecture
Contents· The problem this solves

You have built an agent that can send email to your customers. It works. Then someone asks the obvious question: what stops it emailing four thousand people by mistake? The answer is to make it stop and check with a person first. That pause looks like the simplest thing in the whole system. It is where agents quietly break once you have more than one server.

The problem this solves

An agent that only reads things is easy to live with. It looks things up, it summarises, and the worst it can do is be wrong in a paragraph you ignore.

An agent that does things is different. Sending email. Deleting records. Spending money. Publishing a page. None of those have an undo button.

A request goes to the agent, the agent sends the email, and four thousand inboxes receive it with no way to undo.
No pause anywhere in this line. By the time you find out what the agent decided, it has already happened.

So you add a stop. The agent works out what it wants to do, then waits and shows a person exactly what that is. Nothing happens until someone says yes.

The same flow, but the agent stops at a decision point and shows the person what it is about to do before anything is sent.
The whole feature is the diamond in the middle. Everything in this article is about what that diamond costs you.

People call this a gate, or human in the loop. The name does not matter. The shape is always the same: work, stop, ask, carry on.

Why you cannot just ask at the start

The tempting shortcut is to ask for permission once, up front, and be done with it. It does not work, and it is worth being clear about why.

At the start, the agent does not know what it is going to do yet. That is the entire reason you are using an agent instead of a form. Asking then is asking someone to approve a blank cheque.

Asking at the start gives you a vague question you cannot judge. Asking at the moment of action gives you a specific one you can read.
Same permission, two different moments. Only one of them gives the person something they can actually judge.

The version that works on your laptop

Every agent framework supports this and the code really is short. You stop at the step that matters, save the state, and pick it up again when the answer comes back.

// stop here, and save where we got to
const paused = await graph.invoke(input, { configurable: { thread_id } });

// ... the person reads the card and clicks approve, some time later ...

// pick the same job back up
await graph.invoke(new Command({ resume: true }), {
  configurable: { thread_id },
});
This is not wrong. It is what you should write first.

Run that on your laptop and it is perfect. Approve, and the agent carries on. Reject, and it does not. Ship it.

It worked because of something that code never says out loud: the same program that stopped the job is the one that starts it again. On your laptop that is guaranteed, because there is only one program. In production it is the first thing you lose.

Three things you lose in production

What follows is three separate versions of that same surprise, in the order you will meet them. Each one is harder to spot than the last.

What you loseWhat goes wrongHow easy to spot
1Only one serverThe job silently starts overVery hard. It looks like the agent forgot.
2Instant answersThe agent acts on yesterdayHard. Wrong in a way nobody checks.
3Only one way inAn approved job can reach a tool it should notOnly if you go looking.
None of these throw an error. That is what makes them ship.

One: the approval that gets lost

You know this one already, from real life. You ring a call centre. You explain the whole problem to a person. The line drops. You ring back, and you get somebody else, who has never heard of you, so you start again from the top.

That is exactly what happens here.

Once you run two servers behind a load balancer, which is the most ordinary thing you will ever do to a web app, the job stops on one machine and the approval arrives at whichever machine happens to be free. Clicking approve is a brand new request. Nothing about it says which machine was handling you before.

The approval click goes through the load balancer to machine B, while the half finished job is saved on machine A, so machine B starts the job over.
Machine B is not broken. It was handed a job it has never seen, so it does the sensible thing and begins a new one.

Saving the state means the job survives.

Only if the approval reaches the machine that saved it

The default in most frameworks saves that state inside the program itself. The documentation calls this durable, and it is, in the sense that it survives inside that one machine. That sentence is true and it is the most expensive sentence in this whole subject, because it sounds like a promise and describes a trap.

Now notice what does not happen. Nothing crashes. There is no error to alert on. Machine B was given a job it does not recognise, so it starts fresh, and the person watches the agent redo work it already did while ignoring the decision they just made.

You will not see it in testing if testing runs one server. You will not see it in load tests, because load tests never stop to think. You see it the week you scale up, as scattered reports that the assistant sometimes forgets things.

The fix is a shared notebook

Stop saving the state inside the machine. Put it somewhere every machine can read, which in practice means your database.

Both machines read and write one shared store, so whichever machine receives the approval can find the job and carry on.
The load balancer still picks at random. It stops mattering, which is the point.

Then make the wrong setting impossible

Moving the state is the easy half. The half that matters is making sure nobody can run the unsafe version by accident. A setting that defaults to wrong and is fixed by a line in a README will be wrong again eventually, on a new environment, or a rushed rollback, or a Friday.

So the program refuses to start:

export function assertGraphCheckpointerCoupling(env = process.env) {
  const isProduction = env.NODE_ENV === 'production';
  const usesPostgres = env.AGENT_GRAPH_CHECKPOINTER === 'postgres';

  if (isProduction && !usesPostgres) {
    throw new Error(
      'The LangGraph orchestrator requires AGENT_GRAPH_CHECKPOINTER=postgres ' +
      'in production: the default in-process MemorySaver is single-pod only, ' +
      'so the graph confirm/resume breaks across pods.'
    );
  }
}
It takes the settings as an argument, so you can test it without starting anything.

The message says what breaks, not just what is missing, so somebody who has never met this problem learns it from the crash. One more line goes into the log at startup saying which setting is live. That turns “is this deployed properly” into a glance instead of an investigation.

Then let a queue do the resuming

The shared notebook fixes finding the work. It does not fix doing it. The machine that catches your click is a web server with a browser waiting on it, and replaying an agent turn can take a minute or more. Run the agent right there and the click hangs. If anything times out halfway through, you are left not knowing whether your approval counted.

So the click does two small things and then stops. It writes your answer down, and it puts one job on a queue. Then it answers the browser. The real work happens a moment later, on whichever worker happens to be free.

The click writes the answer to the store and drops one job on a queue. A separate worker takes that job, reads the paused state back out of the store, and carries the agent on.
Two arrows leave the web machine and neither of them is the agent. The click got fast, and the slow work moved to where slow work belongs.

In Poplitu that queue is BullMQ, a thin layer over Redis. Which queue you pick matters far less than people expect, because you are buying the same four things from any of them.

A job that fails gets tried again instead of disappearing with the request that started it. A job that is mid flight survives someone restarting the web servers. A busy hour becomes a longer queue instead of a wall of timeouts. And the length of that queue is a number you can watch, which is usually the first honest sign that something further down is unwell.

There is one thing a queue will do to you if you let it, and it is better to know now than to find out. Almost every queue promises to deliver a job at least once, not exactly once. A worker that dies just before it reports success leaves the job looking undone, so the queue hands it to somebody else. Two workers then hold the same resume. If resuming means sending an email, that is two emails.

The queue cannot fix this for you, because only your database knows whether the work actually happened. The resume has to claim the job in the store before it does anything, in the same write that records it as claimed. The second worker finds nothing left to claim and goes home. This is the one piece of the design you should write a test for before you ship it, because it will not show up until the day a machine dies at exactly the wrong moment.

Two: the note left overnight

Picture a sticky note on a colleague’s desk that says “post this tomorrow”. They are off sick. They find it a week later. If they follow it exactly, they post it a week late and label it with the wrong day.

That is the second problem. A gate is a pause, and a pause means time passes. When the job starts again, almost everything must come back exactly as it was, or the person approved one thing and you did another. Almost everything.

Across the pause, the agent, the request and the card contents are replayed exactly, while the current date is asked again rather than replayed.
Everything is replayed except the one thing that must not be.
When the job restartsReplayed exactly?Why
Which agent is runningYesThey approved this one, not another
What was asked forYesThe approval was for this request
What the card saidYesApproving a card means approving what it showed
What day it isNo, asked againThe pause may have lasted all night
If you replay the clock as well, an approval left overnight books something for a day that has already gone.

The test for this is worth stealing. Run a real two part job, stop it and start it again for real, and use dates far in the past and far in the future rather than “tomorrow”. Otherwise the test quietly depends on a fake clock being installed, and passes for the wrong reason.

Three: the side door

Think of an office with a front desk. Everyone who walks in the front gets their badge checked. There is also a side door for people coming back from lunch, and nobody checks anything there, because they were already checked this morning.

Restarting an approved job is the side door. It is a different piece of code from the one that starts a job, it runs much less often, and it runs right after somebody pressed approve, so whatever it does feels blessed.

A fresh request and a resumed job are two separate ways in. Both reach the same corridor, where every tool call is checked before it can reach email, chat, payments or the database.
Two ways in, one corridor. Check the badge in the corridor, not at each door.

The obvious place to limit what an agent can do is when you hand it its list of tools. Give the support agent three tools and it cannot call a fourth, because it does not know a fourth exists. That is necessary and it is not enough, because the restart path builds that list again, from saved state, in code you tested least.

So the check happens in the corridor instead. Every tool call goes through one place, and that place asks whether this agent is allowed this tool, no matter which door it came in by.

Count the refusals before you start refusing

Switching a new restriction on blind is how you break a working feature for a customer at two in the morning. So the rule runs in two modes. In the quiet mode it notices a violation, counts it, and lets it through. In the strict mode it blocks it.

You run the quiet mode for a week and watch the counter. Zero means it is safe to switch on. Not zero means you have just learned something about your own system that reading the code would never have told you.

What to check in your own agent

Five things. The first one takes ten minutes and is the one that will bite you.

  1. Run two servers and approve something. Not one. Two. If it works, check you were not accidentally being sent back to the same machine every time.
  2. Make the unsafe setting impossible. A default that is wrong in production is an incident with a date on it.
  3. Approve something tomorrow. Leave a gate open overnight and check the restarted job uses today’s date, not the one it stopped on.
  4. Check the side door. Restarting a job is different code from starting one.
  5. Log the settings at startup. One line. It turns a question into a glance.

None of this is really about models, which is the part worth taking away. The hard bits of putting an agent into production are the bits that were hard long before agents existed: work that outlives a single request, work that moves between machines, and a person in the middle who takes their time.