The agent pauses. The human approves. The approval lands on a different machine.
Contents· The problem this solves
You have built an agent that can send email to your customers. It works. Then someone asks the obvious question: what stops it emailing four thousand people by mistake? The answer is to make it stop and check with a person first. That pause looks like the simplest thing in the whole system. It is where agents quietly break once you have more than one server.
The problem this solves
An agent that only reads things is easy to live with. It looks things up, it summarises, and the worst it can do is be wrong in a paragraph you ignore.
An agent that does things is different. Sending email. Deleting records. Spending money. Publishing a page. None of those have an undo button.
So you add a stop. The agent works out what it wants to do, then waits and shows a person exactly what that is. Nothing happens until someone says yes.
People call this a gate, or human in the loop. The name does not matter. The shape is always the same: work, stop, ask, carry on.
Why you cannot just ask at the start
The tempting shortcut is to ask for permission once, up front, and be done with it. It does not work, and it is worth being clear about why.
At the start, the agent does not know what it is going to do yet. That is the entire reason you are using an agent instead of a form. Asking then is asking someone to approve a blank cheque.
The version that works on your laptop
Every agent framework supports this and the code really is short. You stop at the step that matters, save the state, and pick it up again when the answer comes back.
// stop here, and save where we got to
const paused = await graph.invoke(input, { configurable: { thread_id } });
// ... the person reads the card and clicks approve, some time later ...
// pick the same job back up
await graph.invoke(new Command({ resume: true }), {
configurable: { thread_id },
});Run that on your laptop and it is perfect. Approve, and the agent carries on. Reject, and it does not. Ship it.
It worked because of something that code never says out loud: the same program that stopped the job is the one that starts it again. On your laptop that is guaranteed, because there is only one program. In production it is the first thing you lose.
Three things you lose in production
What follows is three separate versions of that same surprise, in the order you will meet them. Each one is harder to spot than the last.
| What you lose | What goes wrong | How easy to spot | |
|---|---|---|---|
| 1 | Only one server | The job silently starts over | Very hard. It looks like the agent forgot. |
| 2 | Instant answers | The agent acts on yesterday | Hard. Wrong in a way nobody checks. |
| 3 | Only one way in | An approved job can reach a tool it should not | Only if you go looking. |
One: the approval that gets lost
You know this one already, from real life. You ring a call centre. You explain the whole problem to a person. The line drops. You ring back, and you get somebody else, who has never heard of you, so you start again from the top.
That is exactly what happens here.
Once you run two servers behind a load balancer, which is the most ordinary thing you will ever do to a web app, the job stops on one machine and the approval arrives at whichever machine happens to be free. Clicking approve is a brand new request. Nothing about it says which machine was handling you before.
Saving the state means the job survives.
Only if the approval reaches the machine that saved it
The default in most frameworks saves that state inside the program itself. The documentation calls this durable, and it is, in the sense that it survives inside that one machine. That sentence is true and it is the most expensive sentence in this whole subject, because it sounds like a promise and describes a trap.
Now notice what does not happen. Nothing crashes. There is no error to alert on. Machine B was given a job it does not recognise, so it starts fresh, and the person watches the agent redo work it already did while ignoring the decision they just made.
You will not see it in testing if testing runs one server. You will not see it in load tests, because load tests never stop to think. You see it the week you scale up, as scattered reports that the assistant sometimes forgets things.
The fix is a shared notebook
Stop saving the state inside the machine. Put it somewhere every machine can read, which in practice means your database.
Then make the wrong setting impossible
Moving the state is the easy half. The half that matters is making sure nobody can run the unsafe version by accident. A setting that defaults to wrong and is fixed by a line in a README will be wrong again eventually, on a new environment, or a rushed rollback, or a Friday.
So the program refuses to start:
export function assertGraphCheckpointerCoupling(env = process.env) {
const isProduction = env.NODE_ENV === 'production';
const usesPostgres = env.AGENT_GRAPH_CHECKPOINTER === 'postgres';
if (isProduction && !usesPostgres) {
throw new Error(
'The LangGraph orchestrator requires AGENT_GRAPH_CHECKPOINTER=postgres ' +
'in production: the default in-process MemorySaver is single-pod only, ' +
'so the graph confirm/resume breaks across pods.'
);
}
}The message says what breaks, not just what is missing, so somebody who has never met this problem learns it from the crash. One more line goes into the log at startup saying which setting is live. That turns “is this deployed properly” into a glance instead of an investigation.
Then let a queue do the resuming
The shared notebook fixes finding the work. It does not fix doing it. The machine that catches your click is a web server with a browser waiting on it, and replaying an agent turn can take a minute or more. Run the agent right there and the click hangs. If anything times out halfway through, you are left not knowing whether your approval counted.
So the click does two small things and then stops. It writes your answer down, and it puts one job on a queue. Then it answers the browser. The real work happens a moment later, on whichever worker happens to be free.
In Poplitu that queue is BullMQ, a thin layer over Redis. Which queue you pick matters far less than people expect, because you are buying the same four things from any of them.
A job that fails gets tried again instead of disappearing with the request that started it. A job that is mid flight survives someone restarting the web servers. A busy hour becomes a longer queue instead of a wall of timeouts. And the length of that queue is a number you can watch, which is usually the first honest sign that something further down is unwell.
There is one thing a queue will do to you if you let it, and it is better to know now than to find out. Almost every queue promises to deliver a job at least once, not exactly once. A worker that dies just before it reports success leaves the job looking undone, so the queue hands it to somebody else. Two workers then hold the same resume. If resuming means sending an email, that is two emails.
The queue cannot fix this for you, because only your database knows whether the work actually happened. The resume has to claim the job in the store before it does anything, in the same write that records it as claimed. The second worker finds nothing left to claim and goes home. This is the one piece of the design you should write a test for before you ship it, because it will not show up until the day a machine dies at exactly the wrong moment.
Two: the note left overnight
Picture a sticky note on a colleague’s desk that says “post this tomorrow”. They are off sick. They find it a week later. If they follow it exactly, they post it a week late and label it with the wrong day.
That is the second problem. A gate is a pause, and a pause means time passes. When the job starts again, almost everything must come back exactly as it was, or the person approved one thing and you did another. Almost everything.
| When the job restarts | Replayed exactly? | Why |
|---|---|---|
| Which agent is running | Yes | They approved this one, not another |
| What was asked for | Yes | The approval was for this request |
| What the card said | Yes | Approving a card means approving what it showed |
| What day it is | No, asked again | The pause may have lasted all night |
The test for this is worth stealing. Run a real two part job, stop it and start it again for real, and use dates far in the past and far in the future rather than “tomorrow”. Otherwise the test quietly depends on a fake clock being installed, and passes for the wrong reason.
Three: the side door
Think of an office with a front desk. Everyone who walks in the front gets their badge checked. There is also a side door for people coming back from lunch, and nobody checks anything there, because they were already checked this morning.
Restarting an approved job is the side door. It is a different piece of code from the one that starts a job, it runs much less often, and it runs right after somebody pressed approve, so whatever it does feels blessed.
The obvious place to limit what an agent can do is when you hand it its list of tools. Give the support agent three tools and it cannot call a fourth, because it does not know a fourth exists. That is necessary and it is not enough, because the restart path builds that list again, from saved state, in code you tested least.
So the check happens in the corridor instead. Every tool call goes through one place, and that place asks whether this agent is allowed this tool, no matter which door it came in by.
Count the refusals before you start refusing
Switching a new restriction on blind is how you break a working feature for a customer at two in the morning. So the rule runs in two modes. In the quiet mode it notices a violation, counts it, and lets it through. In the strict mode it blocks it.
You run the quiet mode for a week and watch the counter. Zero means it is safe to switch on. Not zero means you have just learned something about your own system that reading the code would never have told you.
What to check in your own agent
Five things. The first one takes ten minutes and is the one that will bite you.
- Run two servers and approve something. Not one. Two. If it works, check you were not accidentally being sent back to the same machine every time.
- Make the unsafe setting impossible. A default that is wrong in production is an incident with a date on it.
- Approve something tomorrow. Leave a gate open overnight and check the restarted job uses today’s date, not the one it stopped on.
- Check the side door. Restarting a job is different code from starting one.
- Log the settings at startup. One line. It turns a question into a glance.
None of this is really about models, which is the part worth taking away. The hard bits of putting an agent into production are the bits that were hard long before agents existed: work that outlives a single request, work that moves between machines, and a person in the middle who takes their time.