The Slack alert hit my phone at 2:14am on a Sunday: ‘Ticket #48712 has been unrouted for 41 hours.’ Forty-one hours. A customer named Marion had written in about a double charge, our n8n workflow was supposed to sort her into the billing queue, and instead her ticket slipped out the bottom of a Switch node because I’d never wired up the ‘no match’ output. Nothing errored. Nothing pinged. The workflow just quietly moved on.

That night I stopped defending the way we’d been building routing workflows. So when n8n pushed its native Agent node and my feed filled up with people declaring conditional chains dead, I did what I usually do: built both, pointed live traffic at them, and waited to find out whether I was wrong. I’ve been openly skeptical of the agentic workforce pitch for a couple of years. This seemed like a fair way to check whether any of it survives contact with a real job.

This isn’t a feature tour. It’s a 72-hour head-to-head on live support traffic, with numbers from a real client inbox. Call it an n8n AI agent review if you like. I think of it as the test I should’ve run a year ago.

The old chain: 14 nodes, 8 conditionals, and the ‘where did this branch break’ problem

The workflow in question has been routing support tickets for a 40-person SaaS client since early 2024. We built it in n8n (they picked it over Zapier and Make after working through our comparison of automation platforms, mostly for the Code node and self-hosting), and on paper it looked tidy. Webhook trigger. A Code node that cleans subject lines. An IF for language, a Switch with six keyword outputs, three more IFs for VIP customers and sentiment, then four routes into the helpdesk queues.

Fourteen nodes. Eight conditionals. And a maintenance log that reads like a diary of small embarrassments.

Six months in, I’d made 23 edits to that workflow. Eleven were new keyword branches (‘add chargeback’, ‘add can’t-log-in with the apostrophe’, ‘add the German word for invoice’). That’s roughly nine hours of my billable time, and nearly all of it came after something had already broken. My favorite bug, if that’s the word: the Code node stripped ‘Re:’ from subject lines, but it ran once, so a reply to a reply arrived as ‘RE: RE: invoice question’ and the keyword match failed downstream. Two hours of my life for a regex. I still think about it.

The real problem was never node count. Every branch is a prediction about what humans will type, and humans type creatively in the worst ways. You patch one phrasing, three new ones show up. The chain stays exactly one email behind reality, forever.

The new agent: one node, a system prompt, and three tool schemas

Building the agent version took about three hours, and two of those were the system prompt. The Agent node gets roughly 450 words describing the four queues and the client’s tone, plus three tool schemas: route_to_billing, route_to_technical, route_to_general. Each tool takes the ticket text, a confidence score, and a short reason string. Churn risk isn’t a fourth tool, it’s a flag on whichever lane fires. We pointed it at Claude Haiku because it’s cheap and fast, which matters when a few thousand conversations come through.

Turns out the tool descriptions matter more than I expected. The route_to_billing schema says explicitly what belongs there (billing disputes, double charges, payment failures, refund requests) and what doesn’t (bugs, even payment-related ones). That one description does more routing work than half the keywords in the old chain ever did.

No keywords. No IF nodes. Read the ticket, pick a lane, explain why.

Honest confession from setup, because it nearly killed the test before it started. My first system prompt included ‘when in doubt, escalate to a human,’ which sounded sensible at the time. On dry-run traffic the agent escalated 14 of the first 19 conversations. Fourteen. I sat at my desk genuinely wondering whether I’d wasted my weekend and whether every skeptical thing I’d ever said about agents was about to be proven right in the most annoying possible way. It was my fault, obviously. The prompt invited it. But for a solid ten minutes I did not want it to be my fault, and I think that’s worth admitting.

72 hours of live traffic: the actual numbers

Method first. We mirrored 100% of the client’s inbound support conversations to both patterns, Friday noon to Monday noon: 2,822 conversations, 1,411 through each side. We measured wall-clock latency, failures (execution errors plus misroutes, caught by hand-auditing a 10% sample), and token spend.

Latency

The chain won, and it wasn’t close. Median execution time came to 2.3 seconds for the chain versus 5.6 for the agent. The p95 gap was wider than I expected: 7.4 seconds against 16.8. I’d assumed the chain’s tail would be cleaner, but there’s a retry loop on the helpdesk API node that occasionally spins, and it drags everything behind it. The agent’s tail is mostly the model thinking, plus a second hop now and then when it fills in a missing tool argument.

Whether anyone cares is a different question. Nobody stands at a counter waiting for a ticket to be sorted. Three extra seconds of latency cost this client exactly nothing, and I’d been treating latency like a scoreboard when, for this particular job, it’s decoration.

Failure rate

Then came the part I actually cared about. The chain failed 54 of 1,411 conversations, 3.8%. The agent failed 17, or 1.2%. Most of the chain’s failures were silent misroutes, tickets sitting in the wrong queue with no error anywhere, which is the worst failure mode there is because nothing pings you. Silent failures are the whole reason we wrote about workflow reliability and failure budgets, and this test turned into a live demo of the argument.

The agent’s failures were different in kind. Twelve were malformed tool arguments that n8n’s validation caught and retried; nine of those recovered on the retry. Five were wrong-lane calls, like the customer who mentioned a bug while canceling and landed in technical instead. Annoying, but visible, and fixable with one prompt line. A misroute you can see beats five you can’t.

Cost

Token spend for the chain: zero. The agent averaged 2,400 tokens per conversation, which came to $6.77 for the entire weekend, about half a cent each. Set that against nine hours of my maintenance time on the chain over six months and the economics get kind of absurd. Caveat before someone else makes it for me: token spend scales with volume, my hourly rate does not scale down, and 72 hours is a snapshot. I genuinely don’t know whether that failure gap holds at ten times the traffic. Anyone claiming certainty from one weekend of data is selling something.

Three scenarios where the old chain still wins

I half-hoped the agent would sweep everything so this post would be shorter. It didn’t, and honestly the exceptions matter more than the averages.

Strict compliance paths. GDPR deletion requests have to reach legal every single time, with an audit trail. ‘The IF node matched the word deletion’ is an answer an auditor accepts. ‘The model decided it looked like a compliance issue’ is not, and no amount of prompt tuning makes it one.

Deterministic approvals. Anything with a hard threshold, like refunds over $150 needing human sign-off, belongs in a conditional. An agent with judgment will bend that rule occasionally, maybe once in 200. ‘Occasionally’ is the word that ends meetings with compliance people. Rightfully so.

Zero-LLM fallback. During the big OpenAI outage in November, this client’s chain kept routing while every model-dependent workflow we run sat dead in the water. An agent is only as alive as your LLM provider. If routing is the thing that must never stop, a dumb chain with no external model dependency is genuinely more robust, and I don’t think that’s a minor point.

The verdict: what we spec for new builds now

For new SMB clients (the broader decision framework lives in our no-hype playbook for SMB automation), here’s the pattern I now write into proposals: agent at the front door, chain behind it. The agent reads the messy human input and picks a lane. The moment a tool fires, a boring deterministic chain takes over: API calls, approvals, escalations, all of it. The agent never touches the helpdesk API directly.

That hybrid gets you most of the agent’s flexibility on intent while the compliance, approval, and fallback paths stay exactly as dumb and reliable as they need to be. One agent node. Three tight tool schemas. A cheap model, hard validation on every argument, one retry, and an alert whenever confidence drops below a threshold.

If your instinct for a fresh n8n routing build is eight Switch nodes and a keyword list, think about Marion’s ticket for a second. The chain will work right up until it doesn’t, and it will fail quietly when it does.

Got a routing chain with its own Marion story? I’d genuinely like to hear it. I’m Damian, and if you want the longer version of who I am and what StartMit does, it’s here.

Talk soon,
Damian

Related reading

Leave a Reply

Your email address will not be published. Required fields are marked *