Last March, a quiet model update broke one of our lead intake workflows, and it took us nine days to notice. No errors. No failed executions. The workflow ran green the entire time, cheerfully routing roughly a third of our new leads into the wrong queue.
The deal we lost over it was worth €4,200. The prospect got a reply four days late, from the wrong person, and politely signed with a competitor instead. That was the day prompts stopped being config in my head and became code. Code gets tests. Prompts, up to that point, had gotten vibes.
This post is the setup we landed on after that mess: a prompt regression suite inside n8n that costs us nothing extra to run. It takes the same discipline we already apply to workflow reliability on the API side and points it at the fuzzy, probabilistic part of the stack. The whole trick, honestly, is assuming a green execution tells you nothing about whether the output was correct.
Why prompt regressions hide so well (and cost so much)
A code bug announces itself. Something throws, something goes red, you get an alert at 2am. A prompt regression does the opposite. The workflow completes, the JSON parses, and the output comes back fluent, confident, grammatically perfect, and completely wrong.
Non-determinism makes it worse. You can’t diff last week’s responses against this week’s, because the model never produces the same output twice anyway. Maybe it just felt different on Tuesday. There’s no stack trace for a model suddenly deciding that ‘I want to cancel my subscription’ is a sales inquiry.
So your monitoring becomes your customers, which is the worst possible arrangement. The ones who noticed got weird replies. The ones who didn’t notice just went elsewhere. Nine days of that.
Fixing the prompt after the fact was easy. Believing the fix would hold forever was the actual mistake, and it’s the same trap I wrote about in why teams keep blaming prompt engineering for AI failures. Providers ship point updates. You tweak a system prompt. Someone adds ‘just one more’ few-shot example. Drift is normal. What’s usually missing is anything that notices it.
The stakes keep climbing, too. Most of the AI automation we build for SMB clients starts with an innocent little classification step. But once clients move toward agents that take actions instead of just labeling things, a drifted prompt doesn’t misfile a ticket anymore. It emails the wrong person or quotes the wrong price. I wanted a gate in front of that, not an apology behind it.
The harness: n8n’s Evaluate node and frozen input/output pairs
We self-host n8n on a small Hetzner box we were already paying for, which is a big part of why the suite lands at $0 marginal cost. If you’re still weighing platforms, the comparison we keep updated between Zapier, Make, and n8n covers the cost angle, and it matters double for testing: task-based pricing gets expensive fast when a suite fires a few hundred executions before breakfast.
The setup is genuinely boring, which I mean as a compliment.
The frozen pairs
Our golden examples live in a Google Sheet. Eighteen rows, each holding the raw input, the expected classification, and two or three ‘must contain’ phrases. Why a Sheet? Because it’s free, non-engineers can edit it, and n8n reads it natively. The ‘frozen’ idea started with n8n’s pinned data in the editor, which is still the fastest way to replay one weird input by hand, and grew into a full dataset once we needed coverage instead of curiosity.
The run itself
The test workflow pulls the Sheet, loops the rows, and feeds each input to the exact production workflow, not a copy. This matters more than it sounds. For a while we tested against a duplicated version, and it turns out the copy quietly drifted from the real thing. Test what you ship, or you’re testing a lie.
n8n’s evaluation framework does the heavy lifting. The Evaluate node runs each frozen input through the workflow and scores what comes out, and you can seed the dataset straight from past executions. That’s how our golden set actually started: I pulled the ten weirdest real leads out of the execution log and froze them as-is.
Scoring happens on two levels. Hard assertions in a Code node: does the output contain the expected category, does the JSON parse, does it avoid the forbidden-phrases list (things like promising refunds, which our intake bot has no business doing). Then a soft score from an LLM judge with a tight rubric: given this input and this expected answer, rate whether the actual output conveys the same intent. Anything under 0.8 fails the row.
I was skeptical of LLM-as-judge, and honestly I still am a bit. For tone it’s mush. For classification with a fixed answer key and a strict rubric, it’s consistent enough that I sleep fine. If you want a v1 this week, the hard assertions alone catch most regressions and cost nothing to argue about.
The golden set rule: you need fewer examples than you think
Opinion time. For an SMB workflow, anything past 20 golden examples is mostly ceremony. We run 18. Somewhere along the way, teams started insisting you need hundreds of evals, and maybe you do for a consumer chatbot with a million users. A lead intake flow for a 40-person company does not. It needs the weird cases.
Ours: the message with two intents jammed into one sentence. The angry all-caps one. The Dutch one, because cross-border leads happen. The out-of-scope ‘is this even for me’ question. The message that’s just an emoji and a phone number. Twelve of the eighteen came from things that actually went wrong in production at some point.
My first attempt was a 200-row monster I built over a weekend because a blog post said more data equals better evals. Half the rows were near-duplicates, a third had stale expected outputs within a month, and every suite run took three times longer. I deleted the whole thing and started over with twelve hand-picked examples. I’m still a bit annoyed about that weekend. And I’m genuinely not sure the current 18 are the right 18; every so often an edge case slips through and I have to decide whether it deserves a golden slot or whether I’m just teaching the test to parrot the past. No clean answer yet.
Wiring it into deploys without buying anything
n8n’s built-in source control is an Enterprise feature, so the free path is slightly manual, but it works. We export the workflow JSON with the CLI, keep it in a plain GitHub repo, and treat every prompt or node change as a commit. Before a change goes live, the suite runs against the updated version. If a golden fails, the change doesn’t ship. That’s the entire pipeline, and it’s pretty much the same cheap-gate philosophy behind the idempotency checks we run in front of data pipelines: a small wall before things move, so money and messages don’t leak.
Results land in a Slack channel via webhook, and Slack’s free tier is plenty. Total spend on tooling: €0. Total new infrastructure: one subfolder in a git repo and one extra workflow. If your current deploy pipeline is ‘Damian remembers to test things’, this is a strict upgrade. Though I’ll admit ours spent a few weeks as ‘Damian remembers to run the suite’ before the cron made it automatic, which is a bit embarrassing for a company that sells automation.
The 5-minute morning routine
Provider-side drift doesn’t wait for your deploys, so the suite also runs on a cron at 6:50 every morning. By the time the coffee’s done there’s a Slack message that reads either ’18/18 green’ or something like ‘3 failed: #7, #12, #15. Cancel my subscription classified as sales inquiry.’
On a Tuesday in April, that message saved us. The provider had quietly shipped a point update overnight, no email, no changelog entry we could find, and the classifier drifted with it. I saw the failure at 6:54, reproduced it by 6:58, adjusted the system prompt at 7:15, and had the suite green again before 7:30. Customer exposure: zero, because the drift got caught before the first real lead of the day touched the workflow.
Nine days of silent misrouting in March versus four minutes of diagnosis in April. A full suite run costs about $0.02 in API tokens, and yes, I checked, because I wanted the ‘$0’ claim to survive contact with an invoice. The tooling is free. The tokens are a rounding error next to the coffee I drink while reading the report.
What this setup won’t fix
A regression suite won’t make a bad prompt good. It keeps a good prompt from quietly rotting, which is a different and underrated job. It also can’t catch requirements you never wrote down. We learned that in June, when a customer complained about a response style none of our 18 goldens covered, because style was never something we asserted in the first place. That one’s still open, for what it’s worth.
Still, the math feels lopsided in our favor. One lost €4,200 deal paid for this suite a few thousand times over, and the whole thing took an afternoon once I stopped overcomplicating it. If you run AI workflows for a business, steal it. Golden Sheet, Evaluate node, hard assertions, an optional judge, cron, Slack. That’s the full bill of materials.
If you’re new here, this page explains who I am and what StartMit does.
Talk soon,
Damian
