Last March, halfway through a live demo for a client’s ops team, their OpenAI key hit a hard spending cap and the workflow died in front of nine people. Finance had set the cap at $20 a month. We’d burned $18.40 of it during test runs the night before, and at 2:15 on a Tuesday afternoon the lead-scoring demo 401’d on the projector, mid-sentence.

That story is usually a warning about billing governance. But when n8n shipped v2.36 with the n8n AI Gateway, I realized it’s just as much a story about setup friction. The pitch: buy credits inside n8n, call models from one catalog through one billing relationship, skip the provider accounts, the prepaid invoices, and the Tuesday-afternoon humiliation.

We build most client stacks on n8n, and our comparison of Zapier, Make, and n8n is going to need an update after this release. So instead of poking at the feature for ten minutes, my team and I stress-tested it properly. 2,400 production-shaped prompts. Three model families. Two concurrency levels. Direct API keys as the control group.

This is the honest breakdown, including the parts where credits quietly cost more.

What n8n AI Gateway credits get you in v2.36

The mechanics take thirty seconds to explain. Add the Gateway credential, buy a bucket of credits, and every LLM node or AI Agent node in your workflows can run against the Gateway’s catalog. No OpenAI account, no Anthropic account, no Mistral account, no four separate invoices that a client’s finance team has to approve in four separate meetings. We lean on agents constantly (more on how AI agents actually work in production), and spinning one up without a single vendor signup is genuinely appealing.

For small and mid-sized businesses, this solves a real problem. Half our clients have exactly one person who ‘owns’ the OpenAI account, and that person goes on vacation.

As of the build we tested, the catalog covered what you’d expect: GPT-4o and GPT-4o mini, Claude 3.5 Sonnet and Haiku, Gemini, Llama 3.1 70B and 405B via hosting partners, plus Mistral and DeepSeek. Caveat before I go further: n8n ships fast, and I’d bet money parts of this list have already changed. Check the current docs before you plan around my snapshot.

What you still can’t reach through the Gateway

No Azure OpenAI. No AWS Bedrock, no Vertex AI, no private endpoints, nothing self-hosted like Ollama or vLLM, and no fine-tuned models. If your architecture depends on any of those, the Gateway isn’t an option, full stop. Remember that gap. It comes back later in this post with teeth.

Our test setup: 2,400 prompts, three families, two concurrency levels

Fair warning: this wasn’t a lab. We pulled four task types from real client jobs and anonymized them: support-ticket classification, lead-message extraction, document summarization, and short-form generation. The average prompt ran about 750 input tokens and 180 output tokens. 800 prompts per family across GPT-4o, Claude 3.5 Sonnet, and Llama 3.1 70B, and every prompt executed twice, once through Gateway credits and once through our own keys, in the identical workflow with only the credential swapped.

Half the runs went out serially. The other half ran at 10 parallel executions, which is roughly what a busy SMB workflow does on a Monday morning. For context, we’re a two-person shop, and the longer version of who’s writing this and how we work with clients lives on the about page.

Now a confession, because otherwise the numbers mean nothing. The first 400 executions had prompt caching enabled on the direct side but not on the Gateway side, because I configured the two credentials on different days with different amounts of coffee. Direct API came back looking about 30% cheaper than reality. My error, entirely. I threw the batch out and reran it, then sat there for a while wondering how many benchmark posts on the internet carry exactly this kind of drift and never notice.

One more confession. I went in rooting for the Gateway to fail. I keep a credential folder with 40-odd carefully labeled API keys and take a strange pride in it. Wanting a thing to lose is bad science, so halfway through I had my colleague Miriam recheck the cost math without telling her which column was which. She found one arithmetic error. Also mine.

Cost math: credits vs. direct API billing

Blended across all 2,400 prompts, the Gateway charged us $13.87. The identical workload on our own keys cost $11.42. That’s a 21.4% premium, which, turns out, is lower than I expected going in.

Model familyDirect API (800 prompts)Gateway credits (800 prompts)Premium
GPT-4o$5.90$6.95+18%
Claude 3.5 Sonnet$4.10$5.05+23%
Llama 3.1 70B$1.42$1.87+32%

The pattern worth staring at: the cheaper the model, the bigger the percentage. My guess is there’s a fixed per-call routing overhead baked into credit pricing, which lands proportionally harder on cheap models. I can’t prove that. n8n doesn’t publish pricing internals, so it stays a guess.

Where the premium actually lands depends on volume. If you map this onto the usage tiers in our AI automation playbook for SMBs, the brackets look like this:

  • Low volume (around 5,000 prompts a month): internal tools, prototypes, solo founders. The premium works out to roughly $5 a month, about $61 a year. That’s less than the hour of vendor onboarding you’d otherwise spend. Credits win easily.
  • Mid volume (around 60,000 a month): production SMB workflows. The premium lands near $61 a month, roughly $735 a year. Borderline. If the client manages their own instance and you want zero billing entanglement, it’s defensible. If you’re eating the cost inside a fixed retainer, it stings a bit.
  • High volume (400,000+ a month): the premium crosses $400 a month, close to $4,900 a year. Bring your own keys, no contest. Direct APIs also unlock batch pricing at that scale, 50% off on OpenAI and Anthropic batch endpoints, and as far as we could tell the Gateway doesn’t pass those discounts through. I’d love to be wrong about that one, because cheap batch plus zero setup would be lovely, but our batch runs through the Gateway billed at full rates.

Latency reality: the 1.4-second tax on every Gateway call

Every Gateway call came back slower than its direct twin. Median overhead was 1.4 seconds per call, pretty consistent across families, ranging from 1.1 seconds on GPT-4o to 1.7 on Llama. GPT-4o’s p50 response time went from 2.3 seconds direct to 3.7 through the Gateway. At 10 concurrent executions, the p95 overhead ballooned to 4.1 seconds, which smells like queueing on the routing layer.

Error rates stayed small on both sides: 11 failed calls out of 2,400 on the Gateway versus 5 on direct keys, all recovered on retry. Annoying, not alarming.

Where 1.4 seconds actually hurts

A nightly enrichment job chewing through 400 records barely notices. You add about nine minutes to a job nobody watches. Internal tools where a human clicks a button and waits, also fine.

Two places it genuinely bites. First, anything user-facing: a customer-facing chatbot already takes 1.5 to 3 seconds to first token, and another 1.4 on top makes it feel broken. Second, agentic workflows. Multi-step agents make five to eight model calls per run, so a six-call agent picks up 8.4 seconds of routing delay every single run. Slow calls also sit closer to timeout ceilings, and if you’ve ever watched retries cascade on a Monday morning, you know latency is a reliability problem wearing a costume. Our notes on workflow reliability engineering go deep on that failure mode.

The data residency caveat that killed an EU pilot

This one stung.

A German insurance brokerage, about 40 staff, hired us for a pilot: triage and summarize incoming claims emails before they reach the adjusters. Two weeks of build. Everything worked. The adjusters liked the summaries. Then their data protection officer asked one question on the review call: ‘Where exactly are these prompts processed once they leave our n8n instance, and can you put that in writing?’

With direct keys, the answer is easy. Azure OpenAI in Sweden, or Bedrock in Frankfurt, regional processing guarantees written into the contracts. With Gateway credits, at least as of our testing, it wasn’t. The docs couldn’t give a binding EU-processing guarantee for credit-routed calls, and when we asked support directly, the answer amounted to US-managed routing with EU residency not yet contractually promised. The DPO said no. Pilot dead, roughly €6,000 of planned scope parked until we rebuild it on Azure keys.

To be fair, none of this is n8n’s fault in any dramatic sense. Somebody else’s routing means somebody else’s regions; that’s the structural cost of the convenience. But it’s exactly the kind of caveat that never makes the launch announcement, and for any client under GDPR with sensitive data, it’s the whole ballgame. The Gateway is simultaneously the fastest way to start building and the slowest thing in the room during a security review.

Verdict: when credits win, and when bring-your-own-keys is non-negotiable

After three weeks of spreadsheets, here’s my split.

Use Gateway credits for: prototypes, client demos, workshops, internal tools under roughly 10,000 prompts a month, and anything where setup friction is the actual enemy. That Tuesday-afternoon demo failure from the intro? One billing relationship, one low-balance alert, and it doesn’t happen. For a solo founder who’d rather build than do vendor onboarding, credits are kind of great.

Bring your own keys when: the data is regulated or needs contractual EU residency, the workflow is user-facing and latency-sensitive, volume crosses 50,000-ish prompts a month, or you need fine-tuned models and Azure, Bedrock, or private endpoints. In production for paying clients, that’s us most of the time.

I expected to write this as a hit piece, and honestly, I’ll own that. The Gateway beat my expectations on cost for small usage, and the premium is often cheaper than the unbilled setup hours it replaces. The uncomfortable truth for my credential folder: for most of what SMBs actually build, credits are good enough. I’ve already moved two of our internal tools onto them. Client-facing stuff stays on keys.

Oh, and I finally set a proper balance alert on that OpenAI account. Only took a year.

Talk soon,
Damian

Related reading

Leave a Reply

Your email address will not be published. Required fields are marked *