A client forwarded me a proposal last month for an “autonomous AI agent” that would run their entire customer follow-up. Fourteen pages. Gorgeous slides. The word “harness” appeared zero times. Neither did “retry,” “permissions,” or “what happens when this fails.” What did show up, on roughly every third page, was which model powers the thing.

That proposal is why I’m writing this. If you’re shopping for an AI agent for your small business, you’re shopping at a strange moment: Microsoft just made its agent harness generally available, and Google shipped a faster workhorse model within days. Read the two announcements together and you get the real news, the part nobody puts in a press release. The model is becoming the commodity part of an agent. The harness, meaning the layer that plans steps, picks tools, holds context, and enforces permissions, is the part that decides whether the thing works.

Most SMB buyers I meet still evaluate the model and ignore the harness, which is kind of like test-driving a car by revving the engine in the parking lot. I’ve written about what actually separates an agent from a chatbot before, but the buying side needs its own checklist. Five questions, plain language. Bring them to the demo.

Why the model is the cheapest part of your agent

I’ll admit something embarrassing first. The first “AI agent” I ever paid for cost $240 a month and turned out to be a chatbot in a trench coat: one of those docs-search tools that plugs into your Notion, answers questions about your wiki, and calls itself an agent on the pricing page. It couldn’t remember anything between sessions. It couldn’t take a single action beyond replying. I demoed it for twenty minutes, got charmed, swiped my card, and needed six weeks to notice what I’d actually bought. I wanted an agent so badly that I saw one where there wasn’t one.

That was 2023. I’d love to tell you I’m immune now. Mostly I am, but vendors have gotten much better at blurring the line, and the blur always hides the same layer.

Quick definitions. The model is the thing that reads text and writes text: GPT-4o, Claude, Gemini, whichever. It’s smart, it gets cheaper every quarter, and by the time you sign, it’s probably not your differentiator, because your competitor’s vendor rents the same one. The harness is everything wrapped around the model. It plans the steps, picks the tools, carries the context, enforces the permissions. Engine versus the rest of the car, except everyone in this market rents the same engine at the same falling price.

Some numbers from a bake-off I ran last autumn for a 16-person freight brokerage. Two vendors, both selling “AI agents” for reconciling carrier invoices, both running the same model underneath (GPT-4o, for what it’s worth, which is exactly the point). I loaded 220 real invoices into each: three decimal formats, two PDF templates, and one carrier who emails photos of handwritten delivery notes because of course they do.

Vendor A’s agent finished 131 clean. Then an OCR call came back empty, the agent asked the model what to do about it, and the model decided to skip the row and keep going. Vendor B hit the same OCR failure and fell back on a rule its harness enforced: park the row, flag it, notify a human. It finished 214 of 220 untouched, six flagged for review. Same model, same task, eighty-three invoices of difference, all of it in the layer neither pitch deck mentioned.

Question 1: Who plans the steps, the model or the harness?

In the demo, ask them to run a task with at least ten steps. Then stop the show: who decided step four? Does a plan exist somewhere outside the model’s head, or did the model freestyle its way there?

Both answers can be legitimate, honestly. Some of the best agents I’ve seen let the model plan, but the plan gets checked: each step verified against an allowlist, states written down between steps, weird behavior caught mid-run. The answer that should worry you is the shrug. A sales engineer once told me, cheerfully, “the model just figures it out as it goes, that’s the beauty of it.” He believed it. His product, by his own admission in that same call, had no concept of a step at all. A model, a menu of tools, and hope.

Model-planned steps are flexible and occasionally brilliant, and inconsistent: the agent that handles your weird invoice today may skip it tomorrow because the dice landed differently. Harness-planned steps are boring and repeatable, and boring is what you want in accounts payable. The same logic applies if you’re building in-house instead of buying, by the way. When I compared Zapier, Make, and n8n on their plumbing, the differences that mattered weren’t connector counts. They were how each platform sequences, branches, and retries steps.

Question 2: What happens when a tool call fails?

APIs fail constantly. Rate limits, timeouts, a SaaS vendor ships a schema change on a Friday afternoon. Demo wifi never shows you any of this, which is convenient for everybody except you.

So ask to see a failure, live. Revoke a token mid-demo, or just ask plainly: “Show me what a 429 does.” A chatbot with tools bolted on will either crash the whole run or, worse, let the model improvise past the failure. I’ve watched an agent cheerfully confirm “your email has been sent!” while the API call sat there dead. The model wasn’t lying, exactly. It was doing what models do, which is produce plausible text.

A real harness has the answer before you finish asking: retries with backoff, idempotency so a retry can’t double-fire, a dead-letter queue, an alert when a run dies. Most of my day job is closer to workflow reliability engineering than to prompt writing, and that ratio isn’t an accident.

My favorite scar tissue, so you can buy the lesson cheaper than I did. In early 2024 I built an agent that sent quotes through a client’s HubSpot CRM. One afternoon the API started throwing 429s, and my harness retried three times in fast succession because I hadn’t configured backoff. My fault, all mine. Turned out the first call had landed, just slowly. A few customers got the same quote twice, one got it three times, because a webhook I’d forgotten about was also retrying in the background. The model had been perfect all day. The harness handed out duplicate quotes anyway. Across the eight months that workflow ran, the model misbehaved twice. The plumbing misbehaved dozens of times.

Question 3: Where do permissions actually live?

Ask this one directly: “If I told the agent to delete a customer record, what stops it?” If the answer involves the model choosing not to, you’re being sold a chatbot with good manners. Prompts are advisory. Models drift, users jailbreak, and a sufficiently creative request can talk its way past a paragraph of instructions saying “you may only do X.”

Permissions belong in code: scoped API tokens, per-step tool allowlists, destructive actions that need a human’s click, and an audit log you can actually read. The harness should make it impossible for the agent to touch something it was never granted, not merely discouraged. This is the question vendors dodge hardest, in my experience, because the honest answer is often a bit embarrassing. We’ve written a whole piece on why your agent strategy can quietly become a security liability, and not one failure mode in it comes from the model. They all come from the wrapper.

For a small business this isn’t paranoia, it’s arithmetic. Your agent will hold your Stripe keys, your CRM, your customers’ inboxes. One bad afternoon with write access costs more than a year of slightly dumb replies.

Question 4: How does context survive longer than one chat?

Close the chat window and a chatbot forgets you exist. An agent is supposed to be different, and this is where the trench coat slips most visibly.

The test is almost too simple. Email the agent on a Monday, close everything, come back Thursday in a fresh session, and ask what you two talked about. Then ask where that memory lives. “In the chat” is a chatbot answer. A database, a job queue, a customer record the agent reads before acting: those are agent answers. Conversational memory is only half of it, too. If the agent starts a task Tuesday and the server restarts Wednesday, does the job resume or evaporate?

A customer who wrote in three weeks ago isn’t a stranger, and an agent that treats them like one will feel off in a way customers can’t name but will remember. I’ve watched demos where the agent recalled the demo conversation beautifully, which proved nothing, because the memory was the demo conversation. Turns out remembering something for fifteen minutes and remembering it for fifteen days are different engineering problems, and only one of them shows up on a sales call.

Question 5: What does it cost when it’s wrong?

Vendors love answering the question you didn’t ask. Ask about failure costs and you’ll hear accuracy rates. “It’s 95% accurate” is a deflection, not an answer, and honestly the number is kind of meaningless until you know what the other 5% costs.

The math to ask for

Make them do the arithmetic with you, in the meeting. If the agent runs 300 sends a month and saves you 20 hours, what does one bad run cost? A wrong invoice, an email to the wrong customer, a double charge: put a dollar figure on it, then ask how the harness contains that blast radius. How many runs get human review before the agent goes fully autonomous? What gets logged, and can you undo a run after the fact? A vendor who has shipped real agents has answers, or at least a straight face while writing your question down. A vendor selling a chatbot steers back to the model, every time.

An uncomfortable admission: I still don’t fully know how to price model risk against harness risk for every use case, and I’ve stopped pretending I do. Low-stakes work like drafting social posts? Let the model wander. Anything touching money, legal, or a customer’s inbox? The harness does the heavy lifting, and I need to be able to explain why in one sentence on the day something breaks.

The cheap part and the part that ships

None of this means models don’t matter. A great harness around a bad model still produces a mediocre agent, and I’ve seen rigid harnesses turn brilliant models into very expensive form-fillers. But when two vendors pitch you in the same week, the models will be pretty much the same and the harnesses will not, and only one of those differences shows up in the invoice the way it should.

So bring the five questions to the demo, not to the contract negotiation. Ask them early, while the sales engineer still thinks you’re an easy meeting. Watch the face when you ask what happens when a tool call fails. That expression tells you more than fourteen pages of slides ever will.

If you want the build-side version of this thinking, our no-hype playbook for AI automation in small businesses covers what to automate first and what to leave alone for now. And since you made it this far: I’m Damian. I build and break these systems for small businesses for a living, and the breaking is the more useful half.

Related reading

Talk soon,
Damian

Leave a Reply

Your email address will not be published. Required fields are marked *