NoteBuilding Dohos6 min read

The model should never own business state

An early restaurant test taught me that sounding convincing and taking a correct order are different problems.

I switched off my father's phone line one morning and let an AI answer it instead.

It lasted a few hours. In that time it took orders for things the kitchen does not make, quoted prices that do not exist, and agreed to combinations the store does not allow. The calls sounded convincing. The AI was warm, it was fast, and it never got flustered.

Then the tickets came out of the printer and none of them were real.

What it actually cost

The money was the smallest part. Comped orders and remakes for one morning is a number you can absorb.

What I had actually spent was credibility. My father’s store had told regulars yes, only for a human to tell them no twenty minutes later. I had asked my father to trust me with the front of his business, and the system had let him down. The callers had been promised something the restaurant could not deliver.

That third one is the one nobody budgets for. Every bad automated interaction spends down a shared account, and it is the reason “it demoed well” is not an acceptable standard.

The mistake was structural, not a prompting problem

Here is what I had built first. Speech to text, a language model in the middle, text to speech on the way out, and a database write at the end. Give the model the menu in its prompt. Tell it to be accurate. Ship it.

The reasoning seemed sound. I had given the model the menu. But including it in a prompt was not the same as making the system validate an order against it. A request for half mushroom, half onion, uncut, with jalapeños on the whole thing could get a friendly yes, even if that store did not support those combinations.

And people order in wildly different shapes. “Hey! Can I get a large pizza? Can I get it half mushroom, half onion? And can you add jalapeños on the whole thing?” versus “Hello. Can I get a large jalapeño pizza with half onion and half mushroom, but don't cut it.” That is the same pizza. A human handles both without thinking. A language model also handles both, which is exactly the problem, because it will handle the second one into an order that cannot be made.

I tried to fix it with a better prompt. I listed the rules. I told it what it could not do, in bold, twice. The prompts helped, but the failures kept moving. I needed the important rules enforced outside the conversation.

Give the conversation freedom. Give the order rules.

For this system, giving the model more freedom was not solving the problem. I needed to be precise about which decisions belonged in the conversation and which belonged in application code.

The way to make an AI system credible is to constrain it.

Not to make it dumber, and not to script the conversation. The conversation should stay completely free, because that is the half the model is genuinely extraordinary at. The model owns understanding a caller who changes their mind mid-sentence, handling interruptions, recognising that “actually make that two” refers to the thing mentioned eleven seconds ago. Deterministic code owns what exists, what it costs, what is legal, and whether anything actually happened.

In my system, the model proposes changes using references to catalog items. Application code resolves those references, checks whether the operation is supported, calculates the price, and applies the change. The conversation can be flexible while the application checks what actually happens to the order.

I wanted the important rules enforced in code, where I could test them.

What this costs, honestly

It fails closed. If a caller asks for something the catalog cannot describe, the system does not improvise. It says it cannot do that, or it routes to a human. This means my agent sometimes refuses things a person behind the counter would have said yes to. In a demo that looks like a limitation. In production it is the entire point, because a refusal is recoverable in ten seconds and a silently wrong order is not recoverable at all.

A concrete case I like, because it is small. Somebody says “no mayo.” That is not a priced modifier and it is not on the menu. The lazy options are to drop it silently or to treat it as a topping and charge for it. Both are wrong in ways that surface later at the counter. So it stays attached to the line as an unpriced note and routes the order to manager review. The system does not pretend to know what it does not know.

A few other boundaries fell out of the same principle. Any material change to an order invalidates whatever confirmation came before it. Partial speech can trigger a lookup but can never authorise a write, because reads are free to be wrong and writes are not. An interruption stops the agent talking, but it does not by itself cancel a business action, because those are two different events and treating them as one loses orders.

The general version

If you are putting a model in front of anything transactional, the useful question is not “how do I stop it hallucinating.” It is “which categories of truth must never be model-owned.” My list, after a year of this: price, availability, policy, payment, acceptance, and success. The model may discuss all six. It may assert none of them.

Everything else can stay free, and it should. The tone, the phrasing, the recovery when somebody changes their mind, the ability to sound like a person rather than a phone tree. That is the half worth having, and constraining the other half is what makes it safe to ship.

What I am still not certain about

Every failure mode in this system is one I thought of. That is the limit of it. The failures that will matter are the ones a few hundred real callers find in the first month, and by definition I do not know what those are yet. Bounded proof calls are not a Friday dinner rush.

And the claim I would push hardest on if I were reading this is the catalog one. Configuration living in data rather than code is what makes a second restaurant a data-entry job instead of a rewrite. I believe that. It is proven on one menu. A store with a genuinely different structure is the real test and I have not run it.

Dohos is now in live testing at a restaurant and has handled about 2,000 real calls. The next test is what breaks when it runs at more locations.