A demo goes well. The model answers general questions cleanly, the room relaxes, and then somebody from operations asks what happens when a donor gives to a fund that closed in June. The answer comes back fast, confident, and wrong.
That is not really a failure of the model. Nobody told it. Everything your system needs to know about refund windows, fund designations, approval thresholds, and the six exceptions your controller carries around in her head has to get in there somehow. There are three ways to do that. They cost very different amounts to keep running, and a lot of the confusion in AI buying comes from vendors using the same words for all three.
The three ways to give a model your context
Put it in the instructions
The simplest option is to write your rules into the prompt itself. Every time the system asks the model something, it sends your policy along with the question.
This works, and it works better than people expect for small, stable rule sets. A dozen lines about tone, escalation, and what the system is never allowed to promise will do real work. The trouble starts when the rules grow. You cannot paste a two hundred page policy manual into every request, and the version that matters ends up living in whichever prompt someone edited last, usually without telling anyone.
Look it up when the question is asked
The second option keeps your documents in a searchable index. When a question comes in, the system finds the few passages most likely to answer it and hands those to the model along with the question. People call this retrieval, or RAG.
This is what most operational systems actually need. The model stops guessing and starts reading. When your refund policy changes, you edit the policy document and the next answer reflects it, with no retraining and no developer in the loop. Better still, the system can show which passage it used, which matters enormously the first time somebody disputes an answer.
Train it into the model
The third option is fine-tuning. You take a base model and adjust it with many examples of the input and the output you want, until the behavior you are after is built into the model itself.
Fine-tuning is good at form. If you need every response to follow your claims format exactly, or to sort support tickets into your own twelve categories, or to write in a voice your compliance team has already signed off on, examples teach that better than instructions do. What fine-tuning is not good at is facts that change. A model trained on last quarter's fee schedule will keep quoting last quarter's fee schedule, pleasantly, until you pay to train it again.
What each one costs to own
Buying decisions get clearer when you stop comparing capability and start comparing maintenance. The useful question is not which approach is strongest. It is who edits what, how often, once the project team has gone home.
| Approach | Best at | Who keeps it current | What goes wrong |
|---|---|---|---|
| Rules in the prompt | Small, stable boundaries and tone | A developer, in code | Rules sprawl, versions drift, nobody is sure which prompt is live |
| Retrieval | Answering from documents that change | Whoever owns the document | Weak answers when the source library is stale or contradicts itself |
| Fine-tuning | Consistent format, classification, house voice | A vendor or an ML engineer, on a cycle | Confident answers built on facts that quietly expired |
The same question, three ways
Say a service rep asks the system whether a customer qualifies for a refund on a partial shipment received forty days ago.
With the policy pasted into the instructions, the answer is only as good as whatever text went in, and nobody can tell you when it went in. With retrieval, the system pulls the returns policy and the exceptions memo, answers from both, and shows you the paragraph it used. With a fine-tuned model, the answer arrives in your house format and sounds exactly right, and whether it is right depends on whether the forty-day rule existed when the training examples were written.
In a regulated operation, that difference is the whole ballgame.
If you cannot point to the document an answer came from, you cannot defend the answer.
The reflex to watch for
The most common thing I hear in early conversations is some version of "we will just train it on our data." It is an appealing sentence. It suggests the knowledge goes in once and stays put.
Most of the time it is the wrong instinct, and it is expensive in a way that shows up late. Facts that change on a schedule belong somewhere a person can edit on that schedule. Fee tables, eligibility rules, program designations, approval limits, and vendor terms all move. Putting them inside a model means a change request every time the business changes, and there is no clean way for an auditor, a board member, or an irritated customer to see where an answer came from.
The serious systems I work on usually end up with a mix. Retrieval for anything that is documented. A short set of instructions for boundaries and escalation. Fine-tuning only where consistency of form is worth paying for, which is less often than vendors suggest.
Questions worth asking before you sign
- When our policy changes on a Tuesday, who makes the change, and when does the system start giving the new answer?
- Can the system show the source behind any given answer, and is that source something we control?
- What happens when two of our documents disagree, which they will?
- If we part ways, do we keep the useful part, or does it live inside a model we do not own?
- What does it cost to keep this current in year two, in dollars and in whose time?
A vendor who answers those five plainly is worth more of your attention than one with a better demo.
The hard part is not the technology
Here is what surprises people. Choosing among these three is usually the easy call, and a competent partner can make it in an afternoon once they understand the work.
The hard part is that your policies live in eleven places and quietly disagree with each other. The handbook says one thing, the intake form enforces another, the controller has a standing exception for two large accounts, and the real rule exists only as something everyone learned by watching. No AI approach fixes that. It makes it visible faster, and usually in front of a customer.
So the first useful step is almost never a model. It is getting the actual rules written down, with an owner and a place to live.
If you are weighing an AI system for an operation where the answers have to hold up, that is the kind of thing we work through in discovery. It is a short, structured engagement that maps how the work really runs and what it would take to support it properly, before anyone writes code. You can start that conversation by clicking on the button below. Start With Discovery