Three different questions, not one

"Ground the AI in our own data" is really shorthand for at least three distinct problems, and each has a different fix:

  • "Does it know current facts about our business?" - a retrieval problem. Fix: RAG, or a live tool/MCP connection.
  • "Does it apply our specific judgment and process consistently?" - a policy problem. Fix: a Skill.
  • "Does it behave in a narrow, specialized way we can't get through prompting?" - a model-behavior problem. Fix: fine-tuning, rarely needed.
Most "our AI assistant gives wrong answers" complaints are actually a retrieval or policy gap, not a model-capability gap. Diagnosing which of the three problems you actually have matters more than which technique sounds most sophisticated.

RAG: finding the right facts at query time

Retrieval-augmented generation searches a corpus - your docs, your knowledge base, your product catalog - for the passages relevant to the current question, and gives those to Claude as context before it answers. It's the right tool when the corpus is large, changes often, and the question could be about any part of it. The tradeoff is retrieval quality: RAG only works as well as its search step finds the actually-relevant passage, which is why a "RAG-powered" assistant can still confidently answer from the wrong document if retrieval surfaces the wrong one.

Skills: packaging how to apply judgment

A Skill doesn't retrieve facts from a large corpus - it packages the instructions, escalation logic, and reference material for a specific job, loaded on demand rather than carried in every request. It answers a different question than RAG: not "what does our documentation say," but "given what we know, what should Claude actually do." A refund policy's threshold rules, a support ticket's escalation criteria, a report's required format - these are Skill territory, not retrieval territory.

Reach for RAG when...

The answer could live anywhere in a large, changing corpus, and the job is finding the right passage - a knowledge base, a product catalog, a document archive.

Reach for a Skill when...

The job is applying consistent judgment or process against known rules - escalation logic, formatting standards, policy thresholds that don't change per query.

Not sure whether your AI assistant's problem is retrieval, policy, or something else entirely? Talk to us about a grounding diagnosis before building either.

Fine-tuning: the one most businesses don't need

Fine-tuning changes the model's behavior at the weight level through additional training - a fundamentally different, more expensive, and slower-to-iterate mechanism than either RAG or Skills. It earns its place for narrow, extremely high-volume tasks where prompt-based approaches have genuinely been tried and measurably fall short - a very specific classification task run millions of times a day, for instance. For the large majority of business use cases - customer support, internal knowledge, report generation, document review - RAG and Skills solve the actual problem without touching model weights, at a fraction of the engineering cost and with the ability to update the behavior in minutes instead of a retraining cycle.

How the three actually stack in a real system

LayerQuestion it answersHow often it changes
RAG / live tool connectionWhat does our current data actually say?Continuously - reflects live state
SkillGiven what we know, what should Claude do?As often as policy changes - versioned, not code-deployed
Fine-tuningHow should the model behave on this narrow, high-volume task?Rarely - a retraining cycle, not a config edit

Most production business assistants use the top two layers together and never touch the third. A support assistant, for instance, might use RAG to find the relevant help-center article and a Skill to decide how to phrase the answer and when to escalate - two mechanisms working on the same question from different angles, with fine-tuning never entering the picture.

Key takeaways

  • "Ground it in our data" is three different problems - current facts, applied judgment, and specialized behavior - each with a different fix.
  • RAG retrieves relevant passages from a large, changing corpus at query time; it's only as good as its retrieval step.
  • Skills package judgment and process, loaded on demand - they answer "what should Claude do," not "what does the data say."
  • Fine-tuning changes model weights and is rarely justified outside narrow, extremely high-volume tasks where prompting has genuinely been tried and fallen short.
  • Diagnose which failure mode you actually have before building - most "wrong answer" problems are retrieval or policy gaps, not model-capability gaps.
Which combination fits depends on your actual data shape and failure patterns - this is a general framework, not a one-size answer. Diagnosing the real gap is the right first step.