The layers around the model call

A production LLM application is better thought of as a pipeline with the model in the middle. Requests arrive through an application layer that handles authentication, rate limits, and session state. An orchestration layer assembles the prompt - instructions, policy, retrieved context, conversation history - calls the model, and decides what to do with the result, including calling tools and looping.

Around that sit the layers that make it trustworthy: retrieval to ground answers in your own data, validation and guardrails on inputs and outputs, human approval gates for consequential actions, and observability that records what happened on every request.

LayerJobFailure it prevents
RetrievalSupply current, relevant, permission-checked contextStale or invented answers
OrchestrationAssemble prompts, call tools, manage loopsRunaway loops and brittle prompt strings
GuardrailsValidate inputs and outputs, enforce policyUnsafe or malformed responses reaching users
Approval gatesPause before irreversible actionsCostly autonomous mistakes
ObservabilityTrace every request, token, tool call, and costDebugging blind and surprise bills

Evals are the contract that lets you change anything else

The most underbuilt layer is evaluation. Without a set of representative test cases with expected behavior, every change to a prompt, a retrieval setting, or a model version is a guess. A golden set built from real historical requests lets you measure quality before you ship a change and detect regressions after.

Treat the eval suite as the contract for the whole system. It is what makes it safe to upgrade a model, tighten a prompt, or swap a vector store, because you can show the change did not make things worse. It is also the answer to the question every serious buyer asks: how do you know it works?

Build the eval set before the second prompt revision, not after the first incident. Even thirty well-chosen cases beat intuition.

Observability and cost control are design requirements

Log a trace for every request: the assembled prompt, retrieved sources, tool calls and results, the response, latency, and token usage. When an answer is wrong, the trace tells you whether retrieval, the prompt, or the model was at fault. Aggregate the same data to watch quality, latency, and spend over time.

Cost control is architectural too. Techniques such as caching stable prompt prefixes, routing simple requests to smaller models, and capping tool-loop iterations are far easier to add when the orchestration layer is a distinct component rather than strings scattered across the codebase.

Trace every request

Prompt, context, tool calls, response, latency, and tokens - enough to reconstruct any answer after the fact.

Budget at the orchestration layer

Per-request and per-user limits, model routing, and caching live in one place instead of everywhere.

Moving an LLM prototype toward production? Talk to us about an architecture review before the first real traffic arrives.

Key takeaways

  • The model call is the smallest part of a production LLM system; retrieval, orchestration, guardrails, approval gates, and observability determine whether it holds up.
  • Build an evaluation suite from real historical requests early - it is the contract that makes prompt, retrieval, and model changes safe.
  • Trace every request end to end so a wrong answer can be traced to retrieval, prompt, or model.
  • Keep orchestration as a distinct layer so cost controls like caching, routing, and loop limits can live in one place.