AI System Design Patterns for Production LLM Applications
Routing, retrieval, guardrails, and observability — the system design choices that separate demo chatbots from reliable AI products.
7 min read
Building a demo with an LLM API takes an afternoon. Running that same model in production — where latency budgets, cost ceilings, hallucination risk, and compliance requirements all apply — is a system design problem. The teams shipping reliable AI products think in pipelines, not prompts: where data enters, how requests get routed, what gets logged, and what happens when the model is wrong.
This article covers the architectural patterns that recur across production LLM applications, from customer support bots to internal copilots.
The baseline architecture
Most production LLM apps follow a similar skeleton:
Client → API Gateway → Orchestrator → [Retriever | Tools | Model] → Post-processor → Response
↓
Observability store
The orchestrator is the brain. It decides which model to call, whether to fetch context first, and how to handle failures. Treat it as application code with tests, not as a single prompt string.
Pattern 1: Retrieval-augmented generation (RAG)
When answers must reference company docs, tickets, or policies, pure model knowledge is insufficient. RAG retrieves relevant chunks at query time and injects them into the prompt.
Design decisions that matter:
- Chunk size and overlap — too small loses context; too large wastes tokens.
- Embedding model choice — must match the domain (code vs legal vs support).
- Vector store — Pinecone, pgvector, OpenSearch, or managed options depending on scale.
- Re-ranking — a cross-encoder reranker after vector search often improves precision.
Always return citations. Users trust answers they can verify, and support teams need audit trails.
Pattern 2: Router / multi-model dispatch
Not every request needs your most expensive model. A router classifies intent and sends simple FAQs to a small fast model while reserving large models for complex reasoning.
def route(user_message: str) -> str:
if is_faq(user_message):
return "small-model"
if requires_code_analysis(user_message):
return "code-specialist"
return "general-large"
Routers can be rule-based, embedding similarity, or a fine-tuned classifier. Measure cost per resolved ticket — not cost per token — when evaluating routing logic.
Pattern 3: Tool use and agents (with guardrails)
Tool-calling lets models trigger APIs: search databases, create tickets, run calculations. The risk is unbounded action — an agent that can delete records needs the same authorization checks as your backend.
Principles:
- Whitelist tools explicitly; never expose raw SQL generation without validation.
- Require human approval for irreversible actions (refunds, deployments, data deletion).
- Cap iteration loops — agents that recurse indefinitely burn budget and confuse users.
Start with single-step tool calls before building multi-turn autonomous agents.
Pattern 4: Prompt and context assembly
Production prompts are versioned artifacts, not Slack messages. Store templates with variables, test them in CI, and roll back when quality drops.
Context assembly tips:
- Put the most important instructions at the beginning and end of the prompt (models weight edges heavily).
- Separate system rules from user content with clear delimiters.
- Trim conversation history — summarize older turns instead of sending full transcripts.
Pattern 5: Guardrails and output validation
Assume the model will eventually produce something unsafe or malformed.
Layers that work:
- Input filters — block prompt injection patterns, PII you should not process, off-topic abuse.
- Structured output — request JSON schema and validate with pydantic or similar before acting on results.
- Output filters — regex or classifier checks for policy violations before showing text to users.
- Fallback responses — when validation fails, return a safe generic message and log the incident.
Open-source and hosted guardrail libraries exist, but custom validators for your domain often matter more than generic toxicity classifiers.
Pattern 6: Caching and cost control
LLM inference is expensive at scale. Effective controls:
- Semantic cache — if a new question is near-dupe of a recent one, return the cached answer.
- Prompt caching (where providers support it) — reuse prefix tokens for shared system prompts.
- Batch non-urgent jobs — summarization and indexing do not need real-time APIs.
- Token budgets per request — hard caps prevent runaway context growth.
Track cost per feature, not just per environment.
Pattern 7: Observability you will actually use
Log at minimum:
- Request ID, model, latency, token counts, retrieval sources used
- Full prompt/response pairs in a secure internal store (not in client-visible logs)
- User feedback signals (thumbs up/down, escalation to human)
Build dashboards for quality regressions — when answer acceptance rate drops after a prompt change, you want an alert before customers complain.
Tools like LangSmith, Helicone, or custom OpenTelemetry pipelines are common; the pattern matters more than the vendor.
Pattern 8: Evaluation before deployment
Maintain a golden dataset of real user questions with expected properties (must mention refund policy, must not hallucinate pricing). Run evaluations on every prompt or model change.
Metrics to track:
- Faithfulness — answer grounded in retrieved context?
- Task success — did the user complete their goal?
- Latency p95 — within SLA?
Automated LLM-as-judge scoring is useful for triage; human review remains the ground truth for high-stakes domains.
Failure modes and how to handle them
| Failure | Mitigation |
|---|---|
| Hallucinated facts | RAG + citations + validation |
| Slow responses | Smaller models, streaming, async jobs |
| Prompt injection | Input sanitization, tool permission boundaries |
| Model outage | Fallback model or graceful degradation message |
| Cost spike | Rate limits, routing, caching |
Design for degradation. A support bot that says "I cannot answer that — here is a help article" beats one that invents policy.
FAQ
Do I need an agent framework? Frameworks (LangGraph, CrewAI, etc.) help prototype. In production, many teams use a thin custom orchestrator for predictability.
RAG or fine-tuning? RAG when knowledge changes frequently. Fine-tuning when you need consistent style, format, or domain language and data is stable.
How do I choose models? Benchmark on your tasks. Public leaderboards are a starting point, not a purchase order.
Ship systems, not demos
AI system design is mostly classical engineering: clear boundaries, explicit failure handling, measurable quality, and least-privilege access to tools. The model is one component in a pipeline. Teams that treat orchestration, retrieval, validation, and observability as first-class concerns build products that survive real traffic — not just impressive screenshots.
Deployment and infrastructure choices
Sync vs async inference
Interactive chat needs synchronous low-latency endpoints with streaming enabled (Server-Sent Events or WebSockets) so users see tokens as they generate. Batch workloads — nightly summarization, embedding backfills — should run as async jobs on cheaper compute with retries and dead-letter queues.
Where to run the orchestrator
Small teams often run the orchestrator inside the same API service (FastAPI, Node). At higher scale, separate the orchestration layer from the HTTP edge so you can scale retrieval and model calls independently. Containerize with clear CPU/memory limits; embedding and reranking steps can be memory-heavy.
Data residency and compliance
If you process EU customer data, model API region selection and vector store location matter for GDPR. Some teams route PII-heavy requests to on-prem or regional deployments while sending anonymized prompts to cloud models. Document data flows before legal asks.
Building a roadmap from prototype to production
A realistic rollout:
- Week 1–2 — Single-model Q&A with hardcoded context, manual logging.
- Week 3–4 — Add RAG over one document source, structured output validation.
- Month 2 — Router for cost, golden-set evals in CI, basic dashboards.
- Month 3+ — Tool use with approval flows, semantic cache, on-call runbooks.
Skipping straight to autonomous agents without steps two and three is how teams ship impressive demos that break under the first week of real users.
More in artificial-intelligence
Cubed
Write about the technologies shaping the future.
For developers, founders, and curious minds exploring AI, crypto, Web3, and emerging tech—signal over noise.
One free account across In Plain English, Stackademic, Venture, and Cubed.
How it works- AI, crypto & Web3
- Software & emerging technologies
- Analysis & practical resources
- Thoughtful voices, not hype
Sign in
Google or GitHub
Complete profile
Takes a few minutes
Get approved & publish
Start sharing
Why write for Cubed?
The future deserves thoughtful voices, not just louder headlines.
Comments
Loading comments…