Production LLM Optimisation: Proven Methods for Cost, Latency, Quality and Scale
LLM optimisation in 2026 is no longer simply about reducing token counts.
Production systems combine caching, routing, retrieval, context engineering, asynchronous execution, efficient agent architectures, inference optimisation and workload-aware infrastructure. The largest savings usually come from avoiding unnecessary inference, not from making every remaining call a few percent cheaper.
Executive view
Technical view
The fundamental engineering objective is:
Operationally:
The important word is successful. Cutting cost per request while decreasing task completion is not optimisation. A production team should therefore target:
rather than cost per token.
Part I — The optimisation stack and the baseline
1. The modern LLM optimisation stack
A mature request path is several layers, each of which can refuse to spend:
Optimisation therefore occurs at multiple layers:
- Prompt and context optimisation
- Prompt / KV caching
- Semantic response caching
- Model routing
- Model cascading
- RAG and retrieval optimisation
- Context-window optimisation
- Tool optimisation
- Agent-loop optimisation
- Parallel and asynchronous execution
- Batch inference
- Hosted versus open-weight model selection
- GPU / inference optimisation
- Observability and continuous evaluation
The biggest savings frequently come from avoiding inference altogether.
2. Establish the baseline before optimising
Never start by randomly shortening prompts or replacing models. Instrument the system first.
Quality
Track at least:
- Task completion rate
- Answer correctness
- Groundedness / faithfulness
- Hallucination rate
- Retrieval recall
- Tool success rate
- User satisfaction
- Escalation rate
Optimisations that improve cost while increasing escalations usually fail the unit-economics test. Pair quality gates with AI Evaluation and Quality Assurance.
Latency
Track TTFT, p50, p95, p99, tokens per second, retrieval latency, reranking latency, tool latency and total agent duration. Users experience the full path, not the model vendor's marketing TTFT.
Cost
For agents:
And most importantly:
Part II — Caching
3. Prompt caching
Prompt caching is one of the highest-value optimisations when requests contain large repeated context.
Imagine every request contains:
8,000 tokens system instructions
5,000 tokens policies
4,000 tokens tool definitions
10,000 tokens documentation
200 tokens user question
Without caching, the model repeatedly processes approximately 27,200 tokens. Most of those tokens may be identical. Structure the request so the reusable prefix is stable:
STATIC
System instructions
Policies
Tool definitions
Reference material
---------------- CACHE BOUNDARY ----------------
DYNAMIC
Conversation changes
Retrieved context
Current user question
Amazon Bedrock prompt caching is documented as a mechanism for reducing inference latency and input-token cost by skipping recomputation of reused prefixes. AWS reports reductions of up to 85% latency and 90% input cost for suitable workloads. Treat those figures as best-case results, not universal expectations.
Cache writes are often billed at a premium over uncached input, while cache reads are discounted. A low hit rate can therefore increase spend. Monitor read-to-write ratio; a healthy ratio is typically well above 1:1.
Prefix order matters
Worse — the beginning changes every request:
USER QUESTION
SYSTEM PROMPT
TOOLS
DOCUMENTATION
POLICIES
Better — static content first:
SYSTEM PROMPT
TOOLS
POLICIES
DOCUMENTATION
----------------
USER QUESTION
Many prompt-cache implementations rely on reusable prefixes. Changing something near the beginning can invalidate reuse of everything that follows.
Best candidates
- Coding assistants
- Long system prompts
- Multi-turn agents
- Document Q&A
- Large tool schemas
- Policy-heavy enterprise assistants
- Customer-service assistants
- Repeated analysis of the same document
AWS introduced a one-hour cache TTL for selected Claude models in January 2026, specifically to support longer agent workflows and conversations. Longer TTLs usually cost more to write; use them when idle gaps would otherwise expire a 5-minute cache.
Mechanism depth: LLM Caching Ultimate Guide.
4. Semantic caching
Prompt caching and semantic caching solve different problems.
| Mechanism | What it reuses | When inference still runs |
|---|---|---|
| Prompt / prefix cache | Model computation for a repeated prefix | Yes — a new completion is still generated |
| Semantic cache | A previous answer for an equivalent question | No — on a confident hit |
String comparison sees three different questions:
How do I reset my password?
I forgot my password. How can I change it?
Where can I reset my login password?
Embedding similarity may recognise the same intent.
SIMILARITY_THRESHOLD = 0.95
def answer_with_semantic_cache(query: str) -> str:
query_embedding = embed(query)
match = cache.search(query_embedding)
if match.similarity > SIMILARITY_THRESHOLD:
return match.response
response = llm(query)
cache.store(embedding=query_embedding, response=response)
return response
0.95 is not a universal threshold. Calibrate on your own evaluation set:
0.80 → dangerous false matches
0.88 → some incorrect reuse
0.93 → acceptable for some FAQ traffic
0.96 → very safe but fewer hits
Select the threshold from utility, not hit rate:
For high-risk applications, false cache hits should carry a very high penalty. Exact-match caches remain the safe default; semantic caches need ACL filters, version keys and a two-threshold policy. See LLM response caching for RAG and agents.
5. Cache invalidation is more important than cache creation
Consider:
Customer: "What is your cancellation policy?"
Cached answer: "Cancellation is allowed within 30 days."
The company changes the policy to 14 days. The cached response is now wrong.
Cache keys should incorporate relevant versions:
hash(
normalised_query
+ tenant_id
+ knowledge_version
+ policy_version
+ model_version
+ prompt_version
)
Use TTLs as a second safety mechanism:
| Content class | Typical cache posture |
|---|---|
| Stable FAQ | Hours to 24 hours |
| Product availability | Minutes |
| Account balance | Do not semantic-cache |
| Legal / policy | Version-controlled; invalidate on publish |
| News | Extremely short TTL |
Part III — Routing and cascading
6. Model routing
Do not send every problem to your strongest model.
| Example task | Typical route |
|---|---|
| Classify this ticket | Small model |
| Summarise this email | Small or medium model |
| Analyse this contract | Strong model |
| Design a migration architecture | Strong reasoning model |
Amazon Bedrock Intelligent Prompt Routing predicts which model in a family can satisfy a request while balancing quality and cost. AWS reports potential cost reductions of up to approximately 30% for appropriate mixed-difficulty workloads. Again, treat that as a workload-dependent upper bound.
7. Build an evaluation-driven router
Routing should not be:
LONG_PROMPT_TOKEN_THRESHOLD = 1000
def select_model_by_length(prompt: str) -> str:
if len(prompt) > LONG_PROMPT_TOKEN_THRESHOLD:
return "strong_model"
return "small_model"
Prompt length is not equivalent to reasoning complexity.
Instead create an offline dataset of representative production requests. Run each candidate model. Measure quality, latency, cost and task success. For every request, estimate quality given the model, then choose:
That is a much more useful definition: choose the cheapest model predicted to satisfy the quality requirement.
Negative cases the router must handle:
- Ambiguous intent — default to a safer / stronger model, not the cheapest
- High-risk domains (legal, medical, financial advice) — pin to evaluated models
- Tool-calling or structured-output tasks — route only among models that actually succeed at schema adherence
- Adversarial or injection-heavy inputs — do not let a classifier skip safety layers
8. Cascading
Routing chooses a model before execution. Cascading tries inexpensive processing first and escalates when necessary.
VERIFIER_PASS_THRESHOLD = 0.85
def answer_with_cascade(question: str) -> str:
cheap_answer = cheap_model(question)
score = verifier(question, cheap_answer)
if score >= VERIFIER_PASS_THRESHOLD:
return cheap_answer
return strong_model(question)
This works well when a large percentage of traffic is easy. Suppose 80% of requests cost £0.002 and 20% cost £0.030:
compared with £0.030 if everything used the expensive model.
Include verifier cost and failed first attempts in the expected-cost calculation. If the verifier is itself an LLM call, cascading can lose money on uniformly hard traffic.
Part IV — Retrieval and context
9. Retrieval versus large context windows
Large context windows created an incorrect assumption: RAG is unnecessary because we can put everything into the prompt.
The correct decision is workload-dependent. A knowledge base of 2,000 pages and a question about one paragraph means:
A 2025 clinical study on reasoning over electronic health records found that retrieval could approach full-context performance while requiring drastically fewer input tokens (Hegselmann et al., arXiv:2508.14817). RAG remains valuable even with increasingly large context windows — especially when documents are long, noisy or redundant.
Use long context when the task genuinely needs broad coverage (compare all liability clauses across appendices). Use retrieval when the answer lives in a small subset. Many production systems route between the two.
10. Modern retrieval: do not assume vector search is enough
A stronger production pipeline is often hybrid retrieval plus reranking:
A 2026 benchmark of more than 23,000 financial-document queries over mixed text-and-table reports found that hybrid retrieval followed by neural reranking substantially outperformed the tested single-stage approaches. BM25 also beat the tested state-of-the-art dense retrieval on that dataset (Strich et al., arXiv:2604.01733). "Vector search is always better" is not a safe assumption.
Methodology:
BM25 vs Dense vs Hybrid vs Hybrid + Reranker
Measure Recall@K, MRR, nDCG, answer correctness, latency and retrieval cost — then choose empirically. See RAG architecture.
11. Adaptive retrieval
Not every request needs RAG.
| Request | Retrieval |
|---|---|
| "Hello" | None |
| "What can you do?" | Probably none |
| "What is our refund policy?" | Retrieve |
| "Compare clauses 7 and 14 of my contract." | Deep retrieve |
This prevents unnecessary embedding calls, vector searches, reranking, context tokens and latency. Adaptive retrieval that is wrong in the skip direction (no retrieve when evidence is required) is usually worse than over-retrieving — evaluate both error types.
12. Context compression
Retrieval optimisation does not end after retrieval. Eight chunks of 1,000 tokens each is 8,000 tokens, of which perhaps 1,500 are useful.
8,000 retrieved tokens
↓
context filtering / compression
↓
2,000 relevant tokens
↓
LLM
Do not compress so aggressively that useful evidence disappears. Compression is another model (or extractive filter) on the critical path — include its latency, cost and failure modes in the scorecard.
Part V — Agents and tools
13. Agent optimisation
Agents introduce a different cost function. A naive loop may think, search, think, search, think, call an API, think, search, think, then answer — five reasoning calls where two would have been sufficient.
Reducing the number of model calls often has more impact than shaving tokens inside a single call. See Agentic AI and Building production-grade AI agents.
14. Stop using an LLM for deterministic operations
A common anti-pattern: asking the model to calculate 18492 × 17, sort records, parse deterministic JSON, calculate VAT, or check exact IDs.
Use Python, SQL, APIs, a rules engine, a calculator or a schema validator. This reduces cost and usually improves correctness.
15. Reduce tool definitions
Agents sometimes expose 50–100 tools to every request. Tool schemas consume context, and tool selection becomes harder.
Rather than LLM + 80 tool definitions, retrieve a relevant subset:
This is effectively retrieval over tools. Negative case: a too-aggressive router that hides the one tool the task actually needs. Evaluate tool-recall, not only schema-token savings.
16. Parallel tool execution
If CRM lookup takes 800 ms, order lookup 600 ms and payment lookup 700 ms, sequential execution is 2,100 ms. If independent, parallel execution approaches 800 ms (the slowest of the three) plus orchestration overhead.
async def load_customer_context(customer_id: str) -> CustomerContext:
crm, order, payment = await asyncio.gather(
get_crm(customer_id),
get_order(customer_id),
get_payment(customer_id),
)
return CustomerContext(crm=crm, order=order, payment=payment)
Dependency analysis is a latency optimisation. Do not parallelise calls that mutate shared state or that must fail closed in sequence (authenticate, then fetch).
17. Prevent runaway agent loops
Every production agent should have budgets.
MAX_STEPS = 8
MAX_TOOL_CALLS = 12
MAX_RETRIES = 2
MAX_COST_GBP = 0.20
MAX_RUNTIME_SECONDS = 30
Example envelope:
Agent budget
Tokens 30,000
Cost £0.15
Duration 20 sec
Tool calls 10
Steps 8
Retries 2
When a budget is exceeded, terminate or escalate — do not retry unbounded. This turns uncontrolled agents into bounded production workflows.
Part VI — Execution model
18. Asynchronous processing
Not every LLM operation belongs in the synchronous user path. A 300-page upload should not block on extract, chunk, embed, summarise, classify and metadata generation.
Worse: the user waits for the entire pipeline.
Better:
Upload → validate → store → return "Processing"
→ queue → worker (extract, chunk, embed, summarise, metadata)
→ ready
Good asynchronous candidates: embeddings, document ingestion, evaluation, offline summarisation, report generation, conversation analytics, topic classification, nightly processing, large-scale extraction, indexing.
19. Batch processing
If 100,000 independent documents need classification, sending them individually through interactive infrastructure is usually inefficient.
100,000 documents → queue → batching → batch inference → results store
Batching improves accelerator utilisation. The trade-off is lower cost and higher throughput versus less immediate response.
Use synchronous inference when humans are waiting. Use asynchronous or batch processing when they are not. Separate capacity pools so overnight jobs cannot starve interactive SLAs.
Part VII — Hosting and inference
20. Hosted APIs versus open-weight models
Treat this as an economics problem, not an ideological one. Compare families in the AI model landscape 2026.
Hosted model
Benefits: almost no GPU operations, rapid model upgrades, elastic capacity, managed availability, easier experimentation, low initial infrastructure investment.
Self-hosted model
Idle cost is frequently underestimated:
A GPU that is provisioned but mostly idle still costs money.
21. GPU utilisation changes the economics
Imagine GPU infrastructure at £10/hour.
At 90% useful utilisation, effective cost is approximately £11.11 per productive hour. At 20% utilisation it is £50. The infrastructure cost per unit of productive compute becomes dramatically worse.
Self-hosting does not automatically mean cheaper inference. It becomes attractive when several of these align:
- Sustained high volume
- Predictable demand
- High GPU utilisation
- Suitable open-model quality
- Internal ML platform expertise
- Strong privacy / residency requirements
- Ability to batch requests
- Sufficient scale to amortise operations
22. Continuous batching
Traditional serving can behave like: finish request A, then B, then C. Modern inference engines dynamically combine active requests:
t1: A B C
t2: A B C D
t3: A C D E
t4: C D E
The scheduler continuously adds work as other sequences finish. Benefits: better GPU utilisation, higher throughput, lower serving cost, improved concurrency.
This is one reason serving engines such as vLLM matter when evaluating self-hosted architectures. See Inference engines and Inference optimisation.
23. KV / prefix caching in self-hosted infrastructure
Transformer inference generates key-value states for previously processed tokens. If thousands of requests share a system prompt, tool definitions and company policy, recomputing the same prefix wastes GPU resources.
Systems such as LMCache show how KV caches can be shared or offloaded across inference engines and storage layers; published evaluations report substantial throughput improvements on tested multi-turn and document workloads when combined with vLLM. Benchmark numbers should not be assumed to transfer to every production workload — especially when prefix reuse is low, where offload overhead can slightly hurt throughput.
24. Quantisation
Serving can use lower-precision representations:
FP16 / BF16 → FP8 → INT8 → INT4
Lower precision can reduce memory, bandwidth and hardware requirements, and can increase throughput. The trade-off:
Never deploy quantisation from infrastructure benchmarks alone. Run application-level evaluations of the original model versus FP8, INT8 and INT4 against your production test set. See Quantisation concepts.
25. Speculative decoding
Autoregressive generation normally produces tokens sequentially. Speculative decoding uses a smaller, faster draft mechanism to propose tokens, with the stronger model verifying them.
When predictions are accepted frequently enough, generation throughput can improve without changing the final model distribution under appropriate implementations. It is a serving optimisation, not a prompt-engineering trick. Draft-target mismatch (a weak draft on a domain the strong model knows well) yields low acceptance and little gain.
Part VIII — Output, retries and the critical path
26. Output tokens matter
Teams often obsess over input tokens while allowing verbose responses. If the application only needs:
{
"risk": "medium",
"reason": "Missing approval"
}
require bounded output: JSON only, named fields, reason ≤ 40 words. This improves latency, cost, parsing, predictability and downstream reliability.
Negative case: over-constraining output so the model omits evidence needed for audit or user trust. Bound verbosity; do not bound necessary justification in high-risk flows.
27. Avoid LLM retries where possible
Retries are a hidden production cost. If a request costs £0.01 but 15% of requests retry once and 3% retry twice, real cost is higher than dashboard-level nominal inference cost.
Separate provider retry, timeout retry, schema retry, tool retry, agent retry and quality retry. Eliminate the root cause rather than increasing retry counts. Unbounded client retries multiply both cost and load during incidents.
28. Optimise TTFT separately from total latency
A response that begins in 500 ms and finishes in six seconds can feel faster than a blank four-second wait that finishes at five seconds. Track TTFT separately from end-to-end completion time.
Streaming does not necessarily reduce total computation. It improves perceived latency. Do not stream unvalidated high-risk content; run safety checks on the first tokens or delay the stream until the policy gate passes.
29. Design around the critical path
If authentication is 50 ms, retrieval 400 ms, reranking 300 ms, LLM TTFT 1,200 ms, a tool 2,000 ms and generation 2,500 ms, cutting authentication from 50 ms to 30 ms achieves almost nothing.
The same principle applies to cost. Rank work by expected reduction × traffic ÷ effort, with a quality-regression gate on every change.
Part IX — Operating system for optimisation
30. A production optimisation architecture
Behind this path, run continuously: observability, evaluation, cost attribution, prompt versions, model versions, retrieval versions, tracing, cache metrics and GPU metrics.
31. The optimisation flywheel
Optimisation is not a one-time project.
Production traffic → telemetry → failure dataset → evaluation
→ experiment → A/B test → deploy → production traffic
Every optimisation should answer:
- Did quality change?
- Did p95 improve?
- Did p99 improve?
- Did cost per task improve?
- Did task success improve?
- Did failure rate change?
32. Maintain an optimisation scorecard
| Version | Quality | p50 | p95 | Cost/task | Success |
|---|---|---|---|---|---|
| Baseline | 91% | 3.2s | 7.8s | £0.041 | 88% |
| + Cache | 91% | 2.1s | 5.0s | £0.029 | 88% |
| + Routing | 90.8% | 1.9s | 4.7s | £0.020 | 88% |
| + Better RAG | 93% | 2.2s | 5.1s | £0.022 | 91% |
| + Agent optimisation | 93% | 1.8s | 4.0s | £0.017 | 92% |
This makes optimisation a measurable engineering discipline rather than intuition. Notice that better RAG can slightly raise cost per request while still winning on cost per successful task.
33. Recommended optimisation order
Do not begin with GPU micro-optimisation.
Stage 1 — Measure. Quality, tokens, latency, cost, tool calls, model calls, retries.
Stage 2 — Eliminate unnecessary work. Do we need an LLM, retrieval, a tool, an agent, this context?
Stage 3 — Cache. Prompt, semantic, retrieval, embedding, KV where applicable.
Stage 4 — Route. Small → medium → strong based on evaluation-derived quality requirements.
Stage 5 — Optimise retrieval. BM25, dense, hybrid, reranking, adaptive retrieval, context compression.
Stage 6 — Optimise agents. Steps, tool calls, retries, LLM calls, sequential dependencies.
Stage 7 — Move work off the critical path. Queues, workers, batch processing, async pipelines.
Stage 8 — Optimise inference infrastructure. Continuous batching, quantisation, KV caching, speculative decoding, GPU scheduling, autoscaling.
Part X — Worked examples
34. Example: customer-service agent
Initial architecture: large LLM → retrieve 10 chunks → large LLM → CRM → large LLM → order API → large LLM → answer. Four model calls, two tool calls, ten chunks, large model everywhere.
Redesign:
Customer
→ semantic cache (common FAQ → immediate response)
→ cache miss → small router
├── general question → small model
├── knowledge question → hybrid RAG → reranker → model
└── account operation → relevant tools only
→ CRM + orders in parallel → model
Add prompt caching, bounded output, maximum agent steps, tool timeouts, a cost budget and async analytics. Optimisation is now structural, not merely cheaper tokens.
35. Example: enterprise document assistant
Users analyse 200-page contracts. Do not automatically choose between RAG and a 200-page context. Evaluate:
A Full document
B Dense retrieval top 10
C BM25 top 10
D Hybrid top 10
E Hybrid + reranker top 5
F Hybrid + reranker + selective long context
Questions such as "What is the termination period?" may strongly favour retrieval. "Compare liability clauses across appendices" may require much broader coverage. The optimal architecture may itself route between RAG and long context by query type.
36. Example: production coding agent
Coding agents are particularly suitable for caching because substantial context repeats: system instructions, repository instructions, tool schemas, architecture guidelines, coding conventions.
STATIC
System
Repository instructions
Tool definitions
Coding standards
========== CACHE ==========
SEMI-DYNAMIC
Relevant files
Current task
========== DYNAMIC ==========
Latest tool result
Latest user instruction
Combine prompt caching, tool filtering, repository retrieval, parallel file search, step limits, a small model for classification and a strong model for implementation — rather than repeatedly feeding the entire repository into a strong model.
Part XI — Decision framework
37. Seven questions for every LLM operation
- Can I avoid inference? Cache, rules, code, existing answer.
- Can I reduce the context? Retrieval, reranking, compression, tool filtering.
- Can I reuse computation? Prompt caching, KV caching, embedding caching, semantic caching.
- Can a cheaper model solve it? Routing, cascading.
- Can operations execute concurrently? Parallel retrieval, parallel tools, async I/O.
- Does the user need the result immediately? If not: queue, batch, background worker.
- Is infrastructure itself now the bottleneck? Continuous batching, quantisation, GPU scheduling, KV management, speculative decoding, autoscaling.
38. The most important principle
The strongest production optimisation is rarely "use fewer tokens." It is:
Perform the minimum amount of expensive intelligent computation necessary to achieve the required outcome.
In practical terms:
Don't call the model
↓
If necessary, reuse previous computation
↓
Give it only the information it needs
↓
Use the cheapest capable model
↓
Call tools only when necessary
↓
Parallelise independent work
↓
Move non-interactive work asynchronous
↓
Optimise GPU / model serving
That is the difference between LLM optimisation as prompt engineering and LLM optimisation as production systems engineering. The latter is where the largest improvements in cost, latency, reliability and scalability are usually found.
Discussion
Comments
Share feedback or questions about this page. No account required.
Loading comments…