Skip to main content
AI Playbook author
View all authors

Large Language Model Architectures Explained: Mathematics, Programs, Analogies, and Data Flow

· 47 min read
AI Playbook author

An LLM architecture defines how a model converts language into numbers, moves information between tokens, stores short-term contextual memory, transforms representations through neural layers, predicts or generates tokens, and scales parameters and computation.

Most modern language models belong to one or more of these families:

  • Recurrent architectures: RNN, LSTM, GRU and RWKV
  • Encoder-only Transformers: BERT-style models
  • Decoder-only Transformers: GPT, Llama and Mistral-style models
  • Encoder–decoder Transformers: T5 and BART-style models
  • Sparse Mixture-of-Experts models
  • Local, sparse and linear-attention models
  • Retention and delta-rule architectures
  • State-space models such as Mamba
  • Convolutional sequence models such as Hyena
  • Hybrid attention–recurrent architectures
  • Diffusion language models
  • Multimodal language models
  • Retrieval-augmented language-model systems

LLM Caching Ultimate Guide: KV Cache, PagedAttention, Prefix Caching, and Semantic Caching

· 20 min read
AI Playbook author

LLM inference is expensive because the model repeatedly recomputes attention state over the same tokens. Caching attacks that redundancy at every layer of the stack: inside one decode loop (KV cache), across concurrent requests on a serving GPU (PagedAttention / prefix caching), across API calls for stable prompt prefixes (prompt caching), and finally at the application layer when the answer itself can be reused (exact / semantic response caching).

This guide synthesizes the mechanics from Sebastian Raschka’s KV-cache and architecture work, Sankalp’s PagedAttention deep dive, provider prompt-caching docs, AWS caching patterns, and application-layer guides from Machine Learning Mastery, IBM, Latitude, ngrok, and related sources into one engineering playbook.

LLM Cost Tracking and FinOps Across Azure, AWS, and Google Cloud

· 33 min read
AI Playbook author

Generative AI cost management is more complicated than ordinary cloud cost management. A conventional application might be charged according to CPU hours, memory, storage, and network usage. An LLM application can accumulate costs through input tokens, output tokens, cached tokens, reasoning tokens, embeddings, retrieval, reranking, guardrails, model evaluation, agent tool calls, retries, observability, and the infrastructure supporting the application.

The main FinOps challenge is therefore not simply:

How can we use the cheapest model?

The correct question is:

What is the lowest sustainable cost at which the application can meet its quality, latency, reliability, security, and business-outcome requirements?

LLM Response Caching Strategies for RAG and Agentic Applications

· 30 min read
AI Playbook author

Caching in a conventional application usually means storing the result of a database query or API request. In an LLM application, there are many more opportunities: repeated questions, equivalent paraphrases, identical embeddings, repeated retrieval and reranking, repeated tool calls, stable system prompts, and recurring agent planning steps.

The strongest architecture therefore does not implement one “LLM cache.” It implements a hierarchy of specialised caches across the RAG and agent pipeline.

Security, Compliance and Governance for Open-Source and Closed-Source LLM Deployments

· 39 min read
AI Playbook author

Deploying a large language model is not simply a question of choosing between an open-source model and a commercial API. It is an enterprise risk decision involving:

  • What information the system will process.
  • Where that information will travel.
  • Who can access the model, prompts, outputs and logs.
  • What actions the model can perform.
  • How the organisation will detect failures or attacks.
  • Which party is accountable when something goes wrong.
  • What evidence can be presented to auditors, regulators, customers and executives.

Case A Parent: Meridian Insurance GenAI Productivity

· 2 min read
AI Playbook author

Case A parent. Meridian Insurance (composite) faces rising cost-to-serve. The board wants a GenAI plan this quarter. The COO sponsors; the CRO fears hallucination, privacy and audit gaps. Knowledge is fragmented; there is no enterprise evaluation framework. This overview states the decision and outcome. Expanded articles cover discovery, solution/commercial design and delivery.

Case B Parent: MonGo Bank Customer-Service AI

· 3 min read
AI Playbook author

Case B parent. MonGo Bank is the playbook’s worked retail-and-SME bank: millions of customers, a large contact centre, mixed cloud and legacy platforms, strict conduct and privacy obligations. The executive ask arrives as “build a generative AI chatbot that reduces cost.” This parent states the reframe and points to the expanded Banking CS series.