Skip to main content

Models

Large Language Model Architectures Explained: Mathematics, Programs, Analogies, and Data Flow

A deep architectural guide to LLMs—tokenization, embeddings, RNNs, Transformers, attention, MoE, linear attention, DeltaNet, RetNet, Mamba, RWKV, Hyena, diffusion LMs, multimodal models, RAG, LoRA, KV cache, and production trade-offs—with KaTeX equations and runnable PyTorch sketches.

AI Playbook2026-07-2831 min readadvanced

Executive takeaway

A deep architectural guide to LLMs—tokenization, embeddings, RNNs, Transformers, attention, MoE, linear attention, DeltaNet, RetNet, Mamba, RWKV, Hyena, diffusion LMs, multimodal models, RAG, LoRA, KV cache, and production trade-offs—with KaTeX equations and runnable PyTorch sketches.

In this briefing

  1. Key ideas
  2. Practical implications
  3. What to do next

Why it matters

Use this briefing to decide whether to deep-dive the full article for your current delivery problem.

AI EngineerConsultantStudent

Turn insight into action

Read the full analysis, then continue into a playbook or framework.