Skip to main content

Models

Large Language Model Architectures Explained: Mathematics, Programs, Analogies, and Data Flow

A deep architectural guide to LLMs—tokenization, embeddings, RNNs, Transformers, attention, MoE, linear attention, DeltaNet, RetNet, Mamba, RWKV, Hyena, diffusion LMs, multimodal models, RAG, LoRA, KV cache, and production trade-offs—with KaTeX equations and runnable PyTorch sketches.

AI Playbook2026-07-2831 min readadvanced

Executive takeaway​

A deep architectural guide to LLMs—tokenization, embeddings, RNNs, Transformers, attention, MoE, linear attention, DeltaNet, RetNet, Mamba, RWKV, Hyena, diffusion LMs, multimodal models, RAG, LoRA, KV cache, and production trade-offs—with KaTeX equations and runnable PyTorch sketches.

In this briefing​

  1. Key ideas
  2. Practical implications
  3. What to do next

Why it matters​

Use this briefing to decide whether to deep-dive the full article for your current delivery problem.

AI EngineerConsultantStudent

Turn insight into action​

Read the full analysis, then continue into a playbook or framework.