Skip to main content

66 posts tagged with "Architecture"

Cross-cloud AI architecture decisions and patterns

View All Tags

gpt-oss for Engineers: OpenAI's Apache-Licensed Reasoning Models, and Why That's Newsworthy

· 11 min read
AI Playbook author

The newsworthy fact about gpt-oss isn't a benchmark number — it's that OpenAI, a lab whose entire commercial identity has been built around closed, API-only models, shipped genuinely Apache 2.0-licensed open-weight reasoning models. That's a strategic signal worth reading carefully, and gpt-oss-120b and gpt-oss-20b deserve evaluation on their own technical merits rather than being dismissed as a side project or overhyped as a full strategy reversal.

Qwen 3.5 for Engineers: A 397B Hybrid DeltaNet+MoE Model You Can Actually License Freely

· 10 min read
AI Playbook author

Qwen 3.5 is the model to point to when someone claims "open source" and "commercially free to use without restriction" are the same category as Kimi K3 or Llama 4. They aren't — and Qwen 3.5's Apache 2.0 licence is precisely why it deserves separate treatment from its open-weight cousins. A 397-billion-parameter model with only 17B active per token, genuinely Apache-licensed, is a different legal and commercial proposition than a similarly-sized model under a custom licence, even if the benchmark numbers look comparable.

Grok 4.5 for Engineers: What the Cursor Co-Training Story Actually Means for Your Stack

· 11 min read
AI Playbook author

Grok 4.5 is the first model from xAI trained specifically for coding and agentic work, and the headline story — a joint-training partnership with Cursor using real developer interaction data — is genuinely unusual among frontier labs. Most vendors train on public repositories and synthetic agent trajectories; xAI is claiming an edge from live, in-editor developer behaviour. That's a real architectural difference worth understanding, and also exactly the kind of claim you should verify on your own codebase before routing production traffic to it.

Build a Large Language Model from Scratch: The Complete Engineering Playbook

· 19 min read
AI Playbook author

Building a GPT-style large language model yourself is the fastest way to stop treating foundation models as black boxes. This playbook walks through the full path used in educational GPT implementations: prepare text, implement attention, assemble a decoder-only stack, pretrain (or load weights), then fine-tune for classification or instruction following.

The conceptual sequence mirrors production LLM development at smaller scale. You can run the educational path on a laptop; frontier training still needs datacenter compute. Use this article as an engineering map, then implement with LLMs-from-scratch and Sebastian Raschka’s Build a Large Language Model (From Scratch) (Manning).

Databricks Enterprise GenAI Engineering: AI Search, Unity AI Gateway, MLflow 3, AI Functions, LLMOps and Genie One

· 37 min read
AI Playbook author

Enterprise generative AI engineering is no longer limited to writing prompts and connecting an application to a large language model. A production AI system must combine software engineering, governed data access, model routing, retrieval, tool execution, evaluation, monitoring, security, cost control and continuous delivery.

Databricks addresses these requirements through an integrated set of capabilities covering the complete GenAI lifecycle: querying foundation models and agents, building custom and low-code agents, connecting agents to governed tools, preparing structured and unstructured data, implementing retrieval with AI Search, deploying agents and applications, governing traffic through Unity AI Gateway, tracing with MLflow, evaluating and monitoring quality, operationalising through LLMOps, and delivering governed experiences through Genie One.

Diffusion Models Explained: Theory, Architecture, Training, Sampling, and Hugging Face Diffusers

· 40 min read
AI Playbook author

Diffusion models are generative machine-learning models that learn to create data by reversing a gradual corruption process. During training, clean data—such as an image—is progressively disturbed with Gaussian noise. A neural network then learns how to predict and remove that noise. During generation, the model begins with random noise and repeatedly denoises it until a coherent image, video, audio clip, three-dimensional object, molecular structure, or other output emerges.

Although early diffusion systems operated directly on image pixels, modern systems often work in a compressed latent space and use either a convolutional U-Net or a transformer-based denoising network. Text encoders, cross-attention, classifier-free guidance, ControlNet, LoRA adapters, sophisticated numerical solvers, quantisation, distillation, and flow-matching techniques have made diffusion models more controllable and computationally practical.

Hugging Face Diffusers provides a modular implementation of this ecosystem. It separates pretrained denoising models, schedulers, text encoders, autoencoders, adapters, and pipelines so developers can combine and optimise them independently. The library supports image, video, audio, editing, restoration, super-resolution, depth estimation, and other diffusion-related workflows.

From GPT-2 to Kimi Delta Attention: How Transformers Evolved into Modern AI Memory Systems

· 21 min read
AI Playbook author

Twenty-two thousand five hundred and eighty. That is roughly how many GPT-2-scale models (≈124M parameters) fit inside Kimi K3 (≈2.8T parameters, 2026). Scale mattered—but it is not the whole story.

A parallel evolution happened inside the architecture itself:

AI models gradually changed from systems that repeatedly search every previous token into systems that maintain, update, erase, and selectively retrieve structured internal memories.

This article combines that memory-systems interpretation with the concrete architectural worklog 22580: From GPT2 to Kimi3, Explained by ali (@waterloo_intern), including the original diagrams, plus the research path from full softmax attention → linear attention → DeltaNet → Gated DeltaNet → Kimi Delta Attention → Kimi Linear / Kimi K3.

Large Language Model Architectures Explained: Mathematics, Programs, Analogies, and Data Flow

· 47 min read
AI Playbook author

An LLM architecture defines how a model converts language into numbers, moves information between tokens, stores short-term contextual memory, transforms representations through neural layers, predicts or generates tokens, and scales parameters and computation.

Most modern language models belong to one or more of these families:

  • Recurrent architectures: RNN, LSTM, GRU and RWKV
  • Encoder-only Transformers: BERT-style models
  • Decoder-only Transformers: GPT, Llama and Mistral-style models
  • Encoder–decoder Transformers: T5 and BART-style models
  • Sparse Mixture-of-Experts models
  • Local, sparse and linear-attention models
  • Retention and delta-rule architectures
  • State-space models such as Mamba
  • Convolutional sequence models such as Hyena
  • Hybrid attention–recurrent architectures
  • Diffusion language models
  • Multimodal language models
  • Retrieval-augmented language-model systems

LLM Caching Ultimate Guide: KV Cache, PagedAttention, Prefix Caching, and Semantic Caching

· 20 min read
AI Playbook author

LLM inference is expensive because the model repeatedly recomputes attention state over the same tokens. Caching attacks that redundancy at every layer of the stack: inside one decode loop (KV cache), across concurrent requests on a serving GPU (PagedAttention / prefix caching), across API calls for stable prompt prefixes (prompt caching), and finally at the application layer when the answer itself can be reused (exact / semantic response caching).

This guide synthesizes the mechanics from Sebastian Raschka’s KV-cache and architecture work, Sankalp’s PagedAttention deep dive, provider prompt-caching docs, AWS caching patterns, and application-layer guides from Machine Learning Mastery, IBM, Latitude, ngrok, and related sources into one engineering playbook.