Skip to main content

164 posts tagged with "Playbook"

Posts about the AI Playbook product and practice

View All Tags

DeepSeek for Engineers: MLA, MoE, and GRPO — the Three Ideas Behind the Cheapest Frontier-Class Training Run

· 11 min read
AI Playbook author

DeepSeek's real contribution to the 2026 model landscape isn't any single model — it's three specific technical ideas that the rest of the industry has since absorbed to varying degrees: Multi-head Latent Attention for cheaper KV-cache memory, Mixture-of-Experts done at extreme sparsity, and Group Relative Policy Optimisation as a reinforcement-learning method that made reasoning-focused training dramatically cheaper. If you understand those three ideas, you understand why DeepSeek mattered, and you'll recognise the same ideas (in variant form) inside half the other models in this series.

Designing an EMEA Go-to-Market Strategy and Roadmap for an AI Solution

· 37 min read
AI Playbook author

A go-to-market strategy for an AI solution is not simply a marketing plan. It is the coordinated design of:

GTM=Target market×Urgent problem×Differentiated solution×Commercial model×Trust×Distribution×Adoption\text{GTM} = \text{Target market} \times \text{Urgent problem} \times \text{Differentiated solution} \times \text{Commercial model} \times \text{Trust} \times \text{Distribution} \times \text{Adoption}

For AI products, the “trust” component is particularly important. A technically impressive solution can still fail because the buyer cannot establish:

  • Who is accountable for its outputs.
  • Where customer data is processed.
  • Whether the model can hallucinate.
  • Whether regulators will accept it.
  • Whether employees and customers will use it.
  • Whether its financial benefits exceed implementation and operating costs.

In EMEA, the challenge is greater because EMEA is not one market. An AI solution sold in the UK, Germany, the UAE, Saudi Arabia and South Africa may require different hosting, contracting, languages, regulatory controls, partner models and sales motions.

This guide explains the complete process and then applies it to a detailed hypothetical case study: an AI customer-service platform for regulated banks.

Financial Modelling for AI: From Beginner Fundamentals to Advanced AI Economics

· 42 min read
AI Playbook author

Financial modelling is the process of translating a business idea, investment, product, project, or company into numbers.

A financial model helps decision-makers answer questions such as:

  • How much will the AI solution cost?
  • How will the solution generate financial value?
  • When will the investment break even?
  • How much cash will be required?
  • What happens if adoption is slower than expected?
  • Is it cheaper to build, buy, or partner?
  • Which AI architecture provides the best combination of cost, quality, latency, and risk?
  • What is the company or AI product worth?
  • Should the organisation approve, delay, redesign, or reject the investment?

For an ordinary software project, financial modelling usually connects customers, prices, employees, and infrastructure costs.

For an AI solution, the model must connect several additional variables:

Business Demand→AI Usage→Model Performance→Technical Cost→Business Outcome→Financial Value\text{Business Demand} \rightarrow \text{AI Usage} \rightarrow \text{Model Performance} \rightarrow \text{Technical Cost} \rightarrow \text{Business Outcome} \rightarrow \text{Financial Value}

For example:

Customer enquiries→AI conversations→Model calls and tokens→Resolved cases→Reduced contact-centre cost\text{Customer enquiries} \rightarrow \text{AI conversations} \rightarrow \text{Model calls and tokens} \rightarrow \text{Resolved cases} \rightarrow \text{Reduced contact-centre cost}

The most important principle is:

Do not measure only cost per token, request, model call, or GPU hour. Measure cost and value per successful business outcome.

Examples of meaningful AI financial units include:

  • Cost per successfully resolved customer enquiry
  • Cost per approved insurance claim
  • Cost per qualified sales lead
  • Cost per completed legal review
  • Cost per detected fraud case
  • Cost per accurate document extraction
  • Cost per software feature delivered
  • Cost per clinical document summarised
  • Revenue per AI-assisted customer
  • Gross profit per AI agent session

Cloud FinOps applies the same fundamental relationship to AI as to other cloud services:

Cost=Price×Quantity\text{Cost} = \text{Price} \times \text{Quantity}

However, AI introduces unusual quantities such as input tokens, output tokens, model calls, GPU time, retrieval operations, evaluation runs, agent steps, tool calls, and human-review events. FinOps therefore recommends connecting cloud costs to business-unit economics rather than viewing infrastructure spending in isolation.

From Chatbots to Autonomous Digital Workers: How Frontier AI Models Are Becoming Smarter Over Time

· 33 min read
AI Playbook author

Artificial intelligence models are improving at a remarkable pace. A model released only six months ago can quickly appear less capable, less efficient and less reliable than a newer generation. It is tempting to assume that every new model is smarter because it contains more parameters — but modern frontier systems such as Claude Fable 5, Claude Opus 5, GPT-5.6, Gemini 3.6 Flash and Kimi K3 are improving through a much broader stack of advances.

GLM-5 for Engineers: Built for Agentic Engineering, Not Just Chat

· 10 min read
AI Playbook author

Most models in this series added agentic and coding capability on top of a general-purpose foundation. GLM-5 is positioned the other way around: an architecture and training programme explicitly aimed at agentic software engineering — multi-step tool use, long-running autonomous tasks, and code-heavy reasoning — as the primary target, with general chat capability as a secondary consequence rather than the main event. That inversion of priorities is the thing actually worth evaluating here, not a benchmark score.

Gemini 3.x for Engineers: Flash, Pro and Live — Picking the Right Model in a Fast-Moving Lineup

· 11 min read
AI Playbook author

Google's Gemini strategy in 2026 is breadth over singular flagship supremacy: text, images, audio, video, real-time voice, search grounding and Workspace integration under one API surface, refreshed on a fast Flash cadence while the headline Pro upgrade quietly slips month after month. That's not a criticism — it's a fact about how to plan around this family. If your architecture assumes "we'll upgrade to Gemini 3.5 Pro when it ships," you need a fallback plan, because as of this writing it still hasn't.

Gemma 4 for Engineers: The Apache-Licensed Family Built to Run on Your Laptop, Not Just Your Datacenter

· 11 min read
AI Playbook author

Gemma 4 is the model family to reach for when the question isn't "which frontier model is smartest" but "which model can I actually run on the hardware I already have." Spanning from a 2B-effective-parameter model that runs on a laptop to a 31B model that needs a real GPU, Apache 2.0-licensed and multimodal out of the box, Gemma 4 is Google's answer to on-device and edge deployment — a genuinely different design target from Gemini's cloud-platform breadth.

Kimi K3 for Engineers: A 2.8T-Parameter Open-Weight Model You Can Actually Self-Host

· 11 min read
AI Playbook author

Kimi K3 is the model that forces a lot of "just use the closed frontier API" arguments to actually justify themselves with numbers. It's a 2.8-trillion-parameter Mixture-of-Experts model with only 104B parameters active per token, a genuine 1M-token context window, and weights you can download and run on your own hardware today. It is also not, strictly, open source — and that distinction matters more than most teams evaluating it realise.

Knowledge Distillation for Large Language Models

· 32 min read
AI Playbook author

Knowledge distillation is a model-compression and capability-transfer technique in which a powerful teacher model provides training signals for a smaller student model.

Instead of requiring the student to rediscover every useful behaviour from raw internet-scale pretraining data, the student learns from the teacher’s outputs, probability distributions, internal representations, reasoning demonstrations or preferences.

A distilled model can therefore become:

  • Faster at inference
  • Less expensive to operate
  • Smaller in memory
  • Easier to deploy on limited hardware
  • More specialised for a particular task
  • More consistent than the original general-purpose model for a narrow workflow

However, an important distinction must be made:

The exact internal process used to train current proprietary Claude models is not publicly disclosed in full.

Anthropic publishes model reports, safety research and selected training information, but its public materials do not expose Claude’s weights, token logits, hidden states, training datasets or the complete teacher–student training recipe. Therefore, when people discuss “distilling Claude,” they may be referring to two different things:

  1. Authorised distillation within an official platform, such as Amazon Bedrock’s documented Claude Sonnet-to-Haiku distillation workflow.
  2. Black-box behavioural distillation, where permitted Claude outputs are collected and used to fine-tune another model.

Anthropic and AWS have publicly described an authorised workflow in which Claude 3.5 Sonnet generates synthetic training data, Claude 3 Haiku is trained and evaluated using that data, and the resulting distilled model is hosted for inference. This is primarily output-based behavioural distillation rather than traditional access to Claude’s internal logits or hidden layers.

The 2026 LLM Reading Map: A Practical Order for Working Through the Whole Model Landscape

· 9 min read
AI Playbook author

Fifteen model families, three licensing categories, and at least six distinct architectural ideas (MoE, MLA, KDA, DeltaNet, Mamba/LatentMoE, sparse attention) is a lot to hold in your head at once. This post is the map, not the territory — a condensed way to navigate the full 2026 model landscape without reading fifteen separate briefings cold, plus the one terminology distinction (closed vs. open-weight vs. open-source) that trips up more procurement and architecture conversations than any benchmark score ever does.