Skip to main content

5 posts tagged with "RAG"

Retrieval-augmented generation architectures and production patterns

View All Tags

Architecting Production-Grade AI on AWS: Security, Scale, Governance and Compliance by Design

· 56 min read
AI Playbook author

Security, scale, governance and compliance by design

An impressive AI demonstration can be built in days. A production AI system that is safe, lawful, observable, resilient, affordable and trusted by users is a different engineering problem.

The model is only one component. The complete system also includes identity, authorization, data pipelines, retrieval, orchestration, tools, human approval, policy enforcement, evaluation, audit evidence, incident response and organisational governance. If any of those layers is weak, a highly capable model can make the overall solution less reliable rather than more valuable.

This article presents a detailed AWS reference architecture for a multi-tenant enterprise AI platform that supports conversational AI, retrieval-augmented generation, deterministic workflows and bounded agentic actions. It is designed around:

  • the EU AI Act;
  • the General Data Protection Regulation (GDPR);
  • ISO/IEC 42001:2023 for AI management systems;
  • ISO/IEC 27001:2022 for information security management systems;
  • ISO/IEC 23894:2023 for AI risk management;
  • the NIST AI Risk Management Framework and its Generative AI Profile;
  • the OWASP Top 10 for LLM and GenAI applications;
  • the AWS Well-Architected Framework, its Generative AI Lens and its Agentic AI Lens.

The objective is not to claim that an AWS service creates compliance automatically. It does not. AWS describes security and compliance as a shared responsibility: AWS secures the infrastructure of the cloud, while the customer remains responsible for the design, configuration, data, identities, applications and controls it operates in the cloud. Likewise, ISO/IEC 42001 does not replace law, and certification is performed by an independent certification body rather than by ISO. AWS Shared Responsibility Model, ISO explanation of ISO/IEC 42001

Important: This is technical and governance guidance, not legal advice. Determine the laws, regulatory guidance, sector rules, contractual requirements and AWS service terms that apply to the specific organisation, countries, data, users and intended purpose. Involve legal counsel, the data protection officer, information security, risk, compliance, product owners and affected stakeholders.

Enterprise AI Solution Engineering: An End-to-End Best-Practice Playbook

· 52 min read
AI Playbook author

Enterprise AI solution engineering is not the practice of connecting a user interface to a large language model and calling the result production-ready. It is the discipline of converting a real business problem into an AI-enabled operating capability that is valuable, secure, reliable, measurable, governable and sustainable.

That requires much more than model knowledge. An AI solution engineer must work across business strategy, user experience, process design, data and knowledge architecture, AI engineering, cloud platforms, integration, cybersecurity, privacy, responsible AI, software delivery, evaluation, operations, FinOps and organizational change.

The central question is therefore not:

Which model should we use?

It is:

What business capability are we improving, what evidence will demonstrate success, what is the safest and simplest architecture that can deliver it, and how will the organization operate it responsibly at scale?

This article provides a complete reference for answering that question.

Version note: This article reflects public guidance available on 21 August 2026. Laws, standards, provider services and AI capabilities evolve quickly. Regulatory interpretations should be confirmed with qualified legal, privacy, risk and compliance specialists for the relevant jurisdiction and use case.

Production-Grade AI Architecture: A Practical Blueprint for Secure, Scalable and Governed Enterprise AI

· 68 min read
AI Playbook author

Designing an end-to-end AI platform for the EU AI Act, ISO/IEC 42001, ISO/IEC 27001, GDPR and modern GenAI security

Updated: 21 August 2026

Enterprise AI architecture is not the art of connecting an application to a large language model. A production system must remain useful when the model is uncertain, safe when retrieved content is malicious, compliant when personal data is involved, available when a provider fails, and auditable months after an output or action was produced.

That changes the architectural question from:

Which model should we use?

to:

How do we build a controlled socio-technical system in which models, data, people, policies and software work together within defined limits?

This article answers that question with a vendor-neutral reference architecture for a multi-tenant enterprise AI assistant using retrieval-augmented generation, or RAG, and governed agent tools. It covers the complete lifecycle from use-case discovery and regulatory classification through ingestion, orchestration, evaluation, deployment, monitoring, incident response and retirement.

The legal discussion is an engineering interpretation, not legal advice. The organisation's legal counsel, Data Protection Officer, information-security team and relevant sector specialists should confirm the final obligations for each use case and jurisdiction.

Knowledge Distillation for Large Language Models

· 32 min read
AI Playbook author

Knowledge distillation is a model-compression and capability-transfer technique in which a powerful teacher model provides training signals for a smaller student model.

Instead of requiring the student to rediscover every useful behaviour from raw internet-scale pretraining data, the student learns from the teacher’s outputs, probability distributions, internal representations, reasoning demonstrations or preferences.

A distilled model can therefore become:

  • Faster at inference
  • Less expensive to operate
  • Smaller in memory
  • Easier to deploy on limited hardware
  • More specialised for a particular task
  • More consistent than the original general-purpose model for a narrow workflow

However, an important distinction must be made:

The exact internal process used to train current proprietary Claude models is not publicly disclosed in full.

Anthropic publishes model reports, safety research and selected training information, but its public materials do not expose Claude’s weights, token logits, hidden states, training datasets or the complete teacher–student training recipe. Therefore, when people discuss “distilling Claude,” they may be referring to two different things:

  1. Authorised distillation within an official platform, such as Amazon Bedrock’s documented Claude Sonnet-to-Haiku distillation workflow.
  2. Black-box behavioural distillation, where permitted Claude outputs are collected and used to fine-tune another model.

Anthropic and AWS have publicly described an authorised workflow in which Claude 3.5 Sonnet generates synthetic training data, Claude 3 Haiku is trained and evaluated using that data, and the resulting distilled model is hosted for inference. This is primarily output-based behavioural distillation rather than traditional access to Claude’s internal logits or hidden layers.

LLM Response Caching Strategies for RAG and Agentic Applications

· 30 min read
AI Playbook author

Caching in a conventional application usually means storing the result of a database query or API request. In an LLM application, there are many more opportunities: repeated questions, equivalent paraphrases, identical embeddings, repeated retrieval and reranking, repeated tool calls, stable system prompts, and recurring agent planning steps.

The strongest architecture therefore does not implement one “LLM cache.” It implements a hierarchy of specialised caches across the RAG and agent pipeline.