Skip to main content
AI Playbook author
View all authors

Architecting Production-Grade AI on AWS: Security, Scale, Governance and Compliance by Design

· 56 min read
AI Playbook author

Security, scale, governance and compliance by design

An impressive AI demonstration can be built in days. A production AI system that is safe, lawful, observable, resilient, affordable and trusted by users is a different engineering problem.

The model is only one component. The complete system also includes identity, authorization, data pipelines, retrieval, orchestration, tools, human approval, policy enforcement, evaluation, audit evidence, incident response and organisational governance. If any of those layers is weak, a highly capable model can make the overall solution less reliable rather than more valuable.

This article presents a detailed AWS reference architecture for a multi-tenant enterprise AI platform that supports conversational AI, retrieval-augmented generation, deterministic workflows and bounded agentic actions. It is designed around:

  • the EU AI Act;
  • the General Data Protection Regulation (GDPR);
  • ISO/IEC 42001:2023 for AI management systems;
  • ISO/IEC 27001:2022 for information security management systems;
  • ISO/IEC 23894:2023 for AI risk management;
  • the NIST AI Risk Management Framework and its Generative AI Profile;
  • the OWASP Top 10 for LLM and GenAI applications;
  • the AWS Well-Architected Framework, its Generative AI Lens and its Agentic AI Lens.

The objective is not to claim that an AWS service creates compliance automatically. It does not. AWS describes security and compliance as a shared responsibility: AWS secures the infrastructure of the cloud, while the customer remains responsible for the design, configuration, data, identities, applications and controls it operates in the cloud. Likewise, ISO/IEC 42001 does not replace law, and certification is performed by an independent certification body rather than by ISO. AWS Shared Responsibility Model, ISO explanation of ISO/IEC 42001

Important: This is technical and governance guidance, not legal advice. Determine the laws, regulatory guidance, sector rules, contractual requirements and AWS service terms that apply to the specific organisation, countries, data, users and intended purpose. Involve legal counsel, the data protection officer, information security, risk, compliance, product owners and affected stakeholders.

Designing Production-Grade AI on Azure

· 57 min read
AI Playbook author

An end-to-end architecture for secure, scalable, observable, and governed AI under the EU AI Act, ISO/IEC 42001, ISO/IEC 27001, and GDPR

Status note, 21 August 2026: Microsoft now calls its Azure AI application and model platform Microsoft Foundry. Some older documentation and interfaces still use Azure AI Foundry. This article uses the current name while retaining familiar Azure service names such as Azure OpenAI, Azure AI Search, Azure Machine Learning, and Azure AI Content Safety.

Important: This is an engineering and governance blueprint, not legal advice or a guarantee of certification. Compliance depends on the organisation, use case, contractual roles, operating procedures, evidence, and the configuration actually deployed—not merely on choosing Azure services.

Enterprise AI Solution Engineering: An End-to-End Best-Practice Playbook

· 52 min read
AI Playbook author

Enterprise AI solution engineering is not the practice of connecting a user interface to a large language model and calling the result production-ready. It is the discipline of converting a real business problem into an AI-enabled operating capability that is valuable, secure, reliable, measurable, governable and sustainable.

That requires much more than model knowledge. An AI solution engineer must work across business strategy, user experience, process design, data and knowledge architecture, AI engineering, cloud platforms, integration, cybersecurity, privacy, responsible AI, software delivery, evaluation, operations, FinOps and organizational change.

The central question is therefore not:

Which model should we use?

It is:

What business capability are we improving, what evidence will demonstrate success, what is the safest and simplest architecture that can deliver it, and how will the organization operate it responsibly at scale?

This article provides a complete reference for answering that question.

Version note: This article reflects public guidance available on 21 August 2026. Laws, standards, provider services and AI capabilities evolve quickly. Regulatory interpretations should be confirmed with qualified legal, privacy, risk and compliance specialists for the relevant jurisdiction and use case.

Production-Grade AI Architecture: A Practical Blueprint for Secure, Scalable and Governed Enterprise AI

· 68 min read
AI Playbook author

Designing an end-to-end AI platform for the EU AI Act, ISO/IEC 42001, ISO/IEC 27001, GDPR and modern GenAI security

Updated: 21 August 2026

Enterprise AI architecture is not the art of connecting an application to a large language model. A production system must remain useful when the model is uncertain, safe when retrieved content is malicious, compliant when personal data is involved, available when a provider fails, and auditable months after an output or action was produced.

That changes the architectural question from:

Which model should we use?

to:

How do we build a controlled socio-technical system in which models, data, people, policies and software work together within defined limits?

This article answers that question with a vendor-neutral reference architecture for a multi-tenant enterprise AI assistant using retrieval-augmented generation, or RAG, and governed agent tools. It covers the complete lifecycle from use-case discovery and regulatory classification through ingestion, orchestration, evaluation, deployment, monitoring, incident response and retirement.

The legal discussion is an engineering interpretation, not legal advice. The organisation's legal counsel, Data Protection Officer, information-security team and relevant sector specialists should confirm the final obligations for each use case and jurisdiction.

Agentic Coding with Claude Code: Design the Harness, Not Only the Prompt

· 43 min read
AI Playbook author

Agentic coding is not autocomplete with a better model. A coding agent can inspect a repository, run commands, edit multiple files, execute tests, and continue until it believes the task is complete. That autonomy is valuable, but it changes the engineering problem.

The central challenge is no longer only “How do I write a good prompt?” It is:

How do I design a development environment in which an agent receives the right context, operates within safe boundaries, verifies its work, and leaves behind maintainable decisions?

Financial Services AI Adoption Plan: Strategy, Operating Model and Implementation Roadmap

· 69 min read
AI Playbook author

A financial institution should not treat AI adoption as a race to deploy the largest number of models. It should build a repeatable decision system that selects valuable use cases, supplies governed data, constrains autonomy, proves performance, assigns accountable humans, and retires systems when benefits or controls deteriorate.

Building a Trusted Data-Driven Organisation: Cataloguing, Governance, Literacy, Management, Marketplaces and Privacy

· 20 min read
AI Playbook author

Organisations rarely struggle because they have too little data. They struggle because data is difficult to find, inconsistent in meaning, uneven in quality, poorly understood, hard to access, or used without sufficient control. The result is a familiar paradox: data volumes increase while confidence in data remains low.

Architecting a Production-Grade Multi-Tenant AI Platform on AWS

· 21 min read
AI Playbook author

A platform with one million registered users does not need one-million-user capacity. It needs enough capacity for peak concurrent work, enough isolation to keep one tenant from harming another, and enough evidence to prove that each AI response is safe, grounded, affordable and attributable.

This article is an AWS-native reference architecture with explicit trade-offs — not a one-size-fits-all blueprint. It is written as a field guide for platforms serving 1M+ registered users, where the hard problems are concurrency maths, tenant isolation, inference routing, RAG authorization, unit economics and recoverability.