Skip to main content

Production-Grade AI Architecture: A Practical Blueprint for Secure, Scalable and Governed Enterprise AI

· 68 min read
AI Playbook author

Designing an end-to-end AI platform for the EU AI Act, ISO/IEC 42001, ISO/IEC 27001, GDPR and modern GenAI security

Updated: 21 August 2026

Enterprise AI architecture is not the art of connecting an application to a large language model. A production system must remain useful when the model is uncertain, safe when retrieved content is malicious, compliant when personal data is involved, available when a provider fails, and auditable months after an output or action was produced.

That changes the architectural question from:

Which model should we use?

to:

How do we build a controlled socio-technical system in which models, data, people, policies and software work together within defined limits?

This article answers that question with a vendor-neutral reference architecture for a multi-tenant enterprise AI assistant using retrieval-augmented generation, or RAG, and governed agent tools. It covers the complete lifecycle from use-case discovery and regulatory classification through ingestion, orchestration, evaluation, deployment, monitoring, incident response and retirement.

The legal discussion is an engineering interpretation, not legal advice. The organisation's legal counsel, Data Protection Officer, information-security team and relevant sector specialists should confirm the final obligations for each use case and jurisdiction.


1. What “production-grade” really means​

A successful proof of concept proves that an AI capability can work under favourable conditions. A production architecture must prove something harder: that the overall system behaves acceptably under normal use, edge cases, attack, component failure, policy change and model change.

A production-grade AI solution should have all of the following properties:

  1. Useful: It improves a measurable business or user outcome, not merely an AI metric.
  2. Grounded: It can distinguish source-backed statements from generated synthesis and uncertainty.
  3. Secure: Every user, service, data access and tool action is authenticated, authorised and observable.
  4. Private: Personal data is collected, used, retained and deleted for explicit purposes under a valid legal basis.
  5. Governed: Owners, intended purposes, prohibited uses, risk decisions and approval gates are documented.
  6. Reliable: Timeouts, retries, queues, circuit breakers, fallbacks, backups and recovery targets are designed before failure occurs.
  7. Evaluated: Quality, safety, security, fairness, privacy, performance and cost are tested against versioned release criteria.
  8. Auditable: The organisation can reconstruct which user, policy, data, prompt, model and tool chain produced a material outcome.
  9. Controllable: Humans and operators can review, override, stop, roll back or disable risky capabilities.
  10. Economically sustainable: Capacity and token costs are bounded, allocated and monitored without weakening safety or privacy.

The foundation is therefore not an LLM. It is a controlled platform around one or more models.


2. Start with the intended purpose, not the technology​

Architecture begins with a one-page use-case contract. It should state:

  • the user and affected-person groups;
  • the problem and measurable outcome;
  • the system's intended purpose;
  • decisions the system may support;
  • decisions or actions it must never make;
  • data categories and sources;
  • expected geography and jurisdictions;
  • model, retrieval and tool capabilities;
  • human-oversight points;
  • foreseeable misuse;
  • risk appetite, service targets and exit criteria.

The intended purpose matters technically and legally. A policy-search assistant and a candidate-ranking system can use the same foundation model but present radically different risks. Changing an assistant from “draft a recommendation” to “automatically reject an applicant” is not a minor feature toggle. It changes the system's decision authority, impact, regulatory classification, human-oversight design and evidence requirements.

A reusable reference scenario​

The architecture in this article assumes an enterprise AI service with these capabilities:

  • web and mobile chat;
  • business-to-business multi-tenancy;
  • organisation, user, role and attribute-based access control;
  • RAG over public, enterprise and user-authorised documents;
  • access to more than one approved model provider;
  • read-only tools by default, with tightly controlled write actions;
  • EU data residency for in-scope data;
  • human approval for material or irreversible actions;
  • full prompt, model, retrieval, policy and tool versioning;
  • privacy-aware observability and audit evidence.

Example, non-universal service objectives might be:

DimensionIllustrative target
Availability99.95% monthly for the user-facing service
Time to first tokenp95 below 2.5 seconds for ordinary queries
End-to-end responsep95 below 8 seconds for ordinary RAG queries
Recovery point objective5 minutes for operational data
Recovery time objective30 minutes for a regional service failure
Citation correctnessAt least 95% on the approved evaluation set
Grounded answer rateAt least 90% for answerable knowledge questions
Cross-tenant leakageZero accepted failures in isolation testing
Unauthorised tool actionZero accepted failures
Harmful critical outputZero accepted failures in defined critical categories
CostPer-task budget by use case, model tier and tenant

These are examples, not universal benchmarks. Targets must reflect harm, task complexity, user expectations, regulatory requirements and the organisation's risk appetite.


3. Understand how the obligations fit together​

The EU AI Act, GDPR and ISO standards do different jobs. Treating them as interchangeable creates gaps.

InstrumentPrimary purposeWhat the architecture team should do
EU AI ActRisk-based legal rules for AI systems and general-purpose AI models in the EUClassify the system and operator role; implement applicable risk, transparency, documentation, oversight, robustness and post-market controls
GDPRLegal rules for processing personal dataEstablish purposes and lawful bases; minimise data; enable rights; secure processing; manage processors, transfers, retention, DPIAs and breaches
ISO/IEC 42001:2023Requirements for an AI management systemEstablish repeatable AI policy, accountability, risk, impact, lifecycle, supplier, monitoring, audit and continual-improvement processes
ISO/IEC 27001:2022Requirements for an information-security management systemManage confidentiality, integrity and availability through risk treatment and controlled information-security operations
ISO/IEC 27701:2025Requirements and guidance for a privacy information management systemOperationalise controller and processor privacy responsibilities and integrate them into management processes
NIST AI RMFVoluntary AI risk-management frameworkUse Govern, Map, Measure and Manage to structure practical risk work and evidence
OWASP guidanceApplication, LLM, agent and software-supply-chain securityThreat-model and verify implementation-level controls against current attack patterns

ISO/IEC 42001 specifies requirements for establishing, implementing, maintaining and continually improving an AI management system. ISO/IEC 27001 does the equivalent for an information-security management system. Certification to either standard can demonstrate a managed process, but it does not by itself prove that a particular AI system complies with the AI Act, GDPR or sector law.

Useful supporting standards include ISO/IEC 23894:2023 for AI risk management, ISO/IEC 42005:2025 for AI system impact assessment, ISO/IEC 27701:2025 for privacy management and ISO/IEC 27018:2025 for protection of PII in public-cloud processing.

The current EU AI Act timeline​

As of 21 August 2026, the AI Act's general application date has passed. Most prohibited-practice provisions and the AI-literacy duty began applying on 2 February 2025, governance and general-purpose AI provisions began applying on 2 August 2025, and the Commission and national authorities began broader enforcement from 2 August 2026. Article 50 transparency duties now apply to specified interactive and generative systems. The 2026 amendment includes transitional details: newly added prohibitions concerning specified non-consensual intimate material and child sexual abuse material apply from 2 December 2026, and providers of relevant content-generating systems placed on the market before 2 August 2026 have until 2 December 2026 for the Article 50(2) marking duty. Following Regulation (EU) 2026/1744, the high-risk rules for Annex III standalone systems apply from 2 December 2027, while rules for high-risk systems embedded in Annex I regulated products apply from 2 August 2028. The authoritative sources are the consolidated AI Act, the 2026 amending regulation and the Commission's implementation overview.

The later high-risk dates are preparation time, not architecture exemptions. Systems being designed now may still be operating when those rules apply. Controls that require traceability, data lineage, human oversight or representative evaluation are expensive to retrofit.

Classify the system before building it​

For every use case, record:

  1. whether it is an AI system within the Act's scope;
  2. whether the organisation is a provider, deployer, importer, distributor, authorised representative, product manufacturer, GPAI provider or downstream provider;
  3. whether a prohibited practice is involved;
  4. whether the intended purpose falls into Annex I or Annex III high-risk categories;
  5. whether Article 50 transparency duties apply;
  6. whether a third-party model introduces GPAI value-chain duties;
  7. whether a change in purpose or substantial modification could shift the organisation into provider obligations;
  8. which national, employment, consumer, financial, health, accessibility or product laws also apply.

Do not classify only the model. Classification belongs to the system and its intended purpose. The same model may support a low-impact writing tool, a customer-facing chatbot and a high-risk employment system.


4. The end-to-end reference architecture​

The safest pattern separates four concerns:

  • experience plane: channels and user experience;
  • AI runtime plane: orchestration, retrieval, tools, models and deterministic guardrails;
  • data plane: operational, source, vector, cache and evidence stores;
  • control plane: identity, policy, governance, evaluation, deployment, observability and incident controls.

Trust boundaries​

Document explicit trust boundaries in the threat model:

BoundaryExamplesDefault assumption
User deviceBrowser, mobile app, collaboration clientPotentially compromised; all input is untrusted
EdgeCDN, WAF, gateway, rate limiterInternet exposed; withstand abuse and malformed traffic
ApplicationAPIs, identity context, orchestratorAuthenticated does not mean authorised for every resource
AI runtimeprompts, retrieved content, model outputs, agent plansProbabilistic and potentially attacker-influenced
Datarelational, object, vector, cache, logsSensitive; every access must preserve tenant and entitlement boundaries
Tool and enterprise systemsCRM, HR, payments, email, ticketingHigh impact; require explicit scopes and deterministic validation
Third-party providermodel, embedding, moderation, observability serviceExternal processor or supplier; contractual and technical boundaries apply

The essential rule is simple: natural-language instructions are never a security boundary. A system prompt telling the model not to reveal secrets cannot replace access control, network policy, schema validation or cryptographic isolation.


5. End-to-end request flow​

For a typical RAG and agent request, the controlled path is:

  1. Establish identity. The user authenticates through enterprise SSO with phishing-resistant MFA where appropriate. The gateway creates a short-lived identity context containing organisation, user, session, role, attributes, assurance level and allowed purposes.
  2. Protect the edge. The WAF and API gateway enforce request size, content type, rate, bot, origin and abuse controls. A quota service applies user, tenant, IP, model and tool budgets.
  3. Create a trace context. The service creates a correlation ID, but avoids placing raw personal data in identifiers, metrics or log labels.
  4. Validate and classify input. Deterministic checks reject malformed content, prohibited file types and oversized input. DLP and policy services identify secrets, personal data, special-category data and potential injection patterns. Detection informs controls but should not be treated as perfect.
  5. Resolve policy. A policy engine evaluates the user, tenant, purpose, data class, tool, model, region, risk tier and consent or lawful-basis metadata. The result is a machine-enforced execution policy, not prose embedded in a prompt.
  6. Retrieve authorised evidence. Retrieval applies tenant and document entitlements before search, performs hybrid keyword and vector search, reranks candidates, validates access again and returns small source passages with provenance.
  7. Construct context. The context builder separates system policy, developer instructions, user input, retrieved untrusted content and tool results. It includes only the minimum relevant information.
  8. Select a model. The model gateway selects from an approved catalogue using task, data class, geography, latency, quality, context and cost policies. It must not silently route protected data to an unapproved fallback.
  9. Generate a plan or response. If tools are needed, the model proposes structured actions. The model never holds standing credentials.
  10. Authorise each tool call. A tool gateway validates the schema, user authority, purpose, resource scope, state transition, idempotency key, risk level and approval requirement. High-impact actions require step-up authentication, user confirmation or human approval.
  11. Validate output. The system verifies structure, citations, sensitive-data policy, allowed claims and action results. Deterministic business rules take priority over model instructions.
  12. Return with transparency. The interface identifies AI interaction where required, distinguishes sources from generated text, communicates important limitations, exposes correction or escalation paths and records user feedback.
  13. Record evidence. Privacy-minimised telemetry records versions, decisions, document IDs, policy outcomes, tool calls, safety signals, latency and cost. Raw prompts and responses are not logged by default.

6. Layer-by-layer architecture decisions​

6.1 Experience and edge layer​

The front end should make safe behaviour easy:

  • identify the experience as AI-powered where Article 50 or other transparency rules apply;
  • show citations adjacent to supported claims;
  • distinguish a draft, recommendation and completed action;
  • display uncertainty and refusal in usable language;
  • require confirmation before consequential actions;
  • offer “report”, “correct”, “delete”, “appeal” and “talk to a person” routes where relevant;
  • meet applicable accessibility requirements;
  • avoid dark patterns that pressure users into sharing additional data.

At the edge, use TLS, modern security headers, strict CORS, CSRF protection for browser sessions, WAF rules, bot controls, upload limits, malware scanning and layered rate limiting. Rate limits should operate at several dimensions: network source, user, tenant, endpoint, token volume, concurrent generation, tool calls and daily spend.

6.2 Identity, tenant and policy layer​

Use the organisation's identity provider rather than creating a second identity island. Prefer OIDC or SAML federation, short-lived tokens and centrally managed joiner, mover and leaver processes.

Authentication answers “who are you?” Authorisation must answer:

  • which tenant are you acting in?
  • which data may you retrieve?
  • for which purpose?
  • which tool and action may you invoke?
  • on which resource?
  • under what approval threshold?
  • from which managed device and assurance level?

RBAC is useful for broad job roles. ABAC or relationship-based access control is normally needed for document sensitivity, region, case assignment, legal hold, ownership and purpose. Policy engines such as Open Policy Agent, Cedar-style policy services or an equivalent managed service can make decisions consistent across API, retrieval and tool layers.

For multi-tenant systems:

  • derive tenant_id from verified identity, never from an untrusted request body alone;
  • include tenant context in every primary key, index, cache key, queue message and audit event;
  • enforce database row-level security or separate databases for higher-assurance tenants;
  • use tenant-scoped encryption keys where warranted;
  • run automated cross-tenant isolation tests on every release;
  • reject rather than guess when tenant context is absent.

6.3 AI orchestrator​

The orchestrator is a state machine, not an unconstrained chat loop. It should own:

  • maximum steps, execution time, token use and cost;
  • allowed transitions and terminal states;
  • retry and timeout policy;
  • prompt and workflow versions;
  • policy results;
  • model and tool selection;
  • checkpointing and idempotency;
  • human approval and cancellation;
  • trace correlation.

For complex agents, separate planning from execution. A model may propose a typed plan, but a deterministic controller validates it before executing one bounded step at a time. Do not allow recursive delegation without explicit depth and budget limits.

6.4 Model gateway​

A model gateway prevents each team from implementing provider access differently. It should provide:

  • an approved model and embedding catalogue;
  • provider, region, version and data-class policies;
  • model-specific adapters and structured-output validation;
  • prompt-template and policy version headers;
  • quota, timeout, retry and circuit-breaker controls;
  • request and response size limits;
  • token and cost accounting;
  • content-safety integration;
  • privacy-preserving telemetry;
  • controlled fallback.

Model version aliases should not float silently in production. Pin a tested version when the provider permits it, detect provider-side changes, rerun the evaluation suite and roll out using a canary.

A fallback is a new processing path. Before enabling it, confirm equivalent data residency, DPA terms, retention settings, model-training exclusions, safety controls and task quality. When no compliant fallback exists, degrade gracefully instead of sending data somewhere unapproved.

6.5 RAG ingestion and retrieval​

RAG improves access to current enterprise knowledge, but creates a new supply chain for untrusted content.

Controlled ingestion pipeline​

  1. Register the data owner, source, purpose, lawful basis, region, sensitivity and retention class.
  2. Authenticate the connector with a dedicated least-privilege service identity.
  3. Copy through a quarantine boundary.
  4. Scan for malware, active content and unsupported formats.
  5. classify and DLP-scan the content;
  6. extract text and structure using sandboxed parsers;
  7. preserve document ACLs and security labels;
  8. normalise, chunk and attach metadata;
  9. create embeddings through an approved regional endpoint;
  10. store chunks and vectors with tenant and entitlement metadata;
  11. run quality, poisoning and access-control tests;
  12. publish the index atomically only after approval.

Every chunk should carry at least:

  • tenant and source identifiers;
  • document and version identifiers;
  • source URI or immutable reference;
  • owner and classification;
  • access-control labels;
  • valid-from and valid-to dates;
  • retention and legal-hold tags;
  • content hash and ingestion timestamp;
  • parser, chunker and embedding versions;
  • language and provenance.

Secure retrieval pipeline​

Apply access filters before similarity search wherever the data store supports it. Never retrieve across all tenants and remove unauthorised results afterward; the retrieval itself is already a disclosure surface. Use hybrid search, metadata filtering and reranking, then re-check entitlement before context assembly.

Treat retrieved text as quoted evidence, not instructions. Delimit it, label its provenance, strip or neutralise active content and instruct the runtime that documents cannot change system policy or request tools. Because prompt injection cannot be solved by prompting alone, combine this with tool isolation, least privilege, egress restrictions and deterministic approval.

6.6 Governed tool gateway​

Agent tools turn an information risk into an action risk. Each tool should expose a narrow, typed capability such as get_invoice_status rather than a generic “run SQL” or “call arbitrary URL” function.

Enforce:

  • allowlisted tools per use case and role;
  • narrowly scoped, short-lived credentials;
  • typed JSON schemas and semantic validation;
  • resource-level authorisation;
  • private network paths and egress allowlists;
  • command, URL, SQL and template injection protection;
  • read-only mode by default;
  • transaction limits and velocity rules;
  • preview or dry-run for write actions;
  • step-up authentication or human approval for sensitive actions;
  • idempotency keys and replay protection;
  • before-and-after state capture;
  • compensating action or rollback where possible;
  • an immediate kill switch.

The tool gateway should independently calculate important values. A model should not be trusted to determine a refund amount, beneficiary, eligibility result or access scope from free text.

6.7 Data stores​

Use different stores for different guarantees:

StorePurposeImportant controls
Relational databasetenants, users, permissions, cases, conversation metadata, approvalsrow-level security, encryption, constraints, audit, point-in-time recovery
Object storesource documents, evaluation sets, model artefacts, exportsversioning, malware quarantine, object lock where needed, lifecycle policy, access labels
Vector indexauthorised retrievalpre-filtered tenant and ACL isolation, deletion propagation, versioned indexes, encryption
Cachesessions, rate limits, safe response fragmentsshort TTL, tenant-bound keys, no unapproved sensitive content, encryption where supported
Audit storesecurity and material decision evidenceappend-only access, integrity protection, strict reader roles, separate retention policy
Telemetry platformmetrics, traces and operational eventsredaction, sampling, label controls, regional storage, limited raw-content capture

Encryption at rest and in transit is necessary but insufficient. Design key ownership, rotation, separation of duties, backup encryption, revocation and emergency access. Keep secrets in a managed secret store; never place credentials in prompts, source code, container images or ordinary environment dumps.

6.8 Observability without surveillance​

Operational telemetry should answer:

  • which version was running?
  • which policy decision applied?
  • which documents were retrieved?
  • which model and provider were called?
  • which tools were proposed, approved and executed?
  • how long did each stage take?
  • what did it cost?
  • did quality, safety or security signals change?

Useful event fields include a pseudonymous actor reference, tenant, use case, trace ID, policy version, workflow version, prompt-template version, model version, retrieval index version, source document IDs, tool name, approval outcome, safety decision, latency, token counts and cost.

Do not put raw prompts, responses, access tokens, emails or document passages into metric labels. Default traces to metadata-only. If content capture is temporarily needed for debugging or evaluation, require an approved purpose, sampled scope, restricted access, encryption, short retention and automatic expiry.

Use OpenTelemetry or an equivalent open schema to avoid hard-coding observability into one vendor. Separate operational traces from immutable compliance evidence; their audiences and retention periods are different.


7. GDPR by design for AI systems​

The GDPR is technology-neutral and continues to apply alongside the AI Act. The authoritative text is Regulation (EU) 2016/679.

7.1 Build a data-processing map​

Before development, map each flow:

data subject → channel → application → retrieval → model provider → tool/API → telemetry → backup → deletion

For every hop, record:

  • controller, joint controller and processor roles;
  • purpose and lawful basis under Article 6;
  • any Article 9 special-category condition;
  • data categories and affected people;
  • source, recipients and sub-processors;
  • region and transfer mechanism;
  • retention and deletion behaviour;
  • security measures;
  • rights-handling route.

“Improving the model” is too vague to be a purpose. Separate service delivery, abuse prevention, product analytics, human review, evaluation and training. Disable provider training or retention when it is not explicitly approved and lawful.

7.2 Apply the GDPR principles in architecture​

GDPR principleArchitectural implementation
Lawfulness, fairness and transparencyPurpose and lawful-basis registry; privacy notice; AI disclosure; meaningful explanation and escalation
Purpose limitationPurpose-bound access policies; separate indexes and pipelines; prevent secondary training by default
Data minimisationredact before prompts; retrieve small passages; avoid full records when fields suffice; short context windows by design
Accuracysource ownership, versioning, correction workflow, stale-document expiry, user challenge route
Storage limitationTTLs, retention classes, deletion jobs, backup expiry and vendor deletion verification
Integrity and confidentialityleast privilege, encryption, tenant isolation, secure SDLC, monitoring and incident response
AccountabilityROPA, DPIA, decision log, policy evidence, contracts, test results and audit trail

Pseudonymised data remains personal data when re-identification is reasonably possible. Do not label data anonymous merely because a direct identifier was removed.

7.3 Data protection impact assessment​

A DPIA is required where processing is likely to result in high risk to people's rights and freedoms. The European Commission highlights systematic and extensive evaluation or profiling, large-scale sensitive-data processing and large-scale public monitoring as clear examples. See the Commission's DPIA guidance and GDPR Article 35.

A practical AI DPIA should cover:

  • necessity and proportionality;
  • people and groups affected;
  • data provenance and expectations;
  • model and retrieval uncertainty;
  • hallucination, bias and automation bias;
  • data leakage, memorisation and re-identification;
  • employee or customer power imbalance;
  • solely automated decisions;
  • children's and vulnerable people's interests;
  • security threats and supplier dependencies;
  • human intervention, appeal and redress;
  • residual risk and approval.

Review the DPIA when the purpose, data, model, autonomy, user population, geography or risk changes.

7.4 Automated decisions and meaningful human oversight​

GDPR Article 22 restricts decisions based solely on automated processing that produce legal or similarly significant effects, subject to defined exceptions and safeguards. “A human can review it” is not sufficient if the person merely rubber-stamps the output.

Meaningful oversight requires a reviewer who:

  • has authority and time to change the result;
  • can see relevant source evidence and limitations;
  • is trained to detect automation bias;
  • is not measured in a way that forces automatic agreement;
  • can request more information;
  • records reasons for accepting or overriding the recommendation;
  • provides an accessible challenge and redress route.

The EDPB maintains official automated decision-making and profiling guidance.

7.5 Rights, retention and deletion​

Design data-subject rights as platform capabilities, not manual database exercises. A rights service should locate data across conversations, source documents, relational rows, vector chunks, caches, annotations, evaluation sets, exports and relevant processors.

Deletion should:

  1. verify identity and scope;
  2. apply exceptions such as legal hold;
  3. delete or irreversibly anonymise primary records;
  4. remove vector chunks and rebuild affected indexes if necessary;
  5. purge caches and queued copies;
  6. send deletion requests to processors;
  7. prevent re-ingestion from the source;
  8. record completion evidence without retaining deleted content;
  9. allow encrypted backups to age out under a documented, access-restricted schedule.

7.6 International transfers and suppliers​

For every model, embedding, moderation, OCR and telemetry provider, confirm:

  • processor terms and Article 28 requirements;
  • sub-processor list and change notice;
  • data location and support access;
  • retention, deletion and training settings;
  • security and incident commitments;
  • audit evidence;
  • international-transfer mechanism, including adequacy or Standard Contractual Clauses where applicable;
  • exit and data-portability plan.

The Commission describes adequacy, SCCs and binding corporate rules in its international data protection overview.


8. Operationalising the EU AI Act​

8.1 Controls applicable across many systems​

Even where a use case is not high-risk, teams should address:

  • prohibited-practice screening;
  • AI literacy for staff and people operating AI on the organisation's behalf;
  • Article 50 disclosure for direct AI interaction where applicable;
  • machine-readable marking and labelling duties for applicable synthetic or manipulated content;
  • clear provider/deployer and supplier responsibilities;
  • complaint, incident and authority-cooperation procedures;
  • monitoring for a change of purpose, substantial modification or reclassification.

Article 4 requires providers and deployers to take measures supporting AI literacy appropriate to people's knowledge, experience, training and context. The consolidated Act contains the current Article 4 text. Training records, role-specific curricula and competence checks are useful evidence.

The Commission's Article 50 transparency guidance states that providers of systems interacting directly with people must design them so people are informed that they are interacting with AI, subject to the provision's exceptions and context.

8.2 Engineer the high-risk requirements now​

For systems that are or may become high-risk, design for the Chapter III controls even before their applicable date:

AI Act areaEngineering responseTypical evidence
Article 9 risk managementcontinuous hazard, misuse and residual-risk process across the lifecyclerisk register, test plan, acceptance records
Article 10 data and data governanceprovenance, suitability, representativeness, quality, bias controls and documented preparationdataset records, lineage, quality reports
Article 11 technical documentationversioned system description, design, limitations, validation, monitoring and changessystem card, architecture, model and data documentation
Article 12 record-keepingautomatic logs sufficient for traceability and monitoringprotected logs, log schema, retention policy
Article 13 transparency to deployersclear instructions, purpose, performance, limitations, oversight and maintenanceinstructions for use, known-limitations register
Article 14 human oversighttrained humans can understand, monitor, disregard, override or stop the systemoversight SOP, UI evidence, training, override logs
Article 15 accuracy, robustness and cybersecuritydeclared metrics, resilience, fail-safe behaviour and AI-specific security testingevaluation, resilience and red-team reports
Articles 16–18 provider duties and QMScontrolled development, compliance, documentation and change managementpolicies, approvals, QMS records
Article 26 deployer dutiesuse according to instructions, oversight, monitoring and logsdeployment controls, operating procedures
Article 27 FRIA where applicableassess effects on fundamental rights before deploymentfundamental-rights impact assessment
Post-market monitoring and incidentscollect performance and safety signals, investigate and correctmonitoring plan, incident and corrective-action records

The consolidated Act specifies lifecycle risk management, testing against predefined metrics, data governance, logging, transparency, effective human oversight, and accuracy, robustness and cybersecurity requirements. It expressly identifies data poisoning, model poisoning, adversarial examples, model evasion and confidentiality attacks as relevant AI-security concerns. See Articles 9–15 in the consolidated text.

8.3 Do not confuse a DPIA, FRIA and AI impact assessment​

These assessments overlap but answer different questions:

  • a DPIA focuses on risks from personal-data processing under GDPR;
  • a fundamental-rights impact assessment applies to specified deployers and high-risk uses under Article 27;
  • an AI system impact assessment, such as the method supported by ISO/IEC 42005, can cover wider effects on individuals, groups and society;
  • a security threat model focuses on adversaries, assets, attack paths and controls;
  • a business risk assessment covers financial, operational, strategic and reputational effects.

Use a shared facts section and risk taxonomy, then maintain the legally required decisions and owners separately. One giant generic checklist can obscure which obligation has actually been satisfied.


9. Integrating ISO/IEC 42001 and ISO/IEC 27001​

The cleanest operating model is an integrated management system: one control is designed once, assigned once and evidenced once, then mapped to multiple obligations.

9.1 ISO/IEC 42001 management cycle​

The management-system clauses can be operationalised as:

Management areaPractical implementation
Contextscope, stakeholders, laws, AI inventory, internal and external issues
Leadershipaccountable executive, AI policy, roles, risk appetite and escalation
PlanningAI risks and opportunities, objectives, impact criteria and treatment plans
Supportcompetence, AI literacy, resources, communication and controlled documents
Operationuse-case intake, impact assessment, lifecycle controls, suppliers and change management
Performance evaluationmetrics, monitoring, internal audit and management review
Improvementincidents, nonconformities, corrective action and continual improvement

ISO/IEC 42001 Annex A control themes cover AI policies, internal organisation, resources, impact assessment, system lifecycle, data, information for interested parties, use of AI and third-party relationships. These should appear in the platform's operating model, not only in a certification folder.

9.2 ISO/IEC 27001 security foundation​

ISO/IEC 27001:2022 requires an ISMS with context, leadership, planning, support, operation, performance evaluation and improvement, backed by a risk treatment process and Statement of Applicability. Its 93 Annex A controls are grouped into organisational, people, physical and technological controls.

For an AI platform, the ISMS scope should explicitly cover:

  • model and embedding providers;
  • prompt, workflow and policy repositories;
  • training, evaluation and RAG data;
  • vector databases and indexes;
  • annotation and human-review platforms;
  • agent tool credentials and enterprise APIs;
  • AI telemetry and audit evidence;
  • development, CI/CD and model-evaluation environments;
  • third-party libraries, models and datasets.

9.3 One control, multiple obligations​

For example, an immutable, access-controlled AI audit event can support:

  • AI Act traceability and log requirements;
  • GDPR accountability and security evidence;
  • ISO/IEC 42001 monitoring and controlled records;
  • ISO/IEC 27001 logging, monitoring and incident management;
  • internal investigation and customer assurance.

But design the event to minimise personal data. “Keep everything forever in case of audit” conflicts with storage limitation and increases breach impact.


10. Threat model and security controls​

OWASP's current resources include the Top 10 for LLM Applications 2026, the Top 10 for Agentic Applications 2026, ASVS 5.0 and the Software Component Verification Standard. Use them with ordinary application and cloud threat modelling; AI-specific risk does not replace web, API, identity, supply-chain or infrastructure security.

ThreatExamplePreventive controlsDetective and recovery controls
Direct prompt injectionuser asks the model to ignore policy and expose secretsdeterministic access control, prompt separation, minimal context, no secrets in promptsadversarial tests, injection signals, trace review
Indirect prompt injectionmalicious instruction hidden in a webpage or documenttreat retrieved text as data, sandbox parsing, content provenance, tool isolation, egress allowlistsource anomaly detection, canary documents, quarantine and index rollback
Cross-tenant leakagevector search returns another customer's documentpre-filter by verified tenant and ACL, RLS or physical isolation, tenant-bound cache keysautomated isolation tests, DLP alerts, kill switch
Sensitive information disclosuremodel returns PII, credentials or confidential textminimisation, redaction, data-class policies, retrieval ACLs, provider privacy settingsoutput DLP, access audit, incident workflow
RAG poisoningcompromised source inserts false or hostile contentauthenticated connectors, owner approval, signing or hashing, versioned atomic indexessource-drift monitoring, retrieval-quality alerts, rollback
Model or dataset supply-chain compromisemalicious model, library, adapter or datasetapproved registry, provenance, licence review, signed artefacts, SBOM and AI BOM, isolated scanningvulnerability monitoring, integrity verification, rapid revoke
Improper output handlinggenerated HTML, SQL or command is executedstrict encoding, parameterised operations, typed schemas, never eval, output treated as untrustedapplication security tests and runtime alerts
Excessive agencyagent sends money, email or data without valid approvalnarrow tools, read-only default, least privilege, step and cost limits, human approvalaction audit, reconciliation, transaction alerts, rollback
SSRF and egress abuseagent tool requests cloud metadata or attacker URLno generic fetch tool, URL allowlist, private DNS policy, metadata endpoint protectionnetwork flow logs, egress alerts, credential rotation
Model extraction or inversionrepeated queries recover behaviour or memorised dataauthentication, rate and similarity limits, minimise exposed confidence detail, privacy testsabuse analytics, account suspension, provider coordination
Denial of walletattacker triggers long contexts and recursive toolsrequest, token, concurrency, step and spend budgets; backpressurecost anomaly alerts, circuit breaker, tenant suspension
Hallucinated authorityassistant invents law, policy or account stateretrieval for factual claims, citations, abstention, deterministic source APIssampled review, groundedness and citation metrics
Memory poisoningmalicious content is stored as durable user or agent memoryexplicit memory-write policy, typed memory, provenance, user review, TTLmemory-diff review, delete and rebuild
Insecure human approvalreviewer sees only an AI summary and rubber-stamps itshow source evidence, consequence and changed fields; separation of dutiesoverride/approval analytics and quality review
Telemetry leakageprompts and documents are copied to logsmetadata-only default, redaction, restricted debug captureDLP scanning, log-access monitoring, purge runbook
Insider misuseprivileged operator retrieves customer contentjust-in-time access, dual approval, session recording, segregation of dutiesUEBA, immutable access logs, periodic review

Zero trust for AI​

NIST SP 800-207 moves trust away from network location and toward users, assets and resources. Apply this literally:

  • every service has a workload identity;
  • every call is authenticated and authorised;
  • credentials are short-lived;
  • policies consider user, device, tenant, workload and data sensitivity;
  • model and tool providers are accessed through controlled gateways;
  • production data is not reachable from developer laptops by default;
  • administrative access is just-in-time, approved and recorded;
  • east-west traffic is restricted, not implicitly trusted because it is inside a VPC.

Secure software and AI supply chain​

Follow the final NIST SSDF 1.1 and its AI-specific SP 800-218A profile:

  • protect source and build systems;
  • review code and infrastructure as code;
  • scan dependencies, containers, secrets and licences;
  • produce an SBOM plus an AI bill of materials listing models, embeddings, datasets, prompts, evaluators and providers;
  • sign artefacts and verify provenance at deployment;
  • isolate untrusted model and document-processing code;
  • patch or revoke vulnerable components;
  • maintain reproducible builds and rollback artefacts.

11. Evaluation is the release contract​

LLM quality cannot be represented by one aggregate score. Build a versioned evaluation portfolio.

Evaluation dimensions​

DimensionExample measures
Businesstask completion, first-contact resolution, handling time, conversion, user effort
Answer qualitycorrectness, completeness, relevance, instruction following, calibrated abstention
RAGretrieval recall, precision at K, context relevance, groundedness, citation coverage and citation correctness
Safetypolicy violation rate by severity, vulnerable-group scenarios, harmful completion rate
PrivacyPII leakage, memorisation probes, deletion effectiveness, data-minimisation conformance
Fairnesserror and outcome differences across relevant groups and languages; accessibility impact
Securityprompt injection success, cross-tenant access, tool abuse, exfiltration, poisoning and denial-of-wallet resistance
Agent behaviourvalid tool selection, argument correctness, unauthorised action rate, loop rate, recovery rate
Reliabilityavailability, timeout, dependency failure, fallback correctness, queue age, recovery time
Performancetime to first token, total latency, retrieval latency, tool latency, throughput
Economicstokens, provider cost, compute, storage and support cost per completed task

Build representative test sets​

Include:

  • normal high-volume tasks;
  • rare but high-impact cases;
  • unanswerable and ambiguous questions;
  • stale and conflicting sources;
  • long, multilingual and accessibility-related inputs;
  • relevant demographic and vulnerable-group scenarios;
  • malicious files, direct and indirect injections;
  • permission and cross-tenant tests;
  • provider timeouts and partial failures;
  • unsafe or irreversible tool requests;
  • known production incidents converted into regression tests.

Use expert human review for high-impact outputs. LLM-as-judge can scale comparison, but validate the judge against expert labels, monitor positional and self-preference bias, and never make it the only release authority for critical safety or legal criteria.

Example release gates​

GateExample pass condition
Architectureapproved threat model, data-flow diagram and failure-mode review
Privacylawful basis confirmed; DPIA approved where required; deletion tested end to end
Securityno open critical findings; tenant isolation and tool-authorisation tests pass
Qualityuse-case thresholds pass with confidence intervals and segment breakdowns
Safetyzero critical policy failures in the approved adversarial suite
Human oversightreviewers can understand, override, stop and record reasons
Reliabilityload, chaos, backup restore, failover and rollback tests pass
Complianceinventory, system card, instructions, contracts and evidence pack complete
Operationsdashboards, alerts, on-call, runbooks, SLOs and kill switches verified
Businessnamed owner accepts expected benefit, cost and residual risk

Do not average away a critical failure. A system with excellent overall accuracy but one cross-tenant disclosure has failed release.


12. Production deployment and LLMOps​

A controlled pipeline should promote versioned bundles rather than independent moving parts. A release bundle can contain:

  • application and infrastructure artefact digests;
  • workflow graph version;
  • system and developer prompt versions;
  • policy bundle version;
  • model and embedding versions;
  • retrieval index and source snapshot;
  • tool schemas and allowed scopes;
  • evaluation dataset and result IDs;
  • risk and approval records.

Suggested delivery flow​

  1. Developer checks: formatting, unit, type, secret and dependency checks.
  2. Build: reproducible container and SBOM generation; artefact signing.
  3. Component tests: prompt templates, parsers, policy, retrieval and tool schemas.
  4. Integration tests: approved sandbox models, stores and enterprise APIs.
  5. AI evaluation: quality, safety, privacy, fairness, security, performance and cost.
  6. Security assurance: SAST, DAST, IaC, container and adversarial AI testing.
  7. Approval: business owner, model-risk or responsible-AI function, security, privacy and legal as required by tier.
  8. Staging: production-like data controls with synthetic or minimised data.
  9. Canary: small cohort, low-risk tools, close monitoring and automatic rollback thresholds.
  10. Progressive rollout: expand by tenant or traffic percentage after evidence review.
  11. Post-release review: compare expected and actual quality, safety, latency, cost and incidents.

Separate development, test and production accounts and keys. Production prompts, policies and configurations should be changed through reviewed version control, not a live console without traceability.


13. Reliability, scaling and graceful degradation​

LLM systems combine multiple failure-prone dependencies. Engineer bulkheads so a slow model does not exhaust the entire API worker pool and a failing tool does not block ordinary knowledge queries.

Core resilience patterns​

  • explicit per-stage timeouts;
  • retries only for safe, transient and idempotent operations;
  • exponential backoff with jitter;
  • circuit breakers per provider, model and tool;
  • bounded queues and dead-letter handling;
  • concurrency pools per tenant and workload;
  • backpressure before saturation;
  • idempotency for user requests and write actions;
  • checkpoints for long-running workflows;
  • health probes that test dependencies appropriately;
  • multi-zone deployment and tested regional recovery;
  • versioned backups and restoration exercises.

For streaming responses, distinguish connection success from task success. If a tool or citation check fails after partial generation, the interface needs a clear terminal error rather than silently presenting an incomplete answer as final.

Safe degradation ladder​

An agentic system can degrade in stages:

  1. full agent with approved read and write tools;
  2. read-only tools;
  3. RAG answer with citations;
  4. search results without synthesis;
  5. static help content;
  6. human handoff.

The model gateway may route to another provider only where its legal, privacy, residency, quality and safety profile is already approved.

Privacy-safe caching​

Cache only when the key includes all relevant security and correctness dimensions: tenant, user or entitlement scope, data version, model version, prompt version, policy version and locale. Never use a global semantic cache for responses based on private data. Encrypt sensitive cache content, apply short TTLs and purge entries when source permissions or documents change.


14. Monitoring, incidents and post-market operation​

Monitor four types of drift​

  1. Data drift: user, source, language or document distributions change.
  2. Quality drift: groundedness, citations, refusal or task success changes.
  3. Control drift: policies, permissions, vendors or configurations diverge from approval.
  4. Impact drift: the system is used for a different purpose, population or level of autonomy.

Alert on symptoms that require action, not every metric fluctuation. Examples include a critical safety regression, cross-tenant access attempt, anomalous write-tool volume, abrupt retrieval-quality drop, cost spike, provider-region change, repeated human overrides or unauthorised configuration change.

AI incident runbook​

  1. Detect and preserve evidence. Create an incident ID and protect relevant versions and events without expanding unnecessary personal-data access.
  2. Triage harm and scope. Determine affected people, tenants, actions, data, geography and continuing risk.
  3. Contain. Disable a tool, model, source, index, tenant capability or the whole AI route using tested kill switches.
  4. Provide a safe service. Fall back to read-only, static content or human handling.
  5. Eradicate. Remove poisoned content, rotate credentials, patch components, correct policy or retrain where justified.
  6. Recover. Restore a known-good bundle, canary it and monitor closely.
  7. Notify. Involve security, privacy, legal, the business owner, affected customers and authorities under the applicable rules.
  8. Learn. Complete root-cause analysis, corrective action, risk updates and regression tests.

Under GDPR Article 33, a controller must notify the competent supervisory authority without undue delay and, where feasible, within 72 hours after becoming aware of a personal-data breach, unless it is unlikely to result in a risk to people's rights and freedoms. High-risk cases can also require communication to affected people. See the GDPR breach provisions. Do not wait for perfect technical certainty before involving the DPO and incident team.


15. Required governance artefacts and evidence​

Maintain evidence as part of delivery, not as a retrospective documentation sprint.

ArtefactMinimum contentsPrimary owner
AI system inventory entryowner, purpose, users, regions, model, data, tools, risk tier, statusAI governance
Use-case contractoutcomes, intended purpose, excluded uses, autonomy and human oversightproduct owner
Regulatory classificationAI Act role and risk, GDPR scope, sector obligations, review triggerslegal/compliance
Data-flow and processing recordsystems, data categories, roles, purposes, transfers, retentionprivacy/data owner
DPIA/FRIA/AI impact assessmentaffected groups, risks, mitigations, residual decisionsDPO/responsible AI
System cardarchitecture, versions, capabilities, limitations, metrics and intended usetechnical owner
Data and retrieval recordsources, rights, lineage, quality, ACLs, index and deletiondata owner
Model/provider recordmodel card, terms, location, training, retention, sub-processors, evaluationmodel owner/procurement
Threat modelassets, trust boundaries, abuse cases, controls and residual riskssecurity architect
Evaluation reportdatasets, methods, segment results, failures, thresholds and approvalsevaluation lead
Human-oversight procedurereviewer authority, evidence, override, stop, escalation and trainingoperations owner
Deployment recordsigned bundle, approvals, canary, rollback and change reasonplatform owner
Monitoring planSLOs, risk metrics, thresholds, review cadence and ownersservice owner
Incident and continuity planscenarios, contacts, kill switches, notification, RTO and RPOincident/service owner
Decommission recordaccess removal, data disposal, vendor exit and retained evidencesystem owner

Each artefact should have an owner, version, approver, review date, evidence links and change triggers.


16. End-to-end enterprise case study​

Consider a retail bank building a customer-service AI assistant for EU customers. The assistant answers product and fee questions using approved policy documents, retrieves customer-specific case status after authentication, and can initiate a small set of service workflows. It does not make credit, pricing, fraud, employment or legal decisions.

Phase 1: discovery and classification​

The bank defines three outcomes:

  • reduce average handling time by 20%;
  • achieve at least 90% correct, source-backed answers on supported intents;
  • improve successful self-service without reducing complaint resolution or accessibility.

The use-case contract prohibits lending decisions, financial advice, vulnerability inference, payment initiation and autonomous complaint rejection. Legal and compliance determine that the customer is interacting with an AI system and Article 50 transparency applies. GDPR applies because chat, account and case information contain personal data. A DPIA is initiated because of the scale, monitoring and potential impact.

The classification record says that adding creditworthiness or insurance-pricing decisions would trigger a new assessment and may move the system into an Annex III high-risk use. That change cannot be enabled through ordinary feature configuration.

Phase 2: data and supplier controls​

The team inventories:

  • public product pages;
  • controlled policy manuals;
  • fee tables;
  • customer account and case-status APIs;
  • chat metadata and feedback;
  • the model, embedding, moderation and observability suppliers.

Documents receive an owner, approval status, valid dates, sensitivity, ACL and retention class. Expired prices are automatically removed from the active index. Contracts prohibit provider training on bank prompts and responses, define regional processing, sub-processors, deletion, incident notice and audit evidence.

Phase 3: architecture​

The system uses:

  • enterprise customer identity and step-up authentication;
  • a WAF and API gateway with user and tenant quotas;
  • a stateless conversation API;
  • a workflow orchestrator with bounded execution;
  • a policy service that combines user assurance, purpose, data class and tool risk;
  • separate public and customer-authorised retrieval collections;
  • a model gateway with EU-approved endpoints;
  • a tool gateway exposing only typed bank APIs;
  • metadata-only traces and a separate protected audit store.

The model never receives core-banking credentials. The account-status tool derives the customer relationship from the authenticated session and returns only necessary fields. Address changes require step-up authentication, a preview of the exact change and explicit confirmation. Card freezing is handled by an existing deterministic bank workflow; the AI may navigate the user to it but cannot bypass its controls.

Phase 4: a live request​

A customer asks: “Why was I charged £12 and can you refund it?”

  1. The gateway verifies the session and applies quotas.
  2. The input classifier detects financial and personal context.
  3. The policy engine allows read-only account and fee-policy tools but denies refund execution.
  4. The account API returns a typed transaction category, not an unrestricted statement export.
  5. Retrieval fetches the current fee policy and terms, applying version and regional filters.
  6. The model produces a draft explanation with citations.
  7. The output validator checks that the fee and eligibility claims are supported.
  8. The assistant explains the charge, says it cannot decide the refund, and offers a controlled complaint or human-review route.
  9. The trace records document IDs, policy, model and tool versions without copying the full transaction description into ordinary telemetry.

This is safer than asking one model to read an entire account, decide what happened and directly issue money.

Phase 5: evaluation and rollout​

The bank builds a test set of normal, ambiguous, vulnerable-customer, multilingual, stale-policy, injection and unauthorised-account cases. It measures answer correctness, citation correctness, unsupported claim rate, privacy leakage, tool authorisation, accessibility, latency and cost.

Release requires:

  • zero cross-customer access failures;
  • zero unauthorised write actions;
  • no critical harmful output in the approved test suite;
  • target citation and task-quality thresholds by language and customer segment;
  • successful rollback, provider outage and deletion tests;
  • approved DPIA, operating instructions and incident runbook.

The first release is read-only for a small cohort. Write-capable workflows are introduced separately after additional authentication, authorisation and human-factors testing.

Phase 6: operation​

Dashboards track groundedness samples, user escalation, document staleness, policy denials, human overrides, provider latency, token cost and complaint signals. A model-version change automatically blocks production promotion until the regression suite passes. A source-poisoning drill proves that the team can quarantine a connector and atomically restore the prior index.

The result is not “a chatbot with compliance added.” It is a controlled service in which AI is one component.


17. A practical implementation roadmap​

Weeks 1–2: scope and risk​

  • appoint business, technical, data, security, privacy and operational owners;
  • define outcome, intended purpose, excluded use and authority boundaries;
  • classify under the AI Act and relevant sector rules;
  • map data and suppliers;
  • start DPIA, FRIA or wider impact assessment as applicable;
  • define risk appetite and release tier.

Weeks 3–5: platform foundation​

  • establish SSO, tenant model and policy engine;
  • deploy edge, secrets, key and private-network controls;
  • implement model gateway and approved catalogue;
  • create prompt, policy, model and data registries;
  • establish privacy-safe telemetry and audit schemas;
  • build CI/CD, signing, SBOM and infrastructure-as-code controls.

Weeks 6–8: RAG and tools​

  • implement quarantined ingestion, lineage and document ACL preservation;
  • build secure retrieval, reranking and citations;
  • expose narrow read-only tools through the tool gateway;
  • add bounded orchestration, input and output controls;
  • implement rights, retention and deletion workflows.

Weeks 9–11: assurance​

  • create representative evaluation sets;
  • conduct application, AI and agent threat testing;
  • test tenant isolation, prompt injection, poisoning and denial of wallet;
  • run load, chaos, backup, restore, regional failover and rollback tests;
  • validate human oversight and accessibility;
  • complete supplier, legal, privacy and responsible-AI approvals.

Week 12 and beyond: controlled operation​

  • canary to a small, low-risk cohort;
  • review actual outcomes against assumptions;
  • expand progressively;
  • run periodic access, supplier, model, data, risk and management reviews;
  • convert incidents and near misses into regression tests;
  • decommission data, permissions and suppliers when the use case ends.

18. Final production checklist​

Before launch, a service owner should be able to answer “yes” to all of these:

Purpose and governance​

  • Is the intended purpose precise, approved and visible to the delivery team?
  • Are prohibited uses and escalation triggers enforceable?
  • Are provider, deployer, controller and processor roles documented?
  • Is the AI Act risk classification current and reviewed after material changes?
  • Are AI literacy and operating competence appropriate to each role?

Privacy and data​

  • Are purpose, lawful basis, Article 9 condition and notices confirmed where relevant?
  • Is a DPIA completed where required?
  • Are source rights, provenance, quality, ACLs and retention documented?
  • Is personal data minimised before retrieval and model calls?
  • Do access, correction, objection, restriction, portability and deletion processes reach all stores and processors?
  • Are transfers, sub-processors and provider-training settings approved?

Security​

  • Are user and workload identities authenticated with least privilege?
  • Is tenant isolation enforced in database, vector, cache, queue and tool paths?
  • Are retrieved content and model output treated as untrusted?
  • Are tools narrow, typed, independently authorised and bounded?
  • Are egress, secrets, keys and production administration controlled?
  • Have application, LLM, RAG, agent and supply-chain threats been tested?

Quality and safety​

  • Are evaluation datasets representative, adversarial and versioned?
  • Are critical criteria treated as hard gates rather than averaged scores?
  • Can the system abstain, cite, escalate and fail safely?
  • Is human oversight meaningful and usable?
  • Are limitations communicated to users and operators?

Reliability and operation​

  • Are SLOs, budgets, quotas, timeouts, retries and circuit breakers defined?
  • Are backup, restore, failover, rollback and kill switches tested?
  • Are fallback providers legally and technically approved?
  • Are privacy-safe dashboards, alerts, on-call and runbooks live?
  • Can the team reconstruct a material output or action from protected evidence?
  • Is there a monitored decommission and vendor-exit plan?

19. A concrete, portable implementation stack​

A reference architecture is only useful when the implementation team can map each responsibility to an actual technology decision. The following stack is an example, not a universal prescription.

CapabilityPortable implementation patternWhy it belongs in the architecture
Web applicationReact or Next.js with a server-side backend-for-frontendcontrols session boundaries, response streaming, citation rendering and user confirmation
Public APIFastAPI, Go, Java/Spring or ASP.NET with a versioned OpenAPI contractprovides predictable authentication, request validation, streaming and operational endpoints
Edge and gatewaymanaged API gateway or Envoy/Kong-class gatewaycentralises WAF integration, quotas, request limits and service identity
Identityenterprise OIDC/SAML provider plus workload identitypreserves joiner/mover/leaver processes and avoids long-lived application secrets
Authorisationpolicy engine implementing RBAC plus attribute or relationship checksmakes tenant, purpose, document and tool decisions enforceable outside the model
Orchestrationexplicit state machine; durable workflow engine for long-running jobsbounds autonomy, makes approvals resumable and provides deterministic state transitions
Relational dataPostgreSQL-compatible managed databasesupports transactional integrity, row-level security and a strong operational ecosystem
Vector searchPostgreSQL with pgvector initially; dedicated search/vector engine when requiredbalances operational simplicity against retrieval scale, filtering and hybrid-search needs
Source documentsobject storage with versioning and lifecycle policiespreserves immutable sources, parser inputs and auditable document history
Cache and quotasRedis-compatible managed cachesupports short-lived sessions, distributed rate limits, idempotency and safe caching
Async processingmanaged queue, Kafka-compatible event stream or workflow queueisolates ingestion, evaluation, notification and recovery workloads
Model accessapproved internal model gatewaycentralises model choice, region, retention, cost, safety and failover controls
Tool executionseparate typed tool gateway with dedicated workload identitiesprevents a model from inheriting broad enterprise credentials
ObservabilityOpenTelemetry plus metrics, logs, traces and a protected audit storecorrelates reliability and AI behaviour while preserving portability
Secrets and keysmanaged KMS/HSM and secret managersupports key ownership, rotation, envelope encryption and emergency revocation
InfrastructureTerraform, OpenTofu, Bicep, CloudFormation or equivalentmakes infrastructure reviewable, repeatable and policy-checkable

A single PostgreSQL platform with vector search can be an excellent first production choice when:

  • the corpus is modest enough for the selected database tier;
  • tenant and access filters fit naturally into SQL;
  • the team values transactional consistency between document metadata and embeddings;
  • hybrid search and reranking requirements can be implemented without excessive complexity;
  • operational simplicity is more valuable than an additional specialised datastore.

Move to a dedicated vector or search service when you need materially larger indexes, specialised hybrid ranking, high concurrent query volume, document-level permissions, geographically distributed search, advanced retrieval pipelines or independent scaling of ingestion and serving.

The correct decision is measurable: benchmark retrieval quality, p95/p99 latency, filtered-query performance, tenant isolation, deletion propagation, index rebuild time and total operating cost against a realistic corpus.

Keep orchestration separate from durability​

An agent framework can help express a retrieval or tool graph. A durable workflow engine provides a different guarantee: recovering state across crashes, long approval windows and retries. A mature solution may use both, but should avoid hiding material business transactions inside an unbounded prompt loop.

For example:

  1. the AI graph drafts a refund recommendation;
  2. the durable workflow records the draft and waits for approval;
  3. a human reviews the source evidence and amount;
  4. a deterministic payment service validates policy and executes the approved action;
  5. the workflow records completion or compensates for failure.

20. Cross-cloud implementation mapping​

The logical architecture should remain stable even when the cloud provider changes. Map capabilities rather than assuming similarly named services have identical semantics, regional availability, residency guarantees or compliance features.

Architecture capabilityAWS exampleAzure exampleGoogle Cloud example
Internet edge and WAFCloudFront, AWS WAF and API GatewayAzure Front Door, WAF and API ManagementCloud CDN, Cloud Armor and Apigee or API Gateway
Enterprise identityIAM Identity Center, Amazon Cognito and IAM rolesMicrosoft Entra ID and managed identitiesCloud Identity, Identity Platform and workload identities
Container runtimeAmazon EKS, ECS or FargateAzure Kubernetes Service or Container AppsGoogle Kubernetes Engine or Cloud Run
Foundation-model accessAmazon Bedrock or SageMaker AIMicrosoft Foundry and Azure OpenAI modelsVertex AI and its model catalogue
AI gatewayAPI Gateway, internal gateway or Bedrock AgentCore GatewayAzure API Management AI gatewayApigee or an internal model gateway in front of Vertex AI
Governed agent executionAmazon Bedrock AgentCoreMicrosoft Foundry agent capabilitiesGemini Enterprise Agent Platform
RetrievalBedrock Knowledge Bases, OpenSearch or PostgreSQL vectorsAzure AI Search, Foundry IQ or PostgreSQL vectorsVertex AI retrieval, AlloyDB or a managed search/vector service
Relational databaseAmazon Aurora PostgreSQL or Amazon RDS for PostgreSQLAzure Database for PostgreSQLAlloyDB or Cloud SQL for PostgreSQL
Source documentsAmazon S3Azure Blob StorageCloud Storage
Caching and quotasElastiCache for RedisAzure Managed Redis or an approved Redis-compatible serviceMemorystore for Redis
Asynchronous workAmazon SQS, EventBridge or MSKService Bus, Event Grid or Event HubsPub/Sub or Cloud Tasks
Secrets and encryptionAWS Secrets Manager and AWS KMSKey Vault and managed HSMSecret Manager and Cloud KMS
Private connectivityVPC endpoints and AWS PrivateLinkPrivate Link and private endpointsPrivate Service Connect and VPC Service Controls where applicable
Operational monitoringCloudWatch, X-Ray and OpenTelemetryAzure Monitor, Application Insights and OpenTelemetryCloud Monitoring, Cloud Logging, Cloud Trace and OpenTelemetry
Governance and inventoryAWS Config, Security Hub and model or platform registriesAzure Policy, Microsoft Purview and Foundry governanceCloud Asset Inventory, Security Command Center and Vertex AI governance

Current vendor guidance supports this capability-oriented approach. AWS documents both a production generative-AI architecture and an Agentic AI Lens. Microsoft describes the layered Microsoft Foundry architecture and API Management AI gateway. Google publishes an agentic architecture component-selection guide and a RAG reference architecture.

Watch for platform lifecycle changes​

Cloud products and agent runtimes evolve rapidly. AWS documents that Bedrock Agents Classic is in maintenance mode, while Google documents that Vertex AI Extensions is deprecated and directs customers toward its Gemini Enterprise Agent Platform. New designs should validate product status, contractual commitments, regional availability, identity support, exportability and migration paths before committing.

Example cloud-selection criteria​

Weight the decision using:

  • enterprise identity and existing cloud operating model;
  • required EU processing regions and actual model availability;
  • private connectivity and egress control;
  • model and embedding selection;
  • document-level permissions and retrieval quality;
  • agent identity, tool policy and human approval capabilities;
  • processor, sub-processor and retention commitments;
  • observability export and evidence availability;
  • internal skills, supplier concentration and exit feasibility;
  • three-year total cost of ownership.

Do not select a provider solely because a model demo looks stronger. Architecture quality depends on the complete control and operating model.


21. Data contracts that make governance enforceable​

Every important boundary should exchange a typed contract. The examples below are illustrative and intentionally exclude raw sensitive content from general audit events.

AI request envelope​

{
"request_id": "req_01JEXAMPLE",
"trace_id": "trc_01JEXAMPLE",
"tenant_id": "tenant_eu_123",
"principal": {
"subject_id": "user_458",
"roles": ["support_adviser"],
"assurance_level": "mfa",
"region": "eu-west",
"allowed_purposes": ["customer_support"]
},
"use_case": "banking_customer_support",
"data_classification": "confidential_personal",
"policy_version": "policy-2026-08-21.4",
"workflow_version": "support-rag-v12",
"limits": {
"maximum_input_tokens": 2500,
"maximum_output_tokens": 600,
"maximum_tool_calls": 3,
"maximum_duration_ms": 10000,
"maximum_cost_units": 20
},
"approved_model_region": "eu",
"human_approval_required_for": ["refund", "address_change"]
}

The server derives the identity, tenant and assurance fields from authenticated context. They are not trusted merely because the browser included them in a JSON body.

RAG document metadata contract​

{
"tenant_id": "tenant_eu_123",
"document_id": "doc_fee_policy_2026_08",
"document_version": "8",
"chunk_id": "chunk_015",
"source_uri": "controlled-document-reference",
"owner_id": "bank_policy_team",
"classification": "internal",
"permitted_roles": ["support_adviser"],
"permitted_regions": ["eu"],
"purpose_tags": ["customer_support"],
"valid_from": "2026-08-01T00:00:00Z",
"valid_until": "2026-12-31T23:59:59Z",
"retention_policy": "policy_document_standard",
"content_hash": "sha256:example",
"embedding_model_version": "approved-embedding-v3",
"parser_version": "document-parser-v6"
}

Enforcement should occur before retrieval and again before a passage enters model context. Version dates prevent an expired fee schedule from being presented as current policy.

Tool execution contract​

{
"tool_name": "create_refund_review",
"action_type": "write",
"tenant_id": "tenant_eu_123",
"resource_id": "case_9981",
"requested_by": "user_458",
"business_purpose": "customer_support",
"risk_tier": "material_financial_action",
"idempotency_key": "idem_01JEXAMPLE",
"arguments": {
"transaction_reference": "txn_772",
"requested_amount_minor_units": 1200,
"currency": "GBP"
},
"required_controls": {
"step_up_authentication": true,
"human_approval": true,
"separation_of_duties": true
},
"approval_reference": "approval_227"
}

The target business service must still validate amount, ownership, currency, approval and state transition independently. The tool gateway reduces risk; it does not replace the destination system's own security.

Privacy-minimised audit event​

{
"event_id": "evt_01JEXAMPLE",
"occurred_at": "2026-08-21T12:30:00Z",
"tenant_id": "tenant_eu_123",
"actor_reference": "pseudonymous-user-reference",
"event_type": "tool.execution.denied",
"use_case": "banking_customer_support",
"policy_version": "policy-2026-08-21.4",
"model_version": "approved-chat-model-v18",
"workflow_version": "support-rag-v12",
"tool_name": "execute_refund",
"decision": "deny",
"reason_code": "human_approval_missing",
"trace_id": "trc_01JEXAMPLE",
"integrity_reference": "signed-audit-event-reference"
}

The event provides accountability without recording the customer's complete prompt, bank statement or generated answer.


22. Policy enforcement outside the model​

A production policy decision should combine identity, data, tool, model and transaction attributes. A readable example:

policy_id: banking-support-refund-review
applies_to:
use_case: banking_customer_support
tool: create_refund_review

allow_when:
- principal.tenant_id == resource.tenant_id
- principal.role in [support_adviser, support_supervisor]
- principal.assurance_level in [mfa, phishing_resistant_mfa]
- request.business_purpose == customer_support
- resource.region == approved_processing_region
- request.amount_minor_units <= principal.approval_limit
- request.idempotency_key is present

require_human_approval_when:
- request.amount_minor_units > 10000
- customer.is_vulnerable == true
- tool.action_type == write

deny_when:
- resource.owner_id != verified_customer_id
- model.region not in approved_regions
- provider.data_retention_policy != approved_retention_policy
- tool.destination not in approved_enterprise_endpoints

Treat the snippet as illustrative policy logic, not a production-ready policy language. A real implementation must define evaluation order, fail-closed behaviour, missing-attribute semantics, emergency overrides and test coverage.

The four checks every tool action needs​

  1. Identity: Is the requester authenticated at the required assurance level?
  2. Authority: Is this principal allowed to perform this action on this resource?
  3. Context: Is the purpose, tenant, data class, region and model route permitted?
  4. Consequence: Does the action require confirmation, approval, amount limits, separation of duties or compensation?

If any answer is missing, deny the action. Asking the model to “be careful” is not equivalent.


23. Multi-region deployment and data residency​

For regulated workloads, availability and residency must be designed together. A global user experience does not automatically justify copying customer prompts, source documents or traces to every cloud region.

  • assign each tenant a documented home region;
  • keep prompts, conversation data, source documents, embeddings, traces and backups within the approved regional boundary;
  • replicate policy definitions and deployment artefacts globally only when they do not contain restricted customer content;
  • use a second approved region for recovery when permitted by the customer contract and transfer assessment;
  • route based on verified tenant metadata, not browser location;
  • prevent an external model fallback from escaping the tenant's approved region;
  • document support-person access and remote administrative access, which may themselves create data-transfer concerns.

Active-active versus active-passive​

Active-active improves resilience and latency but complicates consistency, duplication, conflict handling, deletion, region-specific contracts and residency. Active-passive is simpler for many regulated workloads but introduces recovery delay.

PatternAppropriate whenMain trade-off
Single region, multiple zonesthe business tolerates regional interruption and requires simple residencyregional outage interrupts service
Active-passive within approved regionslow data-loss and predictable recovery matter more than zero interruptionfailover and recovery must be tested
Active-active within one legal boundaryhigh availability and local latency justify operational complexityreplication, consistency and deletion are harder
Dedicated regional tenant cellsresidency, isolation or large-customer needs dominatemore infrastructure, deployment and support overhead

Recovery objectives are meaningful only after restoration tests. A backup policy that has never been restored is a claim, not evidence.


24. Capacity planning and performance engineering​

LLM architecture requires sizing by requests, concurrency, tokens, retrieval operations, tool calls and provider quota. Request-per-second figures alone hide the main bottlenecks.

Worked example​

Assume a peak of 50 customer requests per second and an average complete-request service time of 5 seconds. By Little's Law:

average concurrent requests = arrival rate × average service time

Therefore:

50 requests/second × 5 seconds = 250 average in-flight requests

If each request averages 1,800 input tokens and 350 output tokens:

  • input throughput is approximately 90,000 tokens per second;
  • output throughput is approximately 17,500 tokens per second;
  • input demand is approximately 5.4 million tokens per minute;
  • output demand is approximately 1.05 million tokens per minute;
  • an average of two retrieval queries per request creates about 100 retrieval queries per second;
  • 20 reranked passages per request create roughly 1,000 reranked passages per second;
  • a 30% tool-use rate creates approximately 15 tool calls per second.

These are averages, not provisioning guarantees. Include burst factors, p95 and p99 service times, provider throttling, context growth, regional failure and tenant concentration. Model quotas may be stated in requests per minute, input tokens per minute, total tokens per minute or other provider-specific units, so verify the exact quota semantics.

Performance budget​

Allocate an explicit end-to-end latency budget:

StageIllustrative p95 budget
Gateway, authentication and policy150 ms
Query rewriting and metadata filtering150 ms
Search and reranking700 ms
Tool lookup, where applicable900 ms
Model time to first token1,500 ms
Response generation and validation3,500 ms
Network and contingency1,100 ms
Total indicative budget8,000 ms

Do not add independent p95 measurements and claim the sum is the actual end-to-end p95. Use distributed traces and realistic load tests; the table is a design budget, not a statistical identity.

Cost formula​

Calculate unit economics using:

cost per completed task = model input + model output + embeddings + retrieval + tool execution + compute + storage + observability + human review + support

Then calculate:

net business value per task = avoided manual cost + incremental revenue + risk reduction - total task cost

Optimisation order should be:

  1. remove unnecessary or unlawful data;
  2. improve retrieval precision;
  3. reduce context and output budgets;
  4. route easy tasks to approved smaller models;
  5. cache only under entitlement-safe keys;
  6. batch asynchronous work;
  7. reserve or optimise capacity when usage is predictable;
  8. preserve human review and safety controls for high-impact cases.

Never lower cost by moving data to an unapproved provider or removing a control required by the risk assessment.


25. Choosing a tenant-isolation model​

Multi-tenancy is an architecture spectrum rather than a binary choice.

Isolation modelData boundaryTypical fitRisks and obligations
Shared tables with row-level securitylogical tenant filters in one databaseearly-stage or lower-risk SaaSdemands rigorous policy, index, cache and test discipline
Separate schema or database per tenantdatabase-level boundarycustomers with stronger isolation or recovery requirementsincreases migration, backup and connection-management complexity
Dedicated search index per tenantretrieval-level boundaryconfidential document corpora and strict access requirementsmore indexes, ingestion workflows and operational cost
Regional cell per tenant groupseparate runtime and data plane per jurisdiction or segmentdata residency and blast-radius reductionrequires cell-aware deployment and routing
Dedicated customer environmentisolated cloud account/subscription/project and data storesregulated strategic customers with contractual separationhighest operating and provisioning overhead

An effective design often combines models. For example, ordinary customers can share an application cell, while high-assurance customers receive a dedicated regional cell and customer-managed key.

Mandatory isolation test cases​

Include:

  • modified tenant ID in an API body;
  • valid user token against another tenant's document ID;
  • cross-tenant vector nearest-neighbour query;
  • cache hit after changing role or document ACL;
  • queued message replay in the wrong tenant;
  • agent memory referencing another tenant;
  • model fallback that violates residency;
  • administrator search without approved support access;
  • source deletion followed by immediate retrieval;
  • report export with mixed tenant data.

These are release-blocking security tests, not optional quality checks.


26. Compliance traceability down to individual controls​

A control catalogue connects architecture, law, evidence and accountability. The following identifiers are illustrative internal controls; they are not official ISO or legal clause numbers.

Internal controlEngineering implementationRegulatory or standard relationshipEvidence
AI-GOV-01 intended-purpose registerapproved system inventory, excluded uses and change triggersAI Act classification; ISO/IEC 42001 scope and operational governanceapproved use-case record and review history
AI-GOV-02 risk and impact reviewrisk register, DPIA, FRIA where applicable and impact assessmentGDPR Articles 25/35; AI Act Articles 9/27 where applicable; ISO/IEC 42005assessments, treatment actions and residual-risk acceptance
AI-ID-01 verified tenant identityfederated login, short-lived claims and workload identityGDPR Article 32; ISO/IEC 27001 access controlidentity configuration and access tests
AI-ID-02 attribute-based document accesspre-search filtering, row-level security and ACL recheckdata minimisation and confidentiality; high-risk logging and governance where applicablepolicy tests and tenant-isolation reports
AI-DATA-01 source provenancesource owner, rights, hash, validity, sensitivity and document ACLAI Act Article 10 where applicable; ISO/IEC 42001 lifecycle/data controlsdataset and ingestion records
AI-DATA-02 retention and erasureTTLs, deletion propagation and processor deletion workflowGDPR Articles 5 and 17deletion test and retention configuration
AI-MODEL-01 approved provider catalogueallowed model, region, contract, retention and fallback policyGDPR Articles 28 and 44–49; supplier controlsvendor assessment, DPA and transfer record
AI-MODEL-02 versioned release bundlepinned models, prompts, policies, indexes and evaluationsAI Act technical documentation and QMS where applicable; ISO change managementsigned release manifest and approvals
AI-TOOL-01 least-privilege action gatewaytyped tools, scoped credentials and deterministic validationGDPR Article 32; ISO/IEC 27001 secure development and access controltool schema, workload permissions and test results
AI-TOOL-02 meaningful approvalevidence-based human review, override and stop capabilityGDPR Article 22 where applicable; AI Act Article 14 for high-risk systemsapproval UI, reviewer training and override log
AI-EVAL-01 multidimensional evaluationquality, bias, privacy, security and resilience release gatesAI Act Articles 9 and 15 where applicable; ISO/IEC 42001 performance evaluationversioned evaluation report
AI-OBS-01 privacy-minimised traceabilityintegrity-protected event history and limited operational telemetryGDPR accountability/security; AI Act Article 12 for high-risk systemsaudit schema, access logs and retention settings
AI-OPS-01 incident responsekill switch, escalation, breach analysis and corrective actionGDPR Articles 33/34; AI Act incident duties where applicablerunbooks, exercises and incident records
AI-HR-01 AI literacyrole-specific training for developers, operators and oversight staffAI Act Article 4; ISO/IEC 42001 support and competencecurriculum, attendance and competence records

The same implementation can satisfy multiple requirements, but the evidence must show why it is adequate for each obligation. A generic security certificate is not proof that a particular retrieval index preserves document permissions.


27. Operational runbooks for common AI failures​

Runbook A: model quality regression​

Trigger: Groundedness, citation correctness or critical task success drops below its alert threshold.

  1. Confirm whether the change correlates with model, prompt, policy, index or source version.
  2. Compare affected tenants, languages, document types and risk categories.
  3. Block the suspect version from further rollout.
  4. Route to an already approved prior model or workflow version.
  5. Re-evaluate against production-derived and expert-labelled cases.
  6. Correct the root cause and require a fresh release decision.

Runbook B: poisoned or compromised knowledge source​

Trigger: An ingestion integrity alert, suspicious source instructions, false policy answers or anomalous document changes.

  1. Pause the connector and quarantine newly ingested content.
  2. Disable retrieval from the affected source or index partition.
  3. Restore the prior approved index snapshot.
  4. Identify all answers, users and tools influenced by the malicious document.
  5. Assess whether private data was exposed or unauthorised actions occurred.
  6. Notify source owners, security, privacy and affected business teams as required.
  7. Add a regression case and strengthen source approval or provenance checks.

Runbook C: unauthorised agent action​

Trigger: A tool call executes beyond role, amount, purpose, tenant or approval policy.

  1. Disable the affected tool globally or for the tenant.
  2. Revoke the tool workload identity and rotate relevant credentials.
  3. Determine whether the target business service independently enforced authorisation.
  4. Reconcile every transaction and execute compensating actions when possible.
  5. Preserve model, tool, prompt, policy, approval and trace versions.
  6. Assess legal, financial and notification consequences.
  7. Fix the deterministic control before re-enabling the tool.

Runbook D: model-provider outage​

Trigger: Provider errors, throttling, degraded latency or an unavailable approved region.

  1. Open the model-specific circuit breaker.
  2. Check whether a pre-approved fallback matches region, data class and contract.
  3. If approved, shift a small canary cohort and monitor quality.
  4. If not approved, degrade to search-only, static support or human escalation.
  5. Apply backpressure and tenant communication.
  6. Review recovery, cost and downstream queue behaviour before returning traffic.

Runbook E: sudden token or infrastructure spend​

Trigger: Cost per task, token volume, context size or recursive tool use spikes.

  1. Enforce per-user, tenant and model budgets.
  2. Identify oversized documents, recursive loops, abuse or fallback changes.
  3. Restrict maximum context, generation and tool steps within approved quality limits.
  4. Suspend suspicious traffic and alert account owners.
  5. Re-run quality and privacy checks before changing routing or caching.

28. Architecture decisions and implementation handoff​

An enterprise solution architect should leave behind decisions that another team can understand, challenge and operate.

Minimum architecture decision record​

Each architecture decision record should include:

  • decision ID and date;
  • business context and required outcome;
  • options considered;
  • chosen option and reasoning;
  • security, privacy, compliance, cost and availability impacts;
  • rejected alternatives;
  • owner and approvers;
  • assumptions and review triggers;
  • rollback or exit route.

Example ADR: vector storage​

Decision: Start with PostgreSQL and vector search for the first production release.

Why: The corpus is within the benchmarked capacity envelope, document permissions join naturally with relational ownership, and the team can reuse existing backup, encryption and operational controls.

Trade-off: Search and indexing may need a dedicated engine if corpus size, filtered-query latency or multi-tenant concurrency grows.

Review triggers: p95 retrieval above the approved budget, index rebuild above the recovery target, sustained CPU saturation, loss of retrieval quality or customer demand for dedicated index isolation.

Example ADR: provider fallback​

Decision: Disable automatic cross-provider failover for confidential personal data.

Why: The secondary provider has not yet completed the required region, processor, sub-processor and retention approvals.

Trade-off: During the primary provider outage, the product degrades to source search and human escalation.

Review trigger: Completion of equivalent privacy, security, evaluation and contractual assurance for the secondary route.

Handoff package​

A practical delivery repository can be organised around:

architecture/
01-business-context-and-intended-purpose.md
02-system-context-and-trust-boundaries.md
03-component-and-deployment-architecture.md
04-data-flow-and-processing-record.md
05-rag-ingestion-and-retrieval-design.md
06-agent-tools-and-approval-model.md
07-identity-tenant-and-policy-design.md
08-eu-ai-act-gdpr-and-iso-control-crosswalk.md
09-threat-model-and-red-team-plan.md
10-evaluation-strategy-and-release-gates.md
11-observability-incident-and-recovery-runbooks.md
12-capacity-cost-and-commercial-model.md
13-supplier-and-model-due-diligence.md
14-architecture-decision-records.md
15-implementation-backlog-and-acceptance-tests.md

The implementation backlog should translate each control into an owned, testable story. For example:

As a security owner, I need every vector query to include verified tenant and document-entitlement filters before retrieval, so that an authenticated user cannot infer or retrieve another tenant's content.

Acceptance criteria:

  • query execution fails closed when tenant context is missing;
  • the tenant is derived from the authenticated principal;
  • search filters are applied before similarity retrieval;
  • document ACLs are checked again before prompt assembly;
  • cross-tenant tests fail if any content, metadata or existence signal leaks;
  • audit events contain the policy decision but not the document content.

This turns governance from a slide-deck claim into a feature that can be implemented, tested and operated.


Conclusion​

The best AI architecture does not try to make a probabilistic model behave like a perfectly trusted database or policy engine. It gives each component the responsibility it can safely carry.

  • Models generate, classify and propose.
  • Retrieval supplies authorised evidence.
  • Deterministic services enforce identity, policy, limits and transactions.
  • Humans provide judgement, accountability and redress at the right risk points.
  • Management systems make ownership, evidence, audit and improvement repeatable.

EU AI Act, GDPR, ISO/IEC 42001 and ISO/IEC 27001 requirements should therefore not be separate workstreams added after the technical design. They become architecture inputs: tenant isolation, data minimisation, provenance, versioning, meaningful oversight, traceability, supplier controls, evaluation gates and safe failure.

That is the difference between an impressive demonstration and an enterprise AI capability that can be trusted, operated and defended in the real world.


Primary sources and implementation references​

Discussion

Comments​

Share feedback or questions about this page. No account required.

Loading comments…