Production-Grade AI Architecture: A Practical Blueprint for Secure, Scalable and Governed Enterprise AI
Designing an end-to-end AI platform for the EU AI Act, ISO/IEC 42001, ISO/IEC 27001, GDPR and modern GenAI security
Updated: 21 August 2026
Enterprise AI architecture is not the art of connecting an application to a large language model. A production system must remain useful when the model is uncertain, safe when retrieved content is malicious, compliant when personal data is involved, available when a provider fails, and auditable months after an output or action was produced.
That changes the architectural question from:
Which model should we use?
to:
How do we build a controlled socio-technical system in which models, data, people, policies and software work together within defined limits?
This article answers that question with a vendor-neutral reference architecture for a multi-tenant enterprise AI assistant using retrieval-augmented generation, or RAG, and governed agent tools. It covers the complete lifecycle from use-case discovery and regulatory classification through ingestion, orchestration, evaluation, deployment, monitoring, incident response and retirement.
The legal discussion is an engineering interpretation, not legal advice. The organisation's legal counsel, Data Protection Officer, information-security team and relevant sector specialists should confirm the final obligations for each use case and jurisdiction.
1. What “production-grade” really means
A successful proof of concept proves that an AI capability can work under favourable conditions. A production architecture must prove something harder: that the overall system behaves acceptably under normal use, edge cases, attack, component failure, policy change and model change.
A production-grade AI solution should have all of the following properties:
- Useful: It improves a measurable business or user outcome, not merely an AI metric.
- Grounded: It can distinguish source-backed statements from generated synthesis and uncertainty.
- Secure: Every user, service, data access and tool action is authenticated, authorised and observable.
- Private: Personal data is collected, used, retained and deleted for explicit purposes under a valid legal basis.
- Governed: Owners, intended purposes, prohibited uses, risk decisions and approval gates are documented.
- Reliable: Timeouts, retries, queues, circuit breakers, fallbacks, backups and recovery targets are designed before failure occurs.
- Evaluated: Quality, safety, security, fairness, privacy, performance and cost are tested against versioned release criteria.
- Auditable: The organisation can reconstruct which user, policy, data, prompt, model and tool chain produced a material outcome.
- Controllable: Humans and operators can review, override, stop, roll back or disable risky capabilities.
- Economically sustainable: Capacity and token costs are bounded, allocated and monitored without weakening safety or privacy.
The foundation is therefore not an LLM. It is a controlled platform around one or more models.
2. Start with the intended purpose, not the technology
Architecture begins with a one-page use-case contract. It should state:
- the user and affected-person groups;
- the problem and measurable outcome;
- the system's intended purpose;
- decisions the system may support;
- decisions or actions it must never make;
- data categories and sources;
- expected geography and jurisdictions;
- model, retrieval and tool capabilities;
- human-oversight points;
- foreseeable misuse;
- risk appetite, service targets and exit criteria.
The intended purpose matters technically and legally. A policy-search assistant and a candidate-ranking system can use the same foundation model but present radically different risks. Changing an assistant from “draft a recommendation” to “automatically reject an applicant” is not a minor feature toggle. It changes the system's decision authority, impact, regulatory classification, human-oversight design and evidence requirements.
A reusable reference scenario
The architecture in this article assumes an enterprise AI service with these capabilities:
- web and mobile chat;
- business-to-business multi-tenancy;
- organisation, user, role and attribute-based access control;
- RAG over public, enterprise and user-authorised documents;
- access to more than one approved model provider;
- read-only tools by default, with tightly controlled write actions;
- EU data residency for in-scope data;
- human approval for material or irreversible actions;
- full prompt, model, retrieval, policy and tool versioning;
- privacy-aware observability and audit evidence.
Example, non-universal service objectives might be:
| Dimension | Illustrative target |
|---|---|
| Availability | 99.95% monthly for the user-facing service |
| Time to first token | p95 below 2.5 seconds for ordinary queries |
| End-to-end response | p95 below 8 seconds for ordinary RAG queries |
| Recovery point objective | 5 minutes for operational data |
| Recovery time objective | 30 minutes for a regional service failure |
| Citation correctness | At least 95% on the approved evaluation set |
| Grounded answer rate | At least 90% for answerable knowledge questions |
| Cross-tenant leakage | Zero accepted failures in isolation testing |
| Unauthorised tool action | Zero accepted failures |
| Harmful critical output | Zero accepted failures in defined critical categories |
| Cost | Per-task budget by use case, model tier and tenant |
These are examples, not universal benchmarks. Targets must reflect harm, task complexity, user expectations, regulatory requirements and the organisation's risk appetite.
3. Understand how the obligations fit together
The EU AI Act, GDPR and ISO standards do different jobs. Treating them as interchangeable creates gaps.
| Instrument | Primary purpose | What the architecture team should do |
|---|---|---|
| EU AI Act | Risk-based legal rules for AI systems and general-purpose AI models in the EU | Classify the system and operator role; implement applicable risk, transparency, documentation, oversight, robustness and post-market controls |
| GDPR | Legal rules for processing personal data | Establish purposes and lawful bases; minimise data; enable rights; secure processing; manage processors, transfers, retention, DPIAs and breaches |
| ISO/IEC 42001:2023 | Requirements for an AI management system | Establish repeatable AI policy, accountability, risk, impact, lifecycle, supplier, monitoring, audit and continual-improvement processes |
| ISO/IEC 27001:2022 | Requirements for an information-security management system | Manage confidentiality, integrity and availability through risk treatment and controlled information-security operations |
| ISO/IEC 27701:2025 | Requirements and guidance for a privacy information management system | Operationalise controller and processor privacy responsibilities and integrate them into management processes |
| NIST AI RMF | Voluntary AI risk-management framework | Use Govern, Map, Measure and Manage to structure practical risk work and evidence |
| OWASP guidance | Application, LLM, agent and software-supply-chain security | Threat-model and verify implementation-level controls against current attack patterns |
ISO/IEC 42001 specifies requirements for establishing, implementing, maintaining and continually improving an AI management system. ISO/IEC 27001 does the equivalent for an information-security management system. Certification to either standard can demonstrate a managed process, but it does not by itself prove that a particular AI system complies with the AI Act, GDPR or sector law.
Useful supporting standards include ISO/IEC 23894:2023 for AI risk management, ISO/IEC 42005:2025 for AI system impact assessment, ISO/IEC 27701:2025 for privacy management and ISO/IEC 27018:2025 for protection of PII in public-cloud processing.
The current EU AI Act timeline
As of 21 August 2026, the AI Act's general application date has passed. Most prohibited-practice provisions and the AI-literacy duty began applying on 2 February 2025, governance and general-purpose AI provisions began applying on 2 August 2025, and the Commission and national authorities began broader enforcement from 2 August 2026. Article 50 transparency duties now apply to specified interactive and generative systems. The 2026 amendment includes transitional details: newly added prohibitions concerning specified non-consensual intimate material and child sexual abuse material apply from 2 December 2026, and providers of relevant content-generating systems placed on the market before 2 August 2026 have until 2 December 2026 for the Article 50(2) marking duty. Following Regulation (EU) 2026/1744, the high-risk rules for Annex III standalone systems apply from 2 December 2027, while rules for high-risk systems embedded in Annex I regulated products apply from 2 August 2028. The authoritative sources are the consolidated AI Act, the 2026 amending regulation and the Commission's implementation overview.
The later high-risk dates are preparation time, not architecture exemptions. Systems being designed now may still be operating when those rules apply. Controls that require traceability, data lineage, human oversight or representative evaluation are expensive to retrofit.
Classify the system before building it
For every use case, record:
- whether it is an AI system within the Act's scope;
- whether the organisation is a provider, deployer, importer, distributor, authorised representative, product manufacturer, GPAI provider or downstream provider;
- whether a prohibited practice is involved;
- whether the intended purpose falls into Annex I or Annex III high-risk categories;
- whether Article 50 transparency duties apply;
- whether a third-party model introduces GPAI value-chain duties;
- whether a change in purpose or substantial modification could shift the organisation into provider obligations;
- which national, employment, consumer, financial, health, accessibility or product laws also apply.
Do not classify only the model. Classification belongs to the system and its intended purpose. The same model may support a low-impact writing tool, a customer-facing chatbot and a high-risk employment system.
4. The end-to-end reference architecture
The safest pattern separates four concerns:
- experience plane: channels and user experience;
- AI runtime plane: orchestration, retrieval, tools, models and deterministic guardrails;
- data plane: operational, source, vector, cache and evidence stores;
- control plane: identity, policy, governance, evaluation, deployment, observability and incident controls.
Trust boundaries
Document explicit trust boundaries in the threat model:
| Boundary | Examples | Default assumption |
|---|---|---|
| User device | Browser, mobile app, collaboration client | Potentially compromised; all input is untrusted |
| Edge | CDN, WAF, gateway, rate limiter | Internet exposed; withstand abuse and malformed traffic |
| Application | APIs, identity context, orchestrator | Authenticated does not mean authorised for every resource |
| AI runtime | prompts, retrieved content, model outputs, agent plans | Probabilistic and potentially attacker-influenced |
| Data | relational, object, vector, cache, logs | Sensitive; every access must preserve tenant and entitlement boundaries |
| Tool and enterprise systems | CRM, HR, payments, email, ticketing | High impact; require explicit scopes and deterministic validation |
| Third-party provider | model, embedding, moderation, observability service | External processor or supplier; contractual and technical boundaries apply |
The essential rule is simple: natural-language instructions are never a security boundary. A system prompt telling the model not to reveal secrets cannot replace access control, network policy, schema validation or cryptographic isolation.
5. End-to-end request flow
For a typical RAG and agent request, the controlled path is:
- Establish identity. The user authenticates through enterprise SSO with phishing-resistant MFA where appropriate. The gateway creates a short-lived identity context containing organisation, user, session, role, attributes, assurance level and allowed purposes.
- Protect the edge. The WAF and API gateway enforce request size, content type, rate, bot, origin and abuse controls. A quota service applies user, tenant, IP, model and tool budgets.
- Create a trace context. The service creates a correlation ID, but avoids placing raw personal data in identifiers, metrics or log labels.
- Validate and classify input. Deterministic checks reject malformed content, prohibited file types and oversized input. DLP and policy services identify secrets, personal data, special-category data and potential injection patterns. Detection informs controls but should not be treated as perfect.
- Resolve policy. A policy engine evaluates the user, tenant, purpose, data class, tool, model, region, risk tier and consent or lawful-basis metadata. The result is a machine-enforced execution policy, not prose embedded in a prompt.
- Retrieve authorised evidence. Retrieval applies tenant and document entitlements before search, performs hybrid keyword and vector search, reranks candidates, validates access again and returns small source passages with provenance.
- Construct context. The context builder separates system policy, developer instructions, user input, retrieved untrusted content and tool results. It includes only the minimum relevant information.
- Select a model. The model gateway selects from an approved catalogue using task, data class, geography, latency, quality, context and cost policies. It must not silently route protected data to an unapproved fallback.
- Generate a plan or response. If tools are needed, the model proposes structured actions. The model never holds standing credentials.
- Authorise each tool call. A tool gateway validates the schema, user authority, purpose, resource scope, state transition, idempotency key, risk level and approval requirement. High-impact actions require step-up authentication, user confirmation or human approval.
- Validate output. The system verifies structure, citations, sensitive-data policy, allowed claims and action results. Deterministic business rules take priority over model instructions.
- Return with transparency. The interface identifies AI interaction where required, distinguishes sources from generated text, communicates important limitations, exposes correction or escalation paths and records user feedback.
- Record evidence. Privacy-minimised telemetry records versions, decisions, document IDs, policy outcomes, tool calls, safety signals, latency and cost. Raw prompts and responses are not logged by default.
6. Layer-by-layer architecture decisions
6.1 Experience and edge layer
The front end should make safe behaviour easy:
- identify the experience as AI-powered where Article 50 or other transparency rules apply;
- show citations adjacent to supported claims;
- distinguish a draft, recommendation and completed action;
- display uncertainty and refusal in usable language;
- require confirmation before consequential actions;
- offer “report”, “correct”, “delete”, “appeal” and “talk to a person” routes where relevant;
- meet applicable accessibility requirements;
- avoid dark patterns that pressure users into sharing additional data.
At the edge, use TLS, modern security headers, strict CORS, CSRF protection for browser sessions, WAF rules, bot controls, upload limits, malware scanning and layered rate limiting. Rate limits should operate at several dimensions: network source, user, tenant, endpoint, token volume, concurrent generation, tool calls and daily spend.
6.2 Identity, tenant and policy layer
Use the organisation's identity provider rather than creating a second identity island. Prefer OIDC or SAML federation, short-lived tokens and centrally managed joiner, mover and leaver processes.
Authentication answers “who are you?” Authorisation must answer:
- which tenant are you acting in?
- which data may you retrieve?
- for which purpose?
- which tool and action may you invoke?
- on which resource?
- under what approval threshold?
- from which managed device and assurance level?
RBAC is useful for broad job roles. ABAC or relationship-based access control is normally needed for document sensitivity, region, case assignment, legal hold, ownership and purpose. Policy engines such as Open Policy Agent, Cedar-style policy services or an equivalent managed service can make decisions consistent across API, retrieval and tool layers.
For multi-tenant systems:
- derive
tenant_idfrom verified identity, never from an untrusted request body alone; - include tenant context in every primary key, index, cache key, queue message and audit event;
- enforce database row-level security or separate databases for higher-assurance tenants;
- use tenant-scoped encryption keys where warranted;
- run automated cross-tenant isolation tests on every release;
- reject rather than guess when tenant context is absent.
6.3 AI orchestrator
The orchestrator is a state machine, not an unconstrained chat loop. It should own:
- maximum steps, execution time, token use and cost;
- allowed transitions and terminal states;
- retry and timeout policy;
- prompt and workflow versions;
- policy results;
- model and tool selection;
- checkpointing and idempotency;
- human approval and cancellation;
- trace correlation.
For complex agents, separate planning from execution. A model may propose a typed plan, but a deterministic controller validates it before executing one bounded step at a time. Do not allow recursive delegation without explicit depth and budget limits.
6.4 Model gateway
A model gateway prevents each team from implementing provider access differently. It should provide:
- an approved model and embedding catalogue;
- provider, region, version and data-class policies;
- model-specific adapters and structured-output validation;
- prompt-template and policy version headers;
- quota, timeout, retry and circuit-breaker controls;
- request and response size limits;
- token and cost accounting;
- content-safety integration;
- privacy-preserving telemetry;
- controlled fallback.
Model version aliases should not float silently in production. Pin a tested version when the provider permits it, detect provider-side changes, rerun the evaluation suite and roll out using a canary.
A fallback is a new processing path. Before enabling it, confirm equivalent data residency, DPA terms, retention settings, model-training exclusions, safety controls and task quality. When no compliant fallback exists, degrade gracefully instead of sending data somewhere unapproved.
6.5 RAG ingestion and retrieval
RAG improves access to current enterprise knowledge, but creates a new supply chain for untrusted content.
Controlled ingestion pipeline
- Register the data owner, source, purpose, lawful basis, region, sensitivity and retention class.
- Authenticate the connector with a dedicated least-privilege service identity.
- Copy through a quarantine boundary.
- Scan for malware, active content and unsupported formats.
- classify and DLP-scan the content;
- extract text and structure using sandboxed parsers;
- preserve document ACLs and security labels;
- normalise, chunk and attach metadata;
- create embeddings through an approved regional endpoint;
- store chunks and vectors with tenant and entitlement metadata;
- run quality, poisoning and access-control tests;
- publish the index atomically only after approval.
Every chunk should carry at least:
- tenant and source identifiers;
- document and version identifiers;
- source URI or immutable reference;
- owner and classification;
- access-control labels;
- valid-from and valid-to dates;
- retention and legal-hold tags;
- content hash and ingestion timestamp;
- parser, chunker and embedding versions;
- language and provenance.
Secure retrieval pipeline
Apply access filters before similarity search wherever the data store supports it. Never retrieve across all tenants and remove unauthorised results afterward; the retrieval itself is already a disclosure surface. Use hybrid search, metadata filtering and reranking, then re-check entitlement before context assembly.
Treat retrieved text as quoted evidence, not instructions. Delimit it, label its provenance, strip or neutralise active content and instruct the runtime that documents cannot change system policy or request tools. Because prompt injection cannot be solved by prompting alone, combine this with tool isolation, least privilege, egress restrictions and deterministic approval.
6.6 Governed tool gateway
Agent tools turn an information risk into an action risk. Each tool should expose a narrow, typed capability such as get_invoice_status rather than a generic “run SQL” or “call arbitrary URL” function.
Enforce:
- allowlisted tools per use case and role;
- narrowly scoped, short-lived credentials;
- typed JSON schemas and semantic validation;
- resource-level authorisation;
- private network paths and egress allowlists;
- command, URL, SQL and template injection protection;
- read-only mode by default;
- transaction limits and velocity rules;
- preview or dry-run for write actions;
- step-up authentication or human approval for sensitive actions;
- idempotency keys and replay protection;
- before-and-after state capture;
- compensating action or rollback where possible;
- an immediate kill switch.
The tool gateway should independently calculate important values. A model should not be trusted to determine a refund amount, beneficiary, eligibility result or access scope from free text.
6.7 Data stores
Use different stores for different guarantees:
| Store | Purpose | Important controls |
|---|---|---|
| Relational database | tenants, users, permissions, cases, conversation metadata, approvals | row-level security, encryption, constraints, audit, point-in-time recovery |
| Object store | source documents, evaluation sets, model artefacts, exports | versioning, malware quarantine, object lock where needed, lifecycle policy, access labels |
| Vector index | authorised retrieval | pre-filtered tenant and ACL isolation, deletion propagation, versioned indexes, encryption |
| Cache | sessions, rate limits, safe response fragments | short TTL, tenant-bound keys, no unapproved sensitive content, encryption where supported |
| Audit store | security and material decision evidence | append-only access, integrity protection, strict reader roles, separate retention policy |
| Telemetry platform | metrics, traces and operational events | redaction, sampling, label controls, regional storage, limited raw-content capture |
Encryption at rest and in transit is necessary but insufficient. Design key ownership, rotation, separation of duties, backup encryption, revocation and emergency access. Keep secrets in a managed secret store; never place credentials in prompts, source code, container images or ordinary environment dumps.
6.8 Observability without surveillance
Operational telemetry should answer:
- which version was running?
- which policy decision applied?
- which documents were retrieved?
- which model and provider were called?
- which tools were proposed, approved and executed?
- how long did each stage take?
- what did it cost?
- did quality, safety or security signals change?
Useful event fields include a pseudonymous actor reference, tenant, use case, trace ID, policy version, workflow version, prompt-template version, model version, retrieval index version, source document IDs, tool name, approval outcome, safety decision, latency, token counts and cost.
Do not put raw prompts, responses, access tokens, emails or document passages into metric labels. Default traces to metadata-only. If content capture is temporarily needed for debugging or evaluation, require an approved purpose, sampled scope, restricted access, encryption, short retention and automatic expiry.
Use OpenTelemetry or an equivalent open schema to avoid hard-coding observability into one vendor. Separate operational traces from immutable compliance evidence; their audiences and retention periods are different.
7. GDPR by design for AI systems
The GDPR is technology-neutral and continues to apply alongside the AI Act. The authoritative text is Regulation (EU) 2016/679.
7.1 Build a data-processing map
Before development, map each flow:
data subject → channel → application → retrieval → model provider → tool/API → telemetry → backup → deletion
For every hop, record:
- controller, joint controller and processor roles;
- purpose and lawful basis under Article 6;
- any Article 9 special-category condition;
- data categories and affected people;
- source, recipients and sub-processors;
- region and transfer mechanism;
- retention and deletion behaviour;
- security measures;
- rights-handling route.
“Improving the model” is too vague to be a purpose. Separate service delivery, abuse prevention, product analytics, human review, evaluation and training. Disable provider training or retention when it is not explicitly approved and lawful.
7.2 Apply the GDPR principles in architecture
| GDPR principle | Architectural implementation |
|---|---|
| Lawfulness, fairness and transparency | Purpose and lawful-basis registry; privacy notice; AI disclosure; meaningful explanation and escalation |
| Purpose limitation | Purpose-bound access policies; separate indexes and pipelines; prevent secondary training by default |
| Data minimisation | redact before prompts; retrieve small passages; avoid full records when fields suffice; short context windows by design |
| Accuracy | source ownership, versioning, correction workflow, stale-document expiry, user challenge route |
| Storage limitation | TTLs, retention classes, deletion jobs, backup expiry and vendor deletion verification |
| Integrity and confidentiality | least privilege, encryption, tenant isolation, secure SDLC, monitoring and incident response |
| Accountability | ROPA, DPIA, decision log, policy evidence, contracts, test results and audit trail |
Pseudonymised data remains personal data when re-identification is reasonably possible. Do not label data anonymous merely because a direct identifier was removed.
7.3 Data protection impact assessment
A DPIA is required where processing is likely to result in high risk to people's rights and freedoms. The European Commission highlights systematic and extensive evaluation or profiling, large-scale sensitive-data processing and large-scale public monitoring as clear examples. See the Commission's DPIA guidance and GDPR Article 35.
A practical AI DPIA should cover:
- necessity and proportionality;
- people and groups affected;
- data provenance and expectations;
- model and retrieval uncertainty;
- hallucination, bias and automation bias;
- data leakage, memorisation and re-identification;
- employee or customer power imbalance;
- solely automated decisions;
- children's and vulnerable people's interests;
- security threats and supplier dependencies;
- human intervention, appeal and redress;
- residual risk and approval.
Review the DPIA when the purpose, data, model, autonomy, user population, geography or risk changes.
7.4 Automated decisions and meaningful human oversight
GDPR Article 22 restricts decisions based solely on automated processing that produce legal or similarly significant effects, subject to defined exceptions and safeguards. “A human can review it” is not sufficient if the person merely rubber-stamps the output.
Meaningful oversight requires a reviewer who:
- has authority and time to change the result;
- can see relevant source evidence and limitations;
- is trained to detect automation bias;
- is not measured in a way that forces automatic agreement;
- can request more information;
- records reasons for accepting or overriding the recommendation;
- provides an accessible challenge and redress route.
The EDPB maintains official automated decision-making and profiling guidance.
7.5 Rights, retention and deletion
Design data-subject rights as platform capabilities, not manual database exercises. A rights service should locate data across conversations, source documents, relational rows, vector chunks, caches, annotations, evaluation sets, exports and relevant processors.
Deletion should:
- verify identity and scope;
- apply exceptions such as legal hold;
- delete or irreversibly anonymise primary records;
- remove vector chunks and rebuild affected indexes if necessary;
- purge caches and queued copies;
- send deletion requests to processors;
- prevent re-ingestion from the source;
- record completion evidence without retaining deleted content;
- allow encrypted backups to age out under a documented, access-restricted schedule.
7.6 International transfers and suppliers
For every model, embedding, moderation, OCR and telemetry provider, confirm:
- processor terms and Article 28 requirements;
- sub-processor list and change notice;
- data location and support access;
- retention, deletion and training settings;
- security and incident commitments;
- audit evidence;
- international-transfer mechanism, including adequacy or Standard Contractual Clauses where applicable;
- exit and data-portability plan.
The Commission describes adequacy, SCCs and binding corporate rules in its international data protection overview.
8. Operationalising the EU AI Act
8.1 Controls applicable across many systems
Even where a use case is not high-risk, teams should address:
- prohibited-practice screening;
- AI literacy for staff and people operating AI on the organisation's behalf;
- Article 50 disclosure for direct AI interaction where applicable;
- machine-readable marking and labelling duties for applicable synthetic or manipulated content;
- clear provider/deployer and supplier responsibilities;
- complaint, incident and authority-cooperation procedures;
- monitoring for a change of purpose, substantial modification or reclassification.
Article 4 requires providers and deployers to take measures supporting AI literacy appropriate to people's knowledge, experience, training and context. The consolidated Act contains the current Article 4 text. Training records, role-specific curricula and competence checks are useful evidence.
The Commission's Article 50 transparency guidance states that providers of systems interacting directly with people must design them so people are informed that they are interacting with AI, subject to the provision's exceptions and context.
8.2 Engineer the high-risk requirements now
For systems that are or may become high-risk, design for the Chapter III controls even before their applicable date:
| AI Act area | Engineering response | Typical evidence |
|---|---|---|
| Article 9 risk management | continuous hazard, misuse and residual-risk process across the lifecycle | risk register, test plan, acceptance records |
| Article 10 data and data governance | provenance, suitability, representativeness, quality, bias controls and documented preparation | dataset records, lineage, quality reports |
| Article 11 technical documentation | versioned system description, design, limitations, validation, monitoring and changes | system card, architecture, model and data documentation |
| Article 12 record-keeping | automatic logs sufficient for traceability and monitoring | protected logs, log schema, retention policy |
| Article 13 transparency to deployers | clear instructions, purpose, performance, limitations, oversight and maintenance | instructions for use, known-limitations register |
| Article 14 human oversight | trained humans can understand, monitor, disregard, override or stop the system | oversight SOP, UI evidence, training, override logs |
| Article 15 accuracy, robustness and cybersecurity | declared metrics, resilience, fail-safe behaviour and AI-specific security testing | evaluation, resilience and red-team reports |
| Articles 16–18 provider duties and QMS | controlled development, compliance, documentation and change management | policies, approvals, QMS records |
| Article 26 deployer duties | use according to instructions, oversight, monitoring and logs | deployment controls, operating procedures |
| Article 27 FRIA where applicable | assess effects on fundamental rights before deployment | fundamental-rights impact assessment |
| Post-market monitoring and incidents | collect performance and safety signals, investigate and correct | monitoring plan, incident and corrective-action records |
The consolidated Act specifies lifecycle risk management, testing against predefined metrics, data governance, logging, transparency, effective human oversight, and accuracy, robustness and cybersecurity requirements. It expressly identifies data poisoning, model poisoning, adversarial examples, model evasion and confidentiality attacks as relevant AI-security concerns. See Articles 9–15 in the consolidated text.
8.3 Do not confuse a DPIA, FRIA and AI impact assessment
These assessments overlap but answer different questions:
- a DPIA focuses on risks from personal-data processing under GDPR;
- a fundamental-rights impact assessment applies to specified deployers and high-risk uses under Article 27;
- an AI system impact assessment, such as the method supported by ISO/IEC 42005, can cover wider effects on individuals, groups and society;
- a security threat model focuses on adversaries, assets, attack paths and controls;
- a business risk assessment covers financial, operational, strategic and reputational effects.
Use a shared facts section and risk taxonomy, then maintain the legally required decisions and owners separately. One giant generic checklist can obscure which obligation has actually been satisfied.
9. Integrating ISO/IEC 42001 and ISO/IEC 27001
The cleanest operating model is an integrated management system: one control is designed once, assigned once and evidenced once, then mapped to multiple obligations.
9.1 ISO/IEC 42001 management cycle
The management-system clauses can be operationalised as:
| Management area | Practical implementation |
|---|---|
| Context | scope, stakeholders, laws, AI inventory, internal and external issues |
| Leadership | accountable executive, AI policy, roles, risk appetite and escalation |
| Planning | AI risks and opportunities, objectives, impact criteria and treatment plans |
| Support | competence, AI literacy, resources, communication and controlled documents |
| Operation | use-case intake, impact assessment, lifecycle controls, suppliers and change management |
| Performance evaluation | metrics, monitoring, internal audit and management review |
| Improvement | incidents, nonconformities, corrective action and continual improvement |
ISO/IEC 42001 Annex A control themes cover AI policies, internal organisation, resources, impact assessment, system lifecycle, data, information for interested parties, use of AI and third-party relationships. These should appear in the platform's operating model, not only in a certification folder.
9.2 ISO/IEC 27001 security foundation
ISO/IEC 27001:2022 requires an ISMS with context, leadership, planning, support, operation, performance evaluation and improvement, backed by a risk treatment process and Statement of Applicability. Its 93 Annex A controls are grouped into organisational, people, physical and technological controls.
For an AI platform, the ISMS scope should explicitly cover:
- model and embedding providers;
- prompt, workflow and policy repositories;
- training, evaluation and RAG data;
- vector databases and indexes;
- annotation and human-review platforms;
- agent tool credentials and enterprise APIs;
- AI telemetry and audit evidence;
- development, CI/CD and model-evaluation environments;
- third-party libraries, models and datasets.
9.3 One control, multiple obligations
For example, an immutable, access-controlled AI audit event can support:
- AI Act traceability and log requirements;
- GDPR accountability and security evidence;
- ISO/IEC 42001 monitoring and controlled records;
- ISO/IEC 27001 logging, monitoring and incident management;
- internal investigation and customer assurance.
But design the event to minimise personal data. “Keep everything forever in case of audit” conflicts with storage limitation and increases breach impact.
10. Threat model and security controls
OWASP's current resources include the Top 10 for LLM Applications 2026, the Top 10 for Agentic Applications 2026, ASVS 5.0 and the Software Component Verification Standard. Use them with ordinary application and cloud threat modelling; AI-specific risk does not replace web, API, identity, supply-chain or infrastructure security.
| Threat | Example | Preventive controls | Detective and recovery controls |
|---|---|---|---|
| Direct prompt injection | user asks the model to ignore policy and expose secrets | deterministic access control, prompt separation, minimal context, no secrets in prompts | adversarial tests, injection signals, trace review |
| Indirect prompt injection | malicious instruction hidden in a webpage or document | treat retrieved text as data, sandbox parsing, content provenance, tool isolation, egress allowlist | source anomaly detection, canary documents, quarantine and index rollback |
| Cross-tenant leakage | vector search returns another customer's document | pre-filter by verified tenant and ACL, RLS or physical isolation, tenant-bound cache keys | automated isolation tests, DLP alerts, kill switch |
| Sensitive information disclosure | model returns PII, credentials or confidential text | minimisation, redaction, data-class policies, retrieval ACLs, provider privacy settings | output DLP, access audit, incident workflow |
| RAG poisoning | compromised source inserts false or hostile content | authenticated connectors, owner approval, signing or hashing, versioned atomic indexes | source-drift monitoring, retrieval-quality alerts, rollback |
| Model or dataset supply-chain compromise | malicious model, library, adapter or dataset | approved registry, provenance, licence review, signed artefacts, SBOM and AI BOM, isolated scanning | vulnerability monitoring, integrity verification, rapid revoke |
| Improper output handling | generated HTML, SQL or command is executed | strict encoding, parameterised operations, typed schemas, never eval, output treated as untrusted | application security tests and runtime alerts |
| Excessive agency | agent sends money, email or data without valid approval | narrow tools, read-only default, least privilege, step and cost limits, human approval | action audit, reconciliation, transaction alerts, rollback |
| SSRF and egress abuse | agent tool requests cloud metadata or attacker URL | no generic fetch tool, URL allowlist, private DNS policy, metadata endpoint protection | network flow logs, egress alerts, credential rotation |
| Model extraction or inversion | repeated queries recover behaviour or memorised data | authentication, rate and similarity limits, minimise exposed confidence detail, privacy tests | abuse analytics, account suspension, provider coordination |
| Denial of wallet | attacker triggers long contexts and recursive tools | request, token, concurrency, step and spend budgets; backpressure | cost anomaly alerts, circuit breaker, tenant suspension |
| Hallucinated authority | assistant invents law, policy or account state | retrieval for factual claims, citations, abstention, deterministic source APIs | sampled review, groundedness and citation metrics |
| Memory poisoning | malicious content is stored as durable user or agent memory | explicit memory-write policy, typed memory, provenance, user review, TTL | memory-diff review, delete and rebuild |
| Insecure human approval | reviewer sees only an AI summary and rubber-stamps it | show source evidence, consequence and changed fields; separation of duties | override/approval analytics and quality review |
| Telemetry leakage | prompts and documents are copied to logs | metadata-only default, redaction, restricted debug capture | DLP scanning, log-access monitoring, purge runbook |
| Insider misuse | privileged operator retrieves customer content | just-in-time access, dual approval, session recording, segregation of duties | UEBA, immutable access logs, periodic review |
Zero trust for AI
NIST SP 800-207 moves trust away from network location and toward users, assets and resources. Apply this literally:
- every service has a workload identity;
- every call is authenticated and authorised;
- credentials are short-lived;
- policies consider user, device, tenant, workload and data sensitivity;
- model and tool providers are accessed through controlled gateways;
- production data is not reachable from developer laptops by default;
- administrative access is just-in-time, approved and recorded;
- east-west traffic is restricted, not implicitly trusted because it is inside a VPC.
Secure software and AI supply chain
Follow the final NIST SSDF 1.1 and its AI-specific SP 800-218A profile:
- protect source and build systems;
- review code and infrastructure as code;
- scan dependencies, containers, secrets and licences;
- produce an SBOM plus an AI bill of materials listing models, embeddings, datasets, prompts, evaluators and providers;
- sign artefacts and verify provenance at deployment;
- isolate untrusted model and document-processing code;
- patch or revoke vulnerable components;
- maintain reproducible builds and rollback artefacts.
11. Evaluation is the release contract
LLM quality cannot be represented by one aggregate score. Build a versioned evaluation portfolio.
Evaluation dimensions
| Dimension | Example measures |
|---|---|
| Business | task completion, first-contact resolution, handling time, conversion, user effort |
| Answer quality | correctness, completeness, relevance, instruction following, calibrated abstention |
| RAG | retrieval recall, precision at K, context relevance, groundedness, citation coverage and citation correctness |
| Safety | policy violation rate by severity, vulnerable-group scenarios, harmful completion rate |
| Privacy | PII leakage, memorisation probes, deletion effectiveness, data-minimisation conformance |
| Fairness | error and outcome differences across relevant groups and languages; accessibility impact |
| Security | prompt injection success, cross-tenant access, tool abuse, exfiltration, poisoning and denial-of-wallet resistance |
| Agent behaviour | valid tool selection, argument correctness, unauthorised action rate, loop rate, recovery rate |
| Reliability | availability, timeout, dependency failure, fallback correctness, queue age, recovery time |
| Performance | time to first token, total latency, retrieval latency, tool latency, throughput |
| Economics | tokens, provider cost, compute, storage and support cost per completed task |
Build representative test sets
Include:
- normal high-volume tasks;
- rare but high-impact cases;
- unanswerable and ambiguous questions;
- stale and conflicting sources;
- long, multilingual and accessibility-related inputs;
- relevant demographic and vulnerable-group scenarios;
- malicious files, direct and indirect injections;
- permission and cross-tenant tests;
- provider timeouts and partial failures;
- unsafe or irreversible tool requests;
- known production incidents converted into regression tests.
Use expert human review for high-impact outputs. LLM-as-judge can scale comparison, but validate the judge against expert labels, monitor positional and self-preference bias, and never make it the only release authority for critical safety or legal criteria.
Example release gates
| Gate | Example pass condition |
|---|---|
| Architecture | approved threat model, data-flow diagram and failure-mode review |
| Privacy | lawful basis confirmed; DPIA approved where required; deletion tested end to end |
| Security | no open critical findings; tenant isolation and tool-authorisation tests pass |
| Quality | use-case thresholds pass with confidence intervals and segment breakdowns |
| Safety | zero critical policy failures in the approved adversarial suite |
| Human oversight | reviewers can understand, override, stop and record reasons |
| Reliability | load, chaos, backup restore, failover and rollback tests pass |
| Compliance | inventory, system card, instructions, contracts and evidence pack complete |
| Operations | dashboards, alerts, on-call, runbooks, SLOs and kill switches verified |
| Business | named owner accepts expected benefit, cost and residual risk |
Do not average away a critical failure. A system with excellent overall accuracy but one cross-tenant disclosure has failed release.
12. Production deployment and LLMOps
A controlled pipeline should promote versioned bundles rather than independent moving parts. A release bundle can contain:
- application and infrastructure artefact digests;
- workflow graph version;
- system and developer prompt versions;
- policy bundle version;
- model and embedding versions;
- retrieval index and source snapshot;
- tool schemas and allowed scopes;
- evaluation dataset and result IDs;
- risk and approval records.
Suggested delivery flow
- Developer checks: formatting, unit, type, secret and dependency checks.
- Build: reproducible container and SBOM generation; artefact signing.
- Component tests: prompt templates, parsers, policy, retrieval and tool schemas.
- Integration tests: approved sandbox models, stores and enterprise APIs.
- AI evaluation: quality, safety, privacy, fairness, security, performance and cost.
- Security assurance: SAST, DAST, IaC, container and adversarial AI testing.
- Approval: business owner, model-risk or responsible-AI function, security, privacy and legal as required by tier.
- Staging: production-like data controls with synthetic or minimised data.
- Canary: small cohort, low-risk tools, close monitoring and automatic rollback thresholds.
- Progressive rollout: expand by tenant or traffic percentage after evidence review.
- Post-release review: compare expected and actual quality, safety, latency, cost and incidents.
Separate development, test and production accounts and keys. Production prompts, policies and configurations should be changed through reviewed version control, not a live console without traceability.
13. Reliability, scaling and graceful degradation
LLM systems combine multiple failure-prone dependencies. Engineer bulkheads so a slow model does not exhaust the entire API worker pool and a failing tool does not block ordinary knowledge queries.
Core resilience patterns
- explicit per-stage timeouts;
- retries only for safe, transient and idempotent operations;
- exponential backoff with jitter;
- circuit breakers per provider, model and tool;
- bounded queues and dead-letter handling;
- concurrency pools per tenant and workload;
- backpressure before saturation;
- idempotency for user requests and write actions;
- checkpoints for long-running workflows;
- health probes that test dependencies appropriately;
- multi-zone deployment and tested regional recovery;
- versioned backups and restoration exercises.
For streaming responses, distinguish connection success from task success. If a tool or citation check fails after partial generation, the interface needs a clear terminal error rather than silently presenting an incomplete answer as final.
Safe degradation ladder
An agentic system can degrade in stages:
- full agent with approved read and write tools;
- read-only tools;
- RAG answer with citations;
- search results without synthesis;
- static help content;
- human handoff.
The model gateway may route to another provider only where its legal, privacy, residency, quality and safety profile is already approved.
Privacy-safe caching
Cache only when the key includes all relevant security and correctness dimensions: tenant, user or entitlement scope, data version, model version, prompt version, policy version and locale. Never use a global semantic cache for responses based on private data. Encrypt sensitive cache content, apply short TTLs and purge entries when source permissions or documents change.
14. Monitoring, incidents and post-market operation
Monitor four types of drift
- Data drift: user, source, language or document distributions change.
- Quality drift: groundedness, citations, refusal or task success changes.
- Control drift: policies, permissions, vendors or configurations diverge from approval.
- Impact drift: the system is used for a different purpose, population or level of autonomy.
Alert on symptoms that require action, not every metric fluctuation. Examples include a critical safety regression, cross-tenant access attempt, anomalous write-tool volume, abrupt retrieval-quality drop, cost spike, provider-region change, repeated human overrides or unauthorised configuration change.
AI incident runbook
- Detect and preserve evidence. Create an incident ID and protect relevant versions and events without expanding unnecessary personal-data access.
- Triage harm and scope. Determine affected people, tenants, actions, data, geography and continuing risk.
- Contain. Disable a tool, model, source, index, tenant capability or the whole AI route using tested kill switches.
- Provide a safe service. Fall back to read-only, static content or human handling.
- Eradicate. Remove poisoned content, rotate credentials, patch components, correct policy or retrain where justified.
- Recover. Restore a known-good bundle, canary it and monitor closely.
- Notify. Involve security, privacy, legal, the business owner, affected customers and authorities under the applicable rules.
- Learn. Complete root-cause analysis, corrective action, risk updates and regression tests.
Under GDPR Article 33, a controller must notify the competent supervisory authority without undue delay and, where feasible, within 72 hours after becoming aware of a personal-data breach, unless it is unlikely to result in a risk to people's rights and freedoms. High-risk cases can also require communication to affected people. See the GDPR breach provisions. Do not wait for perfect technical certainty before involving the DPO and incident team.
15. Required governance artefacts and evidence
Maintain evidence as part of delivery, not as a retrospective documentation sprint.
| Artefact | Minimum contents | Primary owner |
|---|---|---|
| AI system inventory entry | owner, purpose, users, regions, model, data, tools, risk tier, status | AI governance |
| Use-case contract | outcomes, intended purpose, excluded uses, autonomy and human oversight | product owner |
| Regulatory classification | AI Act role and risk, GDPR scope, sector obligations, review triggers | legal/compliance |
| Data-flow and processing record | systems, data categories, roles, purposes, transfers, retention | privacy/data owner |
| DPIA/FRIA/AI impact assessment | affected groups, risks, mitigations, residual decisions | DPO/responsible AI |
| System card | architecture, versions, capabilities, limitations, metrics and intended use | technical owner |
| Data and retrieval record | sources, rights, lineage, quality, ACLs, index and deletion | data owner |
| Model/provider record | model card, terms, location, training, retention, sub-processors, evaluation | model owner/procurement |
| Threat model | assets, trust boundaries, abuse cases, controls and residual risks | security architect |
| Evaluation report | datasets, methods, segment results, failures, thresholds and approvals | evaluation lead |
| Human-oversight procedure | reviewer authority, evidence, override, stop, escalation and training | operations owner |
| Deployment record | signed bundle, approvals, canary, rollback and change reason | platform owner |
| Monitoring plan | SLOs, risk metrics, thresholds, review cadence and owners | service owner |
| Incident and continuity plan | scenarios, contacts, kill switches, notification, RTO and RPO | incident/service owner |
| Decommission record | access removal, data disposal, vendor exit and retained evidence | system owner |
Each artefact should have an owner, version, approver, review date, evidence links and change triggers.
16. End-to-end enterprise case study
Consider a retail bank building a customer-service AI assistant for EU customers. The assistant answers product and fee questions using approved policy documents, retrieves customer-specific case status after authentication, and can initiate a small set of service workflows. It does not make credit, pricing, fraud, employment or legal decisions.
Phase 1: discovery and classification
The bank defines three outcomes:
- reduce average handling time by 20%;
- achieve at least 90% correct, source-backed answers on supported intents;
- improve successful self-service without reducing complaint resolution or accessibility.
The use-case contract prohibits lending decisions, financial advice, vulnerability inference, payment initiation and autonomous complaint rejection. Legal and compliance determine that the customer is interacting with an AI system and Article 50 transparency applies. GDPR applies because chat, account and case information contain personal data. A DPIA is initiated because of the scale, monitoring and potential impact.
The classification record says that adding creditworthiness or insurance-pricing decisions would trigger a new assessment and may move the system into an Annex III high-risk use. That change cannot be enabled through ordinary feature configuration.
Phase 2: data and supplier controls
The team inventories:
- public product pages;
- controlled policy manuals;
- fee tables;
- customer account and case-status APIs;
- chat metadata and feedback;
- the model, embedding, moderation and observability suppliers.
Documents receive an owner, approval status, valid dates, sensitivity, ACL and retention class. Expired prices are automatically removed from the active index. Contracts prohibit provider training on bank prompts and responses, define regional processing, sub-processors, deletion, incident notice and audit evidence.
Phase 3: architecture
The system uses:
- enterprise customer identity and step-up authentication;
- a WAF and API gateway with user and tenant quotas;
- a stateless conversation API;
- a workflow orchestrator with bounded execution;
- a policy service that combines user assurance, purpose, data class and tool risk;
- separate public and customer-authorised retrieval collections;
- a model gateway with EU-approved endpoints;
- a tool gateway exposing only typed bank APIs;
- metadata-only traces and a separate protected audit store.
The model never receives core-banking credentials. The account-status tool derives the customer relationship from the authenticated session and returns only necessary fields. Address changes require step-up authentication, a preview of the exact change and explicit confirmation. Card freezing is handled by an existing deterministic bank workflow; the AI may navigate the user to it but cannot bypass its controls.
Phase 4: a live request
A customer asks: “Why was I charged £12 and can you refund it?”
- The gateway verifies the session and applies quotas.
- The input classifier detects financial and personal context.
- The policy engine allows read-only account and fee-policy tools but denies refund execution.
- The account API returns a typed transaction category, not an unrestricted statement export.
- Retrieval fetches the current fee policy and terms, applying version and regional filters.
- The model produces a draft explanation with citations.
- The output validator checks that the fee and eligibility claims are supported.
- The assistant explains the charge, says it cannot decide the refund, and offers a controlled complaint or human-review route.
- The trace records document IDs, policy, model and tool versions without copying the full transaction description into ordinary telemetry.
This is safer than asking one model to read an entire account, decide what happened and directly issue money.
Phase 5: evaluation and rollout
The bank builds a test set of normal, ambiguous, vulnerable-customer, multilingual, stale-policy, injection and unauthorised-account cases. It measures answer correctness, citation correctness, unsupported claim rate, privacy leakage, tool authorisation, accessibility, latency and cost.
Release requires:
- zero cross-customer access failures;
- zero unauthorised write actions;
- no critical harmful output in the approved test suite;
- target citation and task-quality thresholds by language and customer segment;
- successful rollback, provider outage and deletion tests;
- approved DPIA, operating instructions and incident runbook.
The first release is read-only for a small cohort. Write-capable workflows are introduced separately after additional authentication, authorisation and human-factors testing.
Phase 6: operation
Dashboards track groundedness samples, user escalation, document staleness, policy denials, human overrides, provider latency, token cost and complaint signals. A model-version change automatically blocks production promotion until the regression suite passes. A source-poisoning drill proves that the team can quarantine a connector and atomically restore the prior index.
The result is not “a chatbot with compliance added.” It is a controlled service in which AI is one component.
17. A practical implementation roadmap
Weeks 1–2: scope and risk
- appoint business, technical, data, security, privacy and operational owners;
- define outcome, intended purpose, excluded use and authority boundaries;
- classify under the AI Act and relevant sector rules;
- map data and suppliers;
- start DPIA, FRIA or wider impact assessment as applicable;
- define risk appetite and release tier.
Weeks 3–5: platform foundation
- establish SSO, tenant model and policy engine;
- deploy edge, secrets, key and private-network controls;
- implement model gateway and approved catalogue;
- create prompt, policy, model and data registries;
- establish privacy-safe telemetry and audit schemas;
- build CI/CD, signing, SBOM and infrastructure-as-code controls.
Weeks 6–8: RAG and tools
- implement quarantined ingestion, lineage and document ACL preservation;
- build secure retrieval, reranking and citations;
- expose narrow read-only tools through the tool gateway;
- add bounded orchestration, input and output controls;
- implement rights, retention and deletion workflows.
Weeks 9–11: assurance
- create representative evaluation sets;
- conduct application, AI and agent threat testing;
- test tenant isolation, prompt injection, poisoning and denial of wallet;
- run load, chaos, backup, restore, regional failover and rollback tests;
- validate human oversight and accessibility;
- complete supplier, legal, privacy and responsible-AI approvals.
Week 12 and beyond: controlled operation
- canary to a small, low-risk cohort;
- review actual outcomes against assumptions;
- expand progressively;
- run periodic access, supplier, model, data, risk and management reviews;
- convert incidents and near misses into regression tests;
- decommission data, permissions and suppliers when the use case ends.
18. Final production checklist
Before launch, a service owner should be able to answer “yes” to all of these:
Purpose and governance
- Is the intended purpose precise, approved and visible to the delivery team?
- Are prohibited uses and escalation triggers enforceable?
- Are provider, deployer, controller and processor roles documented?
- Is the AI Act risk classification current and reviewed after material changes?
- Are AI literacy and operating competence appropriate to each role?
Privacy and data
- Are purpose, lawful basis, Article 9 condition and notices confirmed where relevant?
- Is a DPIA completed where required?
- Are source rights, provenance, quality, ACLs and retention documented?
- Is personal data minimised before retrieval and model calls?
- Do access, correction, objection, restriction, portability and deletion processes reach all stores and processors?
- Are transfers, sub-processors and provider-training settings approved?
Security
- Are user and workload identities authenticated with least privilege?
- Is tenant isolation enforced in database, vector, cache, queue and tool paths?
- Are retrieved content and model output treated as untrusted?
- Are tools narrow, typed, independently authorised and bounded?
- Are egress, secrets, keys and production administration controlled?
- Have application, LLM, RAG, agent and supply-chain threats been tested?
Quality and safety
- Are evaluation datasets representative, adversarial and versioned?
- Are critical criteria treated as hard gates rather than averaged scores?
- Can the system abstain, cite, escalate and fail safely?
- Is human oversight meaningful and usable?
- Are limitations communicated to users and operators?
Reliability and operation
- Are SLOs, budgets, quotas, timeouts, retries and circuit breakers defined?
- Are backup, restore, failover, rollback and kill switches tested?
- Are fallback providers legally and technically approved?
- Are privacy-safe dashboards, alerts, on-call and runbooks live?
- Can the team reconstruct a material output or action from protected evidence?
- Is there a monitored decommission and vendor-exit plan?
19. A concrete, portable implementation stack
A reference architecture is only useful when the implementation team can map each responsibility to an actual technology decision. The following stack is an example, not a universal prescription.
| Capability | Portable implementation pattern | Why it belongs in the architecture |
|---|---|---|
| Web application | React or Next.js with a server-side backend-for-frontend | controls session boundaries, response streaming, citation rendering and user confirmation |
| Public API | FastAPI, Go, Java/Spring or ASP.NET with a versioned OpenAPI contract | provides predictable authentication, request validation, streaming and operational endpoints |
| Edge and gateway | managed API gateway or Envoy/Kong-class gateway | centralises WAF integration, quotas, request limits and service identity |
| Identity | enterprise OIDC/SAML provider plus workload identity | preserves joiner/mover/leaver processes and avoids long-lived application secrets |
| Authorisation | policy engine implementing RBAC plus attribute or relationship checks | makes tenant, purpose, document and tool decisions enforceable outside the model |
| Orchestration | explicit state machine; durable workflow engine for long-running jobs | bounds autonomy, makes approvals resumable and provides deterministic state transitions |
| Relational data | PostgreSQL-compatible managed database | supports transactional integrity, row-level security and a strong operational ecosystem |
| Vector search | PostgreSQL with pgvector initially; dedicated search/vector engine when required | balances operational simplicity against retrieval scale, filtering and hybrid-search needs |
| Source documents | object storage with versioning and lifecycle policies | preserves immutable sources, parser inputs and auditable document history |
| Cache and quotas | Redis-compatible managed cache | supports short-lived sessions, distributed rate limits, idempotency and safe caching |
| Async processing | managed queue, Kafka-compatible event stream or workflow queue | isolates ingestion, evaluation, notification and recovery workloads |
| Model access | approved internal model gateway | centralises model choice, region, retention, cost, safety and failover controls |
| Tool execution | separate typed tool gateway with dedicated workload identities | prevents a model from inheriting broad enterprise credentials |
| Observability | OpenTelemetry plus metrics, logs, traces and a protected audit store | correlates reliability and AI behaviour while preserving portability |
| Secrets and keys | managed KMS/HSM and secret manager | supports key ownership, rotation, envelope encryption and emergency revocation |
| Infrastructure | Terraform, OpenTofu, Bicep, CloudFormation or equivalent | makes infrastructure reviewable, repeatable and policy-checkable |
When PostgreSQL is sufficient for vector search
A single PostgreSQL platform with vector search can be an excellent first production choice when:
- the corpus is modest enough for the selected database tier;
- tenant and access filters fit naturally into SQL;
- the team values transactional consistency between document metadata and embeddings;
- hybrid search and reranking requirements can be implemented without excessive complexity;
- operational simplicity is more valuable than an additional specialised datastore.
Move to a dedicated vector or search service when you need materially larger indexes, specialised hybrid ranking, high concurrent query volume, document-level permissions, geographically distributed search, advanced retrieval pipelines or independent scaling of ingestion and serving.
The correct decision is measurable: benchmark retrieval quality, p95/p99 latency, filtered-query performance, tenant isolation, deletion propagation, index rebuild time and total operating cost against a realistic corpus.
Keep orchestration separate from durability
An agent framework can help express a retrieval or tool graph. A durable workflow engine provides a different guarantee: recovering state across crashes, long approval windows and retries. A mature solution may use both, but should avoid hiding material business transactions inside an unbounded prompt loop.
For example:
- the AI graph drafts a refund recommendation;
- the durable workflow records the draft and waits for approval;
- a human reviews the source evidence and amount;
- a deterministic payment service validates policy and executes the approved action;
- the workflow records completion or compensates for failure.
20. Cross-cloud implementation mapping
The logical architecture should remain stable even when the cloud provider changes. Map capabilities rather than assuming similarly named services have identical semantics, regional availability, residency guarantees or compliance features.
| Architecture capability | AWS example | Azure example | Google Cloud example |
|---|---|---|---|
| Internet edge and WAF | CloudFront, AWS WAF and API Gateway | Azure Front Door, WAF and API Management | Cloud CDN, Cloud Armor and Apigee or API Gateway |
| Enterprise identity | IAM Identity Center, Amazon Cognito and IAM roles | Microsoft Entra ID and managed identities | Cloud Identity, Identity Platform and workload identities |
| Container runtime | Amazon EKS, ECS or Fargate | Azure Kubernetes Service or Container Apps | Google Kubernetes Engine or Cloud Run |
| Foundation-model access | Amazon Bedrock or SageMaker AI | Microsoft Foundry and Azure OpenAI models | Vertex AI and its model catalogue |
| AI gateway | API Gateway, internal gateway or Bedrock AgentCore Gateway | Azure API Management AI gateway | Apigee or an internal model gateway in front of Vertex AI |
| Governed agent execution | Amazon Bedrock AgentCore | Microsoft Foundry agent capabilities | Gemini Enterprise Agent Platform |
| Retrieval | Bedrock Knowledge Bases, OpenSearch or PostgreSQL vectors | Azure AI Search, Foundry IQ or PostgreSQL vectors | Vertex AI retrieval, AlloyDB or a managed search/vector service |
| Relational database | Amazon Aurora PostgreSQL or Amazon RDS for PostgreSQL | Azure Database for PostgreSQL | AlloyDB or Cloud SQL for PostgreSQL |
| Source documents | Amazon S3 | Azure Blob Storage | Cloud Storage |
| Caching and quotas | ElastiCache for Redis | Azure Managed Redis or an approved Redis-compatible service | Memorystore for Redis |
| Asynchronous work | Amazon SQS, EventBridge or MSK | Service Bus, Event Grid or Event Hubs | Pub/Sub or Cloud Tasks |
| Secrets and encryption | AWS Secrets Manager and AWS KMS | Key Vault and managed HSM | Secret Manager and Cloud KMS |
| Private connectivity | VPC endpoints and AWS PrivateLink | Private Link and private endpoints | Private Service Connect and VPC Service Controls where applicable |
| Operational monitoring | CloudWatch, X-Ray and OpenTelemetry | Azure Monitor, Application Insights and OpenTelemetry | Cloud Monitoring, Cloud Logging, Cloud Trace and OpenTelemetry |
| Governance and inventory | AWS Config, Security Hub and model or platform registries | Azure Policy, Microsoft Purview and Foundry governance | Cloud Asset Inventory, Security Command Center and Vertex AI governance |
Current vendor guidance supports this capability-oriented approach. AWS documents both a production generative-AI architecture and an Agentic AI Lens. Microsoft describes the layered Microsoft Foundry architecture and API Management AI gateway. Google publishes an agentic architecture component-selection guide and a RAG reference architecture.
Watch for platform lifecycle changes
Cloud products and agent runtimes evolve rapidly. AWS documents that Bedrock Agents Classic is in maintenance mode, while Google documents that Vertex AI Extensions is deprecated and directs customers toward its Gemini Enterprise Agent Platform. New designs should validate product status, contractual commitments, regional availability, identity support, exportability and migration paths before committing.
Example cloud-selection criteria
Weight the decision using:
- enterprise identity and existing cloud operating model;
- required EU processing regions and actual model availability;
- private connectivity and egress control;
- model and embedding selection;
- document-level permissions and retrieval quality;
- agent identity, tool policy and human approval capabilities;
- processor, sub-processor and retention commitments;
- observability export and evidence availability;
- internal skills, supplier concentration and exit feasibility;
- three-year total cost of ownership.
Do not select a provider solely because a model demo looks stronger. Architecture quality depends on the complete control and operating model.
21. Data contracts that make governance enforceable
Every important boundary should exchange a typed contract. The examples below are illustrative and intentionally exclude raw sensitive content from general audit events.
AI request envelope
{
"request_id": "req_01JEXAMPLE",
"trace_id": "trc_01JEXAMPLE",
"tenant_id": "tenant_eu_123",
"principal": {
"subject_id": "user_458",
"roles": ["support_adviser"],
"assurance_level": "mfa",
"region": "eu-west",
"allowed_purposes": ["customer_support"]
},
"use_case": "banking_customer_support",
"data_classification": "confidential_personal",
"policy_version": "policy-2026-08-21.4",
"workflow_version": "support-rag-v12",
"limits": {
"maximum_input_tokens": 2500,
"maximum_output_tokens": 600,
"maximum_tool_calls": 3,
"maximum_duration_ms": 10000,
"maximum_cost_units": 20
},
"approved_model_region": "eu",
"human_approval_required_for": ["refund", "address_change"]
}
The server derives the identity, tenant and assurance fields from authenticated context. They are not trusted merely because the browser included them in a JSON body.
RAG document metadata contract
{
"tenant_id": "tenant_eu_123",
"document_id": "doc_fee_policy_2026_08",
"document_version": "8",
"chunk_id": "chunk_015",
"source_uri": "controlled-document-reference",
"owner_id": "bank_policy_team",
"classification": "internal",
"permitted_roles": ["support_adviser"],
"permitted_regions": ["eu"],
"purpose_tags": ["customer_support"],
"valid_from": "2026-08-01T00:00:00Z",
"valid_until": "2026-12-31T23:59:59Z",
"retention_policy": "policy_document_standard",
"content_hash": "sha256:example",
"embedding_model_version": "approved-embedding-v3",
"parser_version": "document-parser-v6"
}
Enforcement should occur before retrieval and again before a passage enters model context. Version dates prevent an expired fee schedule from being presented as current policy.
Tool execution contract
{
"tool_name": "create_refund_review",
"action_type": "write",
"tenant_id": "tenant_eu_123",
"resource_id": "case_9981",
"requested_by": "user_458",
"business_purpose": "customer_support",
"risk_tier": "material_financial_action",
"idempotency_key": "idem_01JEXAMPLE",
"arguments": {
"transaction_reference": "txn_772",
"requested_amount_minor_units": 1200,
"currency": "GBP"
},
"required_controls": {
"step_up_authentication": true,
"human_approval": true,
"separation_of_duties": true
},
"approval_reference": "approval_227"
}
The target business service must still validate amount, ownership, currency, approval and state transition independently. The tool gateway reduces risk; it does not replace the destination system's own security.
Privacy-minimised audit event
{
"event_id": "evt_01JEXAMPLE",
"occurred_at": "2026-08-21T12:30:00Z",
"tenant_id": "tenant_eu_123",
"actor_reference": "pseudonymous-user-reference",
"event_type": "tool.execution.denied",
"use_case": "banking_customer_support",
"policy_version": "policy-2026-08-21.4",
"model_version": "approved-chat-model-v18",
"workflow_version": "support-rag-v12",
"tool_name": "execute_refund",
"decision": "deny",
"reason_code": "human_approval_missing",
"trace_id": "trc_01JEXAMPLE",
"integrity_reference": "signed-audit-event-reference"
}
The event provides accountability without recording the customer's complete prompt, bank statement or generated answer.
22. Policy enforcement outside the model
A production policy decision should combine identity, data, tool, model and transaction attributes. A readable example:
policy_id: banking-support-refund-review
applies_to:
use_case: banking_customer_support
tool: create_refund_review
allow_when:
- principal.tenant_id == resource.tenant_id
- principal.role in [support_adviser, support_supervisor]
- principal.assurance_level in [mfa, phishing_resistant_mfa]
- request.business_purpose == customer_support
- resource.region == approved_processing_region
- request.amount_minor_units <= principal.approval_limit
- request.idempotency_key is present
require_human_approval_when:
- request.amount_minor_units > 10000
- customer.is_vulnerable == true
- tool.action_type == write
deny_when:
- resource.owner_id != verified_customer_id
- model.region not in approved_regions
- provider.data_retention_policy != approved_retention_policy
- tool.destination not in approved_enterprise_endpoints
Treat the snippet as illustrative policy logic, not a production-ready policy language. A real implementation must define evaluation order, fail-closed behaviour, missing-attribute semantics, emergency overrides and test coverage.
The four checks every tool action needs
- Identity: Is the requester authenticated at the required assurance level?
- Authority: Is this principal allowed to perform this action on this resource?
- Context: Is the purpose, tenant, data class, region and model route permitted?
- Consequence: Does the action require confirmation, approval, amount limits, separation of duties or compensation?
If any answer is missing, deny the action. Asking the model to “be careful” is not equivalent.
23. Multi-region deployment and data residency
For regulated workloads, availability and residency must be designed together. A global user experience does not automatically justify copying customer prompts, source documents or traces to every cloud region.
Recommended residency model
- assign each tenant a documented home region;
- keep prompts, conversation data, source documents, embeddings, traces and backups within the approved regional boundary;
- replicate policy definitions and deployment artefacts globally only when they do not contain restricted customer content;
- use a second approved region for recovery when permitted by the customer contract and transfer assessment;
- route based on verified tenant metadata, not browser location;
- prevent an external model fallback from escaping the tenant's approved region;
- document support-person access and remote administrative access, which may themselves create data-transfer concerns.
Active-active versus active-passive
Active-active improves resilience and latency but complicates consistency, duplication, conflict handling, deletion, region-specific contracts and residency. Active-passive is simpler for many regulated workloads but introduces recovery delay.
| Pattern | Appropriate when | Main trade-off |
|---|---|---|
| Single region, multiple zones | the business tolerates regional interruption and requires simple residency | regional outage interrupts service |
| Active-passive within approved regions | low data-loss and predictable recovery matter more than zero interruption | failover and recovery must be tested |
| Active-active within one legal boundary | high availability and local latency justify operational complexity | replication, consistency and deletion are harder |
| Dedicated regional tenant cells | residency, isolation or large-customer needs dominate | more infrastructure, deployment and support overhead |
Recovery objectives are meaningful only after restoration tests. A backup policy that has never been restored is a claim, not evidence.
24. Capacity planning and performance engineering
LLM architecture requires sizing by requests, concurrency, tokens, retrieval operations, tool calls and provider quota. Request-per-second figures alone hide the main bottlenecks.
Worked example
Assume a peak of 50 customer requests per second and an average complete-request service time of 5 seconds. By Little's Law:
average concurrent requests = arrival rate × average service time
Therefore:
50 requests/second × 5 seconds = 250 average in-flight requests
If each request averages 1,800 input tokens and 350 output tokens:
- input throughput is approximately 90,000 tokens per second;
- output throughput is approximately 17,500 tokens per second;
- input demand is approximately 5.4 million tokens per minute;
- output demand is approximately 1.05 million tokens per minute;
- an average of two retrieval queries per request creates about 100 retrieval queries per second;
- 20 reranked passages per request create roughly 1,000 reranked passages per second;
- a 30% tool-use rate creates approximately 15 tool calls per second.
These are averages, not provisioning guarantees. Include burst factors, p95 and p99 service times, provider throttling, context growth, regional failure and tenant concentration. Model quotas may be stated in requests per minute, input tokens per minute, total tokens per minute or other provider-specific units, so verify the exact quota semantics.
Performance budget
Allocate an explicit end-to-end latency budget:
| Stage | Illustrative p95 budget |
|---|---|
| Gateway, authentication and policy | 150 ms |
| Query rewriting and metadata filtering | 150 ms |
| Search and reranking | 700 ms |
| Tool lookup, where applicable | 900 ms |
| Model time to first token | 1,500 ms |
| Response generation and validation | 3,500 ms |
| Network and contingency | 1,100 ms |
| Total indicative budget | 8,000 ms |
Do not add independent p95 measurements and claim the sum is the actual end-to-end p95. Use distributed traces and realistic load tests; the table is a design budget, not a statistical identity.
Cost formula
Calculate unit economics using:
cost per completed task = model input + model output + embeddings + retrieval + tool execution + compute + storage + observability + human review + support
Then calculate:
net business value per task = avoided manual cost + incremental revenue + risk reduction - total task cost
Optimisation order should be:
- remove unnecessary or unlawful data;
- improve retrieval precision;
- reduce context and output budgets;
- route easy tasks to approved smaller models;
- cache only under entitlement-safe keys;
- batch asynchronous work;
- reserve or optimise capacity when usage is predictable;
- preserve human review and safety controls for high-impact cases.
Never lower cost by moving data to an unapproved provider or removing a control required by the risk assessment.
25. Choosing a tenant-isolation model
Multi-tenancy is an architecture spectrum rather than a binary choice.
| Isolation model | Data boundary | Typical fit | Risks and obligations |
|---|---|---|---|
| Shared tables with row-level security | logical tenant filters in one database | early-stage or lower-risk SaaS | demands rigorous policy, index, cache and test discipline |
| Separate schema or database per tenant | database-level boundary | customers with stronger isolation or recovery requirements | increases migration, backup and connection-management complexity |
| Dedicated search index per tenant | retrieval-level boundary | confidential document corpora and strict access requirements | more indexes, ingestion workflows and operational cost |
| Regional cell per tenant group | separate runtime and data plane per jurisdiction or segment | data residency and blast-radius reduction | requires cell-aware deployment and routing |
| Dedicated customer environment | isolated cloud account/subscription/project and data stores | regulated strategic customers with contractual separation | highest operating and provisioning overhead |
An effective design often combines models. For example, ordinary customers can share an application cell, while high-assurance customers receive a dedicated regional cell and customer-managed key.
Mandatory isolation test cases
Include:
- modified tenant ID in an API body;
- valid user token against another tenant's document ID;
- cross-tenant vector nearest-neighbour query;
- cache hit after changing role or document ACL;
- queued message replay in the wrong tenant;
- agent memory referencing another tenant;
- model fallback that violates residency;
- administrator search without approved support access;
- source deletion followed by immediate retrieval;
- report export with mixed tenant data.
These are release-blocking security tests, not optional quality checks.
26. Compliance traceability down to individual controls
A control catalogue connects architecture, law, evidence and accountability. The following identifiers are illustrative internal controls; they are not official ISO or legal clause numbers.
| Internal control | Engineering implementation | Regulatory or standard relationship | Evidence |
|---|---|---|---|
| AI-GOV-01 intended-purpose register | approved system inventory, excluded uses and change triggers | AI Act classification; ISO/IEC 42001 scope and operational governance | approved use-case record and review history |
| AI-GOV-02 risk and impact review | risk register, DPIA, FRIA where applicable and impact assessment | GDPR Articles 25/35; AI Act Articles 9/27 where applicable; ISO/IEC 42005 | assessments, treatment actions and residual-risk acceptance |
| AI-ID-01 verified tenant identity | federated login, short-lived claims and workload identity | GDPR Article 32; ISO/IEC 27001 access control | identity configuration and access tests |
| AI-ID-02 attribute-based document access | pre-search filtering, row-level security and ACL recheck | data minimisation and confidentiality; high-risk logging and governance where applicable | policy tests and tenant-isolation reports |
| AI-DATA-01 source provenance | source owner, rights, hash, validity, sensitivity and document ACL | AI Act Article 10 where applicable; ISO/IEC 42001 lifecycle/data controls | dataset and ingestion records |
| AI-DATA-02 retention and erasure | TTLs, deletion propagation and processor deletion workflow | GDPR Articles 5 and 17 | deletion test and retention configuration |
| AI-MODEL-01 approved provider catalogue | allowed model, region, contract, retention and fallback policy | GDPR Articles 28 and 44–49; supplier controls | vendor assessment, DPA and transfer record |
| AI-MODEL-02 versioned release bundle | pinned models, prompts, policies, indexes and evaluations | AI Act technical documentation and QMS where applicable; ISO change management | signed release manifest and approvals |
| AI-TOOL-01 least-privilege action gateway | typed tools, scoped credentials and deterministic validation | GDPR Article 32; ISO/IEC 27001 secure development and access control | tool schema, workload permissions and test results |
| AI-TOOL-02 meaningful approval | evidence-based human review, override and stop capability | GDPR Article 22 where applicable; AI Act Article 14 for high-risk systems | approval UI, reviewer training and override log |
| AI-EVAL-01 multidimensional evaluation | quality, bias, privacy, security and resilience release gates | AI Act Articles 9 and 15 where applicable; ISO/IEC 42001 performance evaluation | versioned evaluation report |
| AI-OBS-01 privacy-minimised traceability | integrity-protected event history and limited operational telemetry | GDPR accountability/security; AI Act Article 12 for high-risk systems | audit schema, access logs and retention settings |
| AI-OPS-01 incident response | kill switch, escalation, breach analysis and corrective action | GDPR Articles 33/34; AI Act incident duties where applicable | runbooks, exercises and incident records |
| AI-HR-01 AI literacy | role-specific training for developers, operators and oversight staff | AI Act Article 4; ISO/IEC 42001 support and competence | curriculum, attendance and competence records |
The same implementation can satisfy multiple requirements, but the evidence must show why it is adequate for each obligation. A generic security certificate is not proof that a particular retrieval index preserves document permissions.
27. Operational runbooks for common AI failures
Runbook A: model quality regression
Trigger: Groundedness, citation correctness or critical task success drops below its alert threshold.
- Confirm whether the change correlates with model, prompt, policy, index or source version.
- Compare affected tenants, languages, document types and risk categories.
- Block the suspect version from further rollout.
- Route to an already approved prior model or workflow version.
- Re-evaluate against production-derived and expert-labelled cases.
- Correct the root cause and require a fresh release decision.
Runbook B: poisoned or compromised knowledge source
Trigger: An ingestion integrity alert, suspicious source instructions, false policy answers or anomalous document changes.
- Pause the connector and quarantine newly ingested content.
- Disable retrieval from the affected source or index partition.
- Restore the prior approved index snapshot.
- Identify all answers, users and tools influenced by the malicious document.
- Assess whether private data was exposed or unauthorised actions occurred.
- Notify source owners, security, privacy and affected business teams as required.
- Add a regression case and strengthen source approval or provenance checks.
Runbook C: unauthorised agent action
Trigger: A tool call executes beyond role, amount, purpose, tenant or approval policy.
- Disable the affected tool globally or for the tenant.
- Revoke the tool workload identity and rotate relevant credentials.
- Determine whether the target business service independently enforced authorisation.
- Reconcile every transaction and execute compensating actions when possible.
- Preserve model, tool, prompt, policy, approval and trace versions.
- Assess legal, financial and notification consequences.
- Fix the deterministic control before re-enabling the tool.
Runbook D: model-provider outage
Trigger: Provider errors, throttling, degraded latency or an unavailable approved region.
- Open the model-specific circuit breaker.
- Check whether a pre-approved fallback matches region, data class and contract.
- If approved, shift a small canary cohort and monitor quality.
- If not approved, degrade to search-only, static support or human escalation.
- Apply backpressure and tenant communication.
- Review recovery, cost and downstream queue behaviour before returning traffic.
Runbook E: sudden token or infrastructure spend
Trigger: Cost per task, token volume, context size or recursive tool use spikes.
- Enforce per-user, tenant and model budgets.
- Identify oversized documents, recursive loops, abuse or fallback changes.
- Restrict maximum context, generation and tool steps within approved quality limits.
- Suspend suspicious traffic and alert account owners.
- Re-run quality and privacy checks before changing routing or caching.
28. Architecture decisions and implementation handoff
An enterprise solution architect should leave behind decisions that another team can understand, challenge and operate.
Minimum architecture decision record
Each architecture decision record should include:
- decision ID and date;
- business context and required outcome;
- options considered;
- chosen option and reasoning;
- security, privacy, compliance, cost and availability impacts;
- rejected alternatives;
- owner and approvers;
- assumptions and review triggers;
- rollback or exit route.
Example ADR: vector storage
Decision: Start with PostgreSQL and vector search for the first production release.
Why: The corpus is within the benchmarked capacity envelope, document permissions join naturally with relational ownership, and the team can reuse existing backup, encryption and operational controls.
Trade-off: Search and indexing may need a dedicated engine if corpus size, filtered-query latency or multi-tenant concurrency grows.
Review triggers: p95 retrieval above the approved budget, index rebuild above the recovery target, sustained CPU saturation, loss of retrieval quality or customer demand for dedicated index isolation.
Example ADR: provider fallback
Decision: Disable automatic cross-provider failover for confidential personal data.
Why: The secondary provider has not yet completed the required region, processor, sub-processor and retention approvals.
Trade-off: During the primary provider outage, the product degrades to source search and human escalation.
Review trigger: Completion of equivalent privacy, security, evaluation and contractual assurance for the secondary route.
Handoff package
A practical delivery repository can be organised around:
architecture/
01-business-context-and-intended-purpose.md
02-system-context-and-trust-boundaries.md
03-component-and-deployment-architecture.md
04-data-flow-and-processing-record.md
05-rag-ingestion-and-retrieval-design.md
06-agent-tools-and-approval-model.md
07-identity-tenant-and-policy-design.md
08-eu-ai-act-gdpr-and-iso-control-crosswalk.md
09-threat-model-and-red-team-plan.md
10-evaluation-strategy-and-release-gates.md
11-observability-incident-and-recovery-runbooks.md
12-capacity-cost-and-commercial-model.md
13-supplier-and-model-due-diligence.md
14-architecture-decision-records.md
15-implementation-backlog-and-acceptance-tests.md
The implementation backlog should translate each control into an owned, testable story. For example:
As a security owner, I need every vector query to include verified tenant and document-entitlement filters before retrieval, so that an authenticated user cannot infer or retrieve another tenant's content.
Acceptance criteria:
- query execution fails closed when tenant context is missing;
- the tenant is derived from the authenticated principal;
- search filters are applied before similarity retrieval;
- document ACLs are checked again before prompt assembly;
- cross-tenant tests fail if any content, metadata or existence signal leaks;
- audit events contain the policy decision but not the document content.
This turns governance from a slide-deck claim into a feature that can be implemented, tested and operated.
Conclusion
The best AI architecture does not try to make a probabilistic model behave like a perfectly trusted database or policy engine. It gives each component the responsibility it can safely carry.
- Models generate, classify and propose.
- Retrieval supplies authorised evidence.
- Deterministic services enforce identity, policy, limits and transactions.
- Humans provide judgement, accountability and redress at the right risk points.
- Management systems make ownership, evidence, audit and improvement repeatable.
EU AI Act, GDPR, ISO/IEC 42001 and ISO/IEC 27001 requirements should therefore not be separate workstreams added after the technical design. They become architecture inputs: tenant isolation, data minimisation, provenance, versioning, meaningful oversight, traceability, supplier controls, evaluation gates and safe failure.
That is the difference between an impressive demonstration and an enterprise AI capability that can be trusted, operated and defended in the real world.
Primary sources and implementation references
- European Union, Consolidated Regulation (EU) 2024/1689, Artificial Intelligence Act, updated 27 July 2026
- European Union, Regulation (EU) 2026/1744, Digital Omnibus on AI
- European Commission, AI Act regulatory framework and implementation timeline
- European Commission, Guidelines on AI Act Article 50 transparency obligations
- European Union, General Data Protection Regulation, Regulation (EU) 2016/679
- European Data Protection Board, Guidelines on automated individual decision-making and profiling
- European Data Protection Board, Guidelines 4/2019 on data protection by design and by default
- ISO, ISO/IEC 42001:2023, AI management systems
- ISO, ISO/IEC 27001:2022, information-security management systems
- ISO, ISO/IEC 27701:2025, privacy information management systems
- ISO, ISO/IEC 23894:2023, AI risk management
- ISO, ISO/IEC 42005:2025, AI system impact assessment
- NIST, AI Risk Management Framework and Playbook
- NIST, Generative AI Profile, NIST AI 600-1
- NIST, SP 800-218 Secure Software Development Framework
- NIST, SP 800-218A Secure Software Development Practices for Generative AI
- NIST, SP 800-207 Zero Trust Architecture
- OWASP, GenAI Security Project
- OWASP, Application Security Verification Standard 5.0
- OWASP, Software Component Verification Standard
- AWS, Architecting generative AI applications for production
- AWS, Well-Architected Agentic AI Lens
- AWS, Amazon Bedrock AgentCore Gateway
- Microsoft, Microsoft Foundry architecture
- Microsoft, AI gateway capabilities in Azure API Management
- Google Cloud, Choose agentic AI architecture components
- Google Cloud, RAG infrastructure and architecture
- Google Cloud, Gemini Enterprise Agent Platform overview
Discussion
Comments
Share feedback or questions about this page. No account required.
Loading comments…