Skip to main content

Architecting Production-Grade AI on AWS: Security, Scale, Governance and Compliance by Design

· 56 min read
AI Playbook author

Security, scale, governance and compliance by design

An impressive AI demonstration can be built in days. A production AI system that is safe, lawful, observable, resilient, affordable and trusted by users is a different engineering problem.

The model is only one component. The complete system also includes identity, authorization, data pipelines, retrieval, orchestration, tools, human approval, policy enforcement, evaluation, audit evidence, incident response and organisational governance. If any of those layers is weak, a highly capable model can make the overall solution less reliable rather than more valuable.

This article presents a detailed AWS reference architecture for a multi-tenant enterprise AI platform that supports conversational AI, retrieval-augmented generation, deterministic workflows and bounded agentic actions. It is designed around:

  • the EU AI Act;
  • the General Data Protection Regulation (GDPR);
  • ISO/IEC 42001:2023 for AI management systems;
  • ISO/IEC 27001:2022 for information security management systems;
  • ISO/IEC 23894:2023 for AI risk management;
  • the NIST AI Risk Management Framework and its Generative AI Profile;
  • the OWASP Top 10 for LLM and GenAI applications;
  • the AWS Well-Architected Framework, its Generative AI Lens and its Agentic AI Lens.

The objective is not to claim that an AWS service creates compliance automatically. It does not. AWS describes security and compliance as a shared responsibility: AWS secures the infrastructure of the cloud, while the customer remains responsible for the design, configuration, data, identities, applications and controls it operates in the cloud. Likewise, ISO/IEC 42001 does not replace law, and certification is performed by an independent certification body rather than by ISO. AWS Shared Responsibility Model, ISO explanation of ISO/IEC 42001

Important: This is technical and governance guidance, not legal advice. Determine the laws, regulatory guidance, sector rules, contractual requirements and AWS service terms that apply to the specific organisation, countries, data, users and intended purpose. Involve legal counsel, the data protection officer, information security, risk, compliance, product owners and affected stakeholders.

1. What “production-grade AI” really means​

A production-grade AI system should satisfy nine properties at the same time.

PropertyProduction expectationTypical evidence
UsefulIt improves a defined business or user outcomeBaseline, target KPI, adoption, completion rate, benefit realisation
SafeForeseeable harms and misuse are controlledHazard log, abuse cases, red-team results, residual-risk acceptance
LawfulProcessing has a valid basis and the AI Act role and risk class are knownLegal assessment, AI classification, DPIA, records of processing
SecureIdentities, data, models, tools and supply chain are protectedThreat model, access reviews, test results, security findings
ReliableIt degrades safely and meets agreed service objectivesSLOs, load tests, failover tests, runbooks, error budgets
Accurate enoughQuality is measured for the intended purpose and populationGolden dataset, retrieval and generation metrics, subgroup tests
ControllableHumans can understand, challenge, stop or reverse consequential actionsApproval records, override path, kill switch, action ledger
ObservableA decision or action can be reconstructed without indiscriminate data collectionCorrelation IDs, policy decisions, model and prompt versions, audit logs
Economically sustainableCost, latency and environmental impact are actively managedUnit economics, token budgets, capacity plan, model-routing policy

These properties are system properties. “The foundation model is secure” or “the model scored well on a benchmark” is not sufficient evidence for the application that surrounds it.

2. Start with the decision, not the model​

The first architecture document should be a use-case contract, not a model comparison spreadsheet.

Define the following before selecting Amazon Bedrock, Amazon SageMaker AI, an agent framework or a vector database:

  1. User and problem: Who experiences the problem, and what task should improve?
  2. Intended purpose: What exactly is the system designed to do, for whom, in what context and with what exclusions?
  3. Decision impact: Does it merely draft or retrieve information, or influence employment, credit, education, health, safety, legal rights or access to essential services?
  4. Automation boundary: What may the system recommend, draft, decide, execute or never do?
  5. Data contract: Which personal, special-category, confidential, copyrighted, regulated or residency-bound data can enter each component?
  6. Quality threshold: What accuracy, groundedness, coverage, fairness and abstention thresholds are necessary?
  7. Human control: Who reviews exceptions and consequential actions? Can that person understand, reject, reverse and stop the system?
  8. Operational target: What are the availability, latency, recovery, capacity and support objectives?
  9. Economic target: What is the maximum cost per completed task and the expected benefit?
  10. Exit condition: What evidence would cause the use case to be paused, redesigned or retired?

A running example​

The reference architecture will use Atlas, a fictional multi-tenant enterprise knowledge and workflow assistant. Atlas:

  • answers questions from approved corporate documents with citations;
  • summarises cases and drafts responses;
  • retrieves only documents the current user is authorised to see;
  • can create a service ticket after explicit confirmation;
  • requires human approval before changing a customer record or sending an external communication;
  • never makes final decisions on recruitment, credit, education admission, medical treatment or legal entitlement;
  • serves EU users from EU AWS Regions and supports tenant-specific retention and encryption policies.

Atlas is not automatically a high-risk AI system. Classification depends on its intended purpose and use. If the same platform is configured to rank job applicants, score examinations or decide access to credit, the legal and control profile changes substantially.

3. Classify the AI system before building it​

The EU AI Act uses a risk-based approach. As of 21 August 2026, the European Commission describes four levels: unacceptable risk, high risk, transparency risk and minimal or no risk. It also distinguishes roles including provider, deployer, importer, distributor and provider of a general-purpose AI model. Your organisation can have different roles for different components. European Commission AI Act overview

3.1 A practical classification gate​

Use a mandatory architecture gate with the following questions:

GateQuestionRequired outcome
ScopeIs the component an AI system or a GPAI model under the applicable definitions?Documented rationale
TerritoryIs it placed on the EU market, put into service in the EU, or are its outputs used in the EU?Jurisdiction map
RoleAre we provider, deployer, importer, distributor, product manufacturer or GPAI provider?Role and responsibility matrix
ProhibitionDoes intended use or foreseeable misuse match a prohibited practice?Reject or obtain authoritative legal determination before proceeding
High riskIs it a safety component under Annex I or an Annex III use case, considering applicable exceptions?High-risk control profile and conformity plan
TransparencyIs it a chatbot, synthetic-content system, emotion-recognition system, biometric-categorisation system or deepfake/public-interest text use case covered by transparency rules?Disclosure and marking controls
GPAIAre we merely integrating a third-party GPAI model, substantially modifying one, or placing our own GPAI model on the market?GPAI responsibility assessment
Personal dataIs personal data processed anywhere, including prompts, embeddings, telemetry, feedback or evaluation datasets?Lawful-basis assessment, ROPA entry and DPIA screening
Significant decisionIs a person subject to solely automated processing with legal or similarly significant effects?GDPR Article 22 assessment and human safeguards
SectorDo financial services, health, employment, consumer, accessibility, product-safety, NIS2, DORA or other rules apply?Sector-control overlay

The Commission states that transparency requirements for relevant systems came into effect in August 2026, while the amended dates for Annex III high-risk systems and high-risk systems embedded in regulated products are 2 December 2027 and 2 August 2028 respectively. Earlier obligations, including many prohibited practices and AI literacy, already apply. Treat the date as a planning deadline, not as permission to postpone good engineering. EU AI Act application timeline, EU AI Act enforcement framework

3.2 Classify configurations, not only platforms​

A reusable platform can host use cases with different risk profiles. Maintain an AI system inventory where each deployed configuration has its own:

  • system ID, owner and business sponsor;
  • intended purpose and prohibited uses;
  • provider and deployer roles;
  • risk classification and rationale;
  • affected people and fundamental rights;
  • model, prompt, guardrail, knowledge base and tool versions;
  • data categories, lawful basis, residency and retention;
  • external suppliers and subprocessors;
  • human-oversight design;
  • evaluation thresholds and latest results;
  • approvals, incidents, corrective actions and current lifecycle state.

Do not label the entire platform “low risk” and allow teams to configure high-impact workflows underneath it.

4. Architecture principles that prevent expensive failures​

4.1 Separate the control plane from the data plane​

The data plane handles user requests, retrieval, model inference and tool execution. The control plane governs which models, prompts, policies, datasets, tools and versions may be promoted. Production workloads must not modify their own guardrails or approval policies.

4.2 Treat the model as an untrusted probabilistic component​

The model may hallucinate, follow malicious instructions, expose context, choose the wrong tool or produce malformed output. Therefore:

  • authenticate and authorise outside the model;
  • never treat natural-language model output as a security decision;
  • validate structured output against a strict schema;
  • bind every data query to server-derived tenant and user permissions;
  • allowlist tools and arguments;
  • require deterministic policy checks before every action;
  • make high-impact or irreversible actions require human approval;
  • fail closed when identity, policy or safety controls are unavailable.

4.3 Use bounded autonomy​

AWS’s Agentic AI Lens highlights bounded autonomy, traceability, human oversight and goal-alignment evaluation because agents can perform multiple model calls, retain memory and act on external systems. Autonomy should therefore be earned per action, not granted to an entire agent. AWS Well-Architected Agentic AI Lens

Use four action tiers:

TierExampleControl
0: ReadSearch approved documentsAutomatic, still subject to authorisation and logging
1: DraftDraft an email or case noteUser reviews before use
2: Reversible writeCreate a draft ticket or add a non-final noteExplicit confirmation, idempotency and rollback
3: Consequential or irreversibleSend externally, change entitlement, transfer funds, delete recordsNamed human approval; often two-person control; strong authentication; policy and transaction limits

4.4 Make compliance executable​

A policy in a document is useful; the same policy expressed as a preventive control is stronger. Examples include:

  • AWS Organizations service control policies blocking unapproved Regions;
  • CI/CD rules preventing deployment without current evaluation evidence;
  • Cedar policies denying tools outside a user’s authority;
  • S3 bucket policies denying unencrypted uploads;
  • AWS Config rules detecting public storage or unapproved configurations;
  • a model gateway rejecting unapproved model, prompt or guardrail versions;
  • EventBridge rules opening a remediation workflow when a control drifts.

4.5 Preserve evidence as a product feature​

For every consequential response or action, be able to reconstruct:

  • who requested it and on whose behalf;
  • tenant, purpose and authorisation decision;
  • input classification and redaction events;
  • retrieved document IDs, versions and access filters;
  • model, inference profile, prompt and guardrail versions;
  • tool requests, validation, policy decisions and approvals;
  • output checks and final disposition;
  • latency, tokens, cost and errors;
  • whether the user accepted, corrected, appealed or overrode the result.

Store enough evidence to demonstrate control, but not unnecessary raw personal data. Traceability and data minimisation must be designed together.

5. End-to-end AWS reference architecture​

The following design is deliberately modular. Managed services can reduce operational burden, while the policy boundaries remain explicit and portable.

5.1 AWS multi-account foundation​

Use AWS Organizations and AWS Control Tower to create separation of duties and reduce blast radius. AWS’s Security Reference Architecture recommends a multi-account approach and central security capabilities, tailored to the organisation’s risks. AWS Security Reference Architecture

Recommended accounts and organisational units:

Account or OUPurposeImportant controls
ManagementOrganisation-level administration onlyNo workloads; protected root; tightly restricted access
Log ArchiveCentral, immutable security and audit logsSeparate KMS keys; S3 Object Lock where appropriate; restricted deletion
Security ToolingDelegated administration and responseSecurity Hub CSPM, GuardDuty, Macie, Inspector, Detective, Config, Firewall Manager
Identity and Shared ServicesIAM Identity Center, directory, DNS, shared network servicesFederated access, MFA, short-lived roles, central egress controls
AI GovernanceAI inventory, evaluation evidence, model/prompt approval metadataIndependence from workload deployers; controlled promotion workflow
Data PlatformCurated source data, catalogues and governed ingestionLake Formation, Glue, Macie, KMS, data-owner approvals
Non-production WorkloadsDevelopment, test and evaluationSynthetic or masked data by default; no route to production data
Production WorkloadsRuntime APIs, retrieval and approved modelsDedicated VPCs, private endpoints, least privilege, no interactive changes
SandboxTime-limited experimentationBudget limits, blocked production data, restricted external connectivity

Apply service control policies to deny disabled Regions, public S3 access, unauthorised Bedrock models, unapproved cross-Region inference profiles, long-lived IAM users and departure from approved encryption controls. SCPs provide permission boundaries; they do not grant permissions and should be tested carefully to avoid operational lockout.

5.2 Network and edge layer​

  1. Use Route 53 for DNS and health-aware routing.
  2. Place CloudFront in front of public web applications where appropriate.
  3. Use AWS WAF managed rules plus application-specific rules for injection, bots, abusive automation, payload size and rate limits.
  4. Use AWS Shield Standard by default and consider Shield Advanced for workloads with material DDoS risk and support requirements.
  5. Terminate TLS with AWS Certificate Manager.
  6. Use API Gateway for managed APIs, quotas and authorisers, or an Application Load Balancer for containerised streaming workloads.
  7. Keep application compute, data stores and orchestration in private subnets across at least two Availability Zones.
  8. Use VPC endpoints and AWS PrivateLink for supported AWS services, including Bedrock runtime endpoints, rather than routing service traffic through the public internet. PrivateLink resolves supported service endpoints to private addresses in the VPC. AWS PrivateLink documentation
  9. Centralise and allowlist outbound internet access. An agent must not receive general internet access merely because one tool needs a specific external API.
  10. Use separate security groups and roles for the API, orchestrator, retrieval service and each tool class.

5.3 Identity, tenant isolation and authorisation​

Use Amazon Cognito or an enterprise identity provider for customer identity and IAM Identity Center for workforce access. Enforce MFA or phishing-resistant authentication based on risk. Use short-lived credentials and workload roles rather than embedded access keys.

Authentication is only the first step. Implement authorisation at four boundaries:

  1. API: Can this identity call this route for this tenant?
  2. Application: Can the user perform this business action on this resource?
  3. Retrieval: Can the user see this document, chunk and metadata field?
  4. Tool: May the agent invoke this exact capability with these arguments, value limits and approval state?

Amazon Verified Permissions can externalise fine-grained application authorisation using Cedar. For managed agent tools, Amazon Bedrock AgentCore Policy can evaluate deterministic Cedar policies at the gateway boundary and uses default-deny semantics. Amazon Verified Permissions, AgentCore Policy

For every request, derive tenant_id, subject_id, roles and entitlements from a verified token and authoritative directory. Never accept a tenant ID from the prompt or trust a model-generated filter. Propagate the trusted context through signed internal claims and enforce it again at every downstream resource.

Isolation choices should match risk:

PatternAppropriate useTrade-off
Shared service and shared index with tenant filtersLower-sensitivity SaaS at high scaleLowest cost, but filter defects can cause cross-tenant exposure
Shared service with per-tenant index or collectionModerate/high sensitivityBetter isolation with operational overhead
Per-tenant account or deployment cellHighly regulated or very large tenantsStrongest boundary and clearer residency, higher cost and management complexity

For high-impact or confidential deployments, prefer structural isolation over relying only on metadata filters.

5.4 AI application gateway​

The gateway is a security and governance enforcement point, not a thin proxy. Run it on Lambda for short stateless requests, or ECS on Fargate/EKS for long-running, streaming or framework-heavy workloads.

Its responsibilities should include:

  • request size and content-type validation;
  • session and tenant resolution;
  • purpose and consent checks where applicable;
  • rate, concurrency, token and monetary budgets;
  • input PII detection, masking or tokenisation;
  • prompt-injection and abuse screening;
  • approved model, prompt and guardrail selection;
  • routing to RAG, deterministic workflow or agent mode;
  • correlation ID and evidence event generation;
  • output schema and policy validation;
  • safe error handling without leaking prompts, policies or internal topology.

Do not allow application teams to call foundation models directly from browsers or arbitrary services. Route inference through a controlled model gateway with model allowlists and telemetry.

5.5 Workflow and agent orchestration​

Use deterministic code for deterministic work. Step Functions is well suited to explicit, auditable state transitions, retries, timeouts, approvals and compensation. Use an LLM only where semantic interpretation or generation adds value.

Three viable orchestration patterns are:

PatternBest fitRecommendation
Step Functions plus Lambda/ECSRegulated workflows and predictable state machinesDefault for consequential business processes
Custom agent runtime on ECS/EKSMaximum framework control and portabilityUse when runtime behaviour or networking needs customisation
Amazon Bedrock AgentCoreManaged runtime, identity, gateway, memory, observability, evaluation and policyUse after verifying Region, data-flow, feature and control requirements

Regardless of runtime, enforce:

  • a maximum number of reasoning steps and tool calls;
  • wall-clock, token and financial budgets;
  • per-step timeouts and circuit breakers;
  • typed state rather than uncontrolled conversation text;
  • explicit success, abstain, escalate and fail-safe terminal states;
  • no recursive self-delegation without a bounded plan;
  • no dynamic tool installation in production;
  • memory retention rules and user-visible controls;
  • idempotency keys and compensating actions for writes.

5.6 Retrieval-augmented generation layer​

RAG can improve relevance and provide citations, but it does not eliminate hallucination or prompt injection. OWASP notes that RAG and fine-tuning do not fully mitigate prompt injection. Treat retrieved content as untrusted data, not higher-priority instructions. OWASP LLM01: Prompt Injection

Use one of two approaches:

  • Amazon Bedrock Knowledge Bases for a managed retrieval and generation path;
  • a custom retrieval service using S3, OpenSearch Serverless or Aurora PostgreSQL with pgvector when fine-grained control, hybrid retrieval, complex permissions or portability is required.

A strong seven-stage retrieval pipeline is:

  1. Query understanding and safe rewriting.
  2. Mandatory tenant, resource and entitlement filters.
  3. Hybrid lexical and vector candidate retrieval.
  4. Metadata, freshness and policy filtering.
  5. Reranking using an approved model.
  6. Context assembly with source boundaries and token budget.
  7. Answer generation with inline citations, confidence/abstention policy and post-generation verification.

Measure retrieval separately from generation. Useful metrics include recall@k, precision@k, mean reciprocal rank, nDCG, permission-filter accuracy, citation correctness, groundedness, completeness, answer relevance and abstention accuracy. Amazon Bedrock supports RAG evaluation for Knowledge Bases and external RAG sources; use automated evaluation for scale and human expert review for material decisions. Amazon Bedrock RAG evaluation

5.7 Foundation-model layer​

Amazon Bedrock provides managed access to multiple foundation-model providers, while SageMaker AI is appropriate when an organisation needs to train, customise or host models with deeper control. Keep the calling application independent of provider-specific payloads through an internal model interface.

The model gateway should select models using an approved routing policy based on:

  • task type and quality threshold;
  • data classification and permitted Regions;
  • model risk and supplier approval;
  • latency and context requirements;
  • cost ceiling;
  • availability and quota state;
  • language and accessibility requirements.

Use the smallest model that consistently passes the task-specific evaluation. A cheaper model that causes more rework or unsafe escalations is not cheaper in business terms.

5.7.1 Prompting, RAG, fine-tuning or custom training?​

Choose the least complex method that meets the requirement:

NeedPreferred starting pointReason
Change instructions, tone or output formatVersioned prompting and structured outputFastest, cheapest and easiest to reverse
Add current or private knowledge with citationsRAGKnowledge can change without retraining and remains traceable to sources
Make behaviour more consistent for a repeated taskFine-tuning, after prompting and RAG have been measuredCan improve task behaviour, but adds dataset, evaluation and lifecycle risk
Solve a specialised predictive or classification problemSageMaker AI training or an approved pre-trained modelSupports controlled feature, training, registry and deployment workflows
Meet strict model-control or hosting requirementsApproved model imported or hosted using SageMaker AIGreater control with greater security and operational responsibility

For custom or fine-tuned models, create a separate MLOps path:

Separate training, validation and test data; record lineage and licences; scan artefacts; pin container digests; make experiments reproducible; test subgroup performance; and require independent approval before promotion. SageMaker Model Registry can catalogue versions, metadata, lineage and approval status, while Model Cards can preserve intended use, limitations, risk, training and evaluation information. SageMaker Model Registry, SageMaker Model Cards

Monitor custom models using task-appropriate scheduled evaluation, data-quality jobs, application outcomes and CloudWatch telemetry. Do not assume that statistical drift always means harm, or that the absence of drift means the system remains safe. A model can fail because the business process, population, policy, upstream system or meaning of the target changed.

Apply Amazon Bedrock Guardrails to both relevant inputs and outputs for content categories, denied topics, sensitive information, prompt attacks and grounding checks. Guardrails are one layer and must be tested for the actual languages and modalities in scope. AWS documentation notes, for example, that sensitive-information filters do not detect PII inside tool-use parameters, so tool arguments require separate schema and DLP validation. Amazon Bedrock Guardrails, Bedrock sensitive-information filters

5.8 Tool broker and human approval​

The agent should not hold broad credentials to enterprise systems. Put a tool broker between the orchestrator and every external action.

Each tool definition must include:

  • immutable tool ID and version;
  • business owner and technical owner;
  • typed input and output schema;
  • allowed callers and resources;
  • data classification and destinations;
  • maximum transaction or record scope;
  • approval tier;
  • idempotency behaviour;
  • timeout, retry and compensation policy;
  • audit fields;
  • abuse cases and tests;
  • expiry and recertification date.

Before execution, the broker should:

  1. validate the identity and tenant context;
  2. parse arguments against a strict schema;
  3. reject hidden, unexpected or overlong fields;
  4. run deterministic authorisation and business rules;
  5. classify data leaving the boundary;
  6. obtain explicit user confirmation or human approval where required;
  7. exchange for a narrow, short-lived credential;
  8. execute with an idempotency key;
  9. verify the result and record a tamper-evident action event;
  10. trigger compensation or escalation if the outcome is uncertain.

Never let the model construct raw SQL, shell commands, IAM policies or arbitrary HTTP destinations for direct execution. Use parameterised operations and allowlisted resources.

5.9 Data stores​

Use data stores according to access pattern, not fashion:

DataAWS optionControl notes
Source documentsS3Versioning, KMS, bucket-owner enforcement, public-access block, lifecycle and Object Lock where justified
Relational metadata and business transactionsAurora PostgreSQLMulti-AZ, TLS, RDS Proxy, row-level controls, backups and deletion workflow
Session and workflow stateDynamoDBTTL, conditional writes, point-in-time recovery, tenant-key design
Cache and rate stateElastiCache for Redis/ValkeyEncryption, authentication, private subnets; do not make cache the system of record
Vectors and hybrid searchOpenSearch Serverless or Aurora pgvectorDocument-level ACLs, tenant isolation, index lifecycle and deletion propagation
SecretsSecrets ManagerRotation and workload-role access; never store secrets in prompts
ConfigurationAppConfig and Parameter StoreVersion, validate and progressively deploy safety-critical configuration
Evidence and auditCentral S3/CloudWatch plus governance storeSeparate raw security logs from privacy-minimised AI decision records

Encrypt data in transit and at rest. Use separate customer-managed KMS keys where separation, revocation, customer control or audit needs justify it. Key policies, grants, rotation, recovery and separation of key administrators from data users are part of the control.

6. Secure document-ingestion architecture​

Documents are a major indirect-prompt-injection and data-leakage path. Ingestion must be a controlled supply chain.

Ingestion steps​

  1. Source registration: Record owner, authority, lawful basis, licence, classification, geography, refresh and deletion policy.
  2. Quarantine: Upload to a non-public, non-searchable S3 prefix with an object-level correlation ID.
  3. Malware scan: GuardDuty Malware Protection for S3 can scan newly uploaded objects; quarantine remains closed until the result is clean. GuardDuty Malware Protection for S3
  4. File validation: Verify actual type, extension, size, archive depth, macros, embedded objects and decompression ratio.
  5. Sensitive-data discovery: Use Macie for S3 estate discovery plus purpose-specific detection, because generic PII detection is not a substitute for a business data inventory. Amazon Macie automated discovery
  6. Parsing: Run Textract or isolated parsers with resource and time limits. Treat extracted text and metadata as untrusted.
  7. Content safety: Detect hidden instructions, poisoned passages, malicious links, executable content and attempts to override system policy. Tag risk; do not assume perfect classification.
  8. Normalisation and redaction: Remove prohibited fields, tokenise identifiers where required and preserve a secure mapping only when necessary.
  9. Chunking: Use structure-aware chunks with document ID, version, page or section, effective date and security labels.
  10. Embedding: Generate embeddings only after access labels and residency policy are attached.
  11. Indexing: Use a staging index, enforce tenant and document ACLs, and make publication an explicit transaction.
  12. Evaluation: Test retrieval, permissions, citations, poisoning resistance and deletion propagation.
  13. Approval: Data owner or delegated steward approves promotion.
  14. Publication: Atomically switch an index alias or version pointer.
  15. Change and deletion: Re-ingest on source change; tombstone and delete chunks, vectors, caches and derived evaluation copies when the source is withdrawn or a valid erasure request applies.

Maintain lineage from every chunk and embedding back to the exact source object, version and authority. Without lineage, correction, deletion and evidence production become slow and unreliable.

7. Runtime request and action flow​

For each user request, the recommended sequence is:

  1. Authenticate: Validate issuer, audience, signature, expiry, authentication strength and session risk.
  2. Resolve context: Load tenant, roles, entitlements, locale, purpose and applicable policy profile from trusted systems.
  3. Authorise: Check route and resource access. Deny by default.
  4. Validate: Enforce payload type, length, file rules and rate limits.
  5. Classify and minimise: Detect sensitive data, remove fields not needed for the task and replace identifiers with reversible tokens only when necessary.
  6. Screen: Apply abuse, prompt-attack and prohibited-use controls.
  7. Plan: Select direct answer, RAG, deterministic workflow or bounded agent. The model may propose a plan; deterministic code approves executable steps.
  8. Retrieve: Apply server-side entitlements before similarity search and again before context assembly.
  9. Invoke: Use an approved model, prompt and guardrail version through a private endpoint.
  10. Validate output: Parse structured output, verify citations, DLP-scan, enforce business rules and abstain when evidence is insufficient.
  11. Approve action: Present the exact proposed action, material inputs, recipient and consequences to the authorised human.
  12. Execute: Tool broker checks authorisation again, uses short-lived credentials and records the result.
  13. Respond: Label the AI interaction where required, show sources and limitations, and provide correction, escalation or appeal routes.
  14. Record evidence: Write privacy-minimised trace and audit events with model, prompt, retrieval and policy versions.
  15. Learn safely: Send sampled or user-corrected interactions to a governed evaluation queue. Never automatically train on all production conversations.

8. EU AI Act controls translated into architecture​

For high-risk AI systems, Articles 9–15 of the AI Act cover lifecycle risk management, data governance, technical documentation, automatic logging, information to deployers, human oversight, and accuracy, robustness and cybersecurity. Article 26 establishes important deployer obligations, including competent human oversight, monitoring and log retention under the deployer’s control. Official text of Regulation (EU) 2024/1689

AI Act areaArchitecture and process controlEvidence
Risk management, Article 9Continuous hazard, misuse and fundamental-rights assessment; control testing; residual-risk approval; post-market feedbackAI risk register, test plan, red-team report, approval and review history
Data governance, Article 10Dataset provenance, suitability, representativeness, error and bias checks; geographic/context validation; protected-data controlsDataset card, lineage, quality report, bias analysis, data-owner approval
Technical documentation, Article 11 and Annex IVVersioned system description, architecture, intended purpose, metrics, dependencies, changes and limitationsGenerated technical file tied to release ID
Record keeping, Article 12Automatic event logging, model/prompt/tool versions, relevant inputs and outputs, oversight and incidentsImmutable or tamper-evident audit records and retention policy
Transparency to deployers, Article 13Instructions for use, capability and limitation statements, required inputs, expected accuracy, oversight and maintenanceDeployer guide, model/system card, release notes, training record
Human oversight, Article 14Trained and authorised reviewers; reject/override/reverse/stop controls; automation-bias trainingApproval logs, competency record, override and kill-switch tests
Accuracy, robustness and cybersecurity, Article 15Task thresholds, adversarial and resilience tests, redundancy, safe fallback, poisoning and prompt-injection controlsEvaluation reports, SLO results, threat model, penetration and red-team reports
Provider obligationsQuality management, conformity workflow, registration where required, post-market monitoring and serious-incident handlingQMS procedures, declaration and assessment records, monitoring plan
Deployer obligations, Article 26Use within instructions, qualified oversight, relevant input data, monitoring, suspension and incident escalationOperating procedure, training, monitoring dashboard and incident record
Transparency, Article 50AI interaction disclosure and synthetic-content marking/labels when applicableUI test, content metadata, disclosure copy and monitoring

Practical design rule​

Do not build an “AI Act dashboard” that is disconnected from delivery. Link every obligation to:

  • a control owner;
  • an implementation mechanism;
  • a test;
  • a threshold;
  • an evidence source;
  • a review frequency;
  • a remediation SLA;
  • a release or runtime blocking rule.

9. ISO/IEC 42001 as the organisational operating system​

ISO/IEC 42001:2023 is an AI management system standard. ISO explains that it uses Plan–Do–Check–Act and covers governance of AI-related risks and opportunities across the organisation. It is not only a technical checklist for one application. ISO/IEC 42001 overview

Plan​

  • Define the AIMS scope, interested parties, internal and external context.
  • Approve an AI policy and measurable objectives.
  • Establish an inventory and classification method.
  • Assign accountability, competence and escalation routes.
  • Assess AI risks and impacts, including affected groups and foreseeable misuse.
  • Define a statement of applicability for chosen controls and exclusions.
  • Set supplier, data, lifecycle, transparency and incident requirements.

Do​

  • Operate approved data, development, evaluation and release processes.
  • Train developers, reviewers, operators and business users.
  • Apply lifecycle controls to models, prompts, retrieval, tools and configurations.
  • Maintain communication and user transparency.
  • Control documents, records, suppliers and operational changes.

Check​

  • Monitor AI quality, safety, fairness, security, complaints and business impact.
  • Run internal audits and independent assessments.
  • Review control exceptions and corrective-action ageing.
  • Conduct management reviews with evidence and resource decisions.

Act​

  • Contain incidents and nonconformities.
  • Identify root causes rather than merely tuning prompts.
  • Implement corrective and preventive actions.
  • Update risk assessments, policies, tests and training.
  • Retire systems whose residual risk or value is no longer acceptable.

ISO/IEC 42001 evidence pack​

The AI Governance account or integrated GRC platform should preserve:

  • AI policy and objectives;
  • AIMS scope and responsibility matrix;
  • AI system inventory;
  • risk and impact assessments;
  • data and model/system cards;
  • supplier assessments and contracts;
  • evaluation and red-team results;
  • human-oversight design and competence records;
  • approvals and release manifests;
  • monitoring, complaints, incidents and corrective actions;
  • internal audit and management-review records;
  • continual-improvement backlog.

AWS Audit Manager can automate collection of some AWS configuration evidence, while human and governance evidence must still be added from authoritative business processes. Automated collection is evidence support, not proof that the control is effective in context. AWS Audit Manager evidence collection

10. ISO/IEC 27001 integrated with AI security​

ISO/IEC 27001:2022 defines requirements for an information security management system and preserves confidentiality, integrity and availability through risk management. AI should normally sit inside the existing ISMS rather than creating a disconnected security programme. ISO/IEC 27001 overview

ISMS concernAI-specific interpretationAWS implementation examples
Asset managementModels, prompts, embeddings, datasets, tools, policies and evaluation sets are assetsResource tags, inventory, S3 catalogues, model/prompt registry
Access controlUsers and agents need least privilege and segregation of dutiesIdentity Center, IAM roles, permission boundaries, Verified Permissions, SCPs
CryptographyProtect data, evidence and model artefacts with controlled keysKMS, ACM, TLS, Secrets Manager, key separation
Secure developmentAI changes require threat modelling and evaluation, not only unit testsCodePipeline/CodeBuild, IaC tests, SBOM, ECR scanning, evaluation gates
Supplier securityFoundation-model, data and tool providers affect confidentiality and resilienceSupplier register, DPA, SCC/TIA where applicable, SLA and exit plan
Logging and monitoringTrace prompts, retrieval and tools while minimising personal dataCloudTrail, CloudWatch, OpenTelemetry, central S3 and privacy filters
Vulnerability managementInclude containers, libraries, parsers, agent tools and model supply chainInspector, ECR enhanced scanning, Dependabot or equivalent, signed artefacts
Incident managementAI incidents include harmful output, data leakage, poisoning, runaway actions and model regressionsSecurity Hub, EventBridge, Step Functions, runbooks and forensic evidence
ContinuityModel or Region failure must not force unsafe behaviourMulti-AZ, approved fallback model, queues, cached safe responses, DR tests
ComplianceLegal, regulatory and contractual obligations must be traceable to controlsAudit Manager, Config conformance packs, evidence repository

Maintain one integrated risk register where possible, with links between information-security, privacy, AI, operational and third-party risks. Duplicated registers often drift and obscure ownership.

11. GDPR: privacy engineering for AI systems​

GDPR applies to personal data in obvious places such as user profiles and in less obvious places such as free-text prompts, retrieved passages, embeddings, chat history, telemetry, human feedback, evaluation datasets and security logs.

11.1 Translate principles into defaults​

Article 5 requires lawfulness, fairness, transparency, purpose limitation, data minimisation, accuracy, storage limitation, integrity/confidentiality and accountability. Article 25 requires data protection by design and by default; Article 32 requires risk-appropriate security, resilience, restoration and regular testing. GDPR official text

GDPR principle or dutyEngineering control
Lawfulness, fairness and transparencyRecord purpose and lawful basis; provide layered notices; test disparate impacts; disclose AI interaction where applicable
Purpose limitationBind datasets and APIs to approved purpose tags; prohibit silent reuse of conversations for training
Data minimisationRemove unneeded fields before model calls; retrieve only the minimum passages; do not log raw content by default
AccuracyShow citations, support correction, avoid writing generated facts back to systems of record without verification
Storage limitationPer-data-class TTLs, S3 lifecycle, DynamoDB TTL, deletion jobs and backup-retention rules
Integrity and confidentialityEncryption, fine-grained access, private networking, DLP, tamper-evident logs and tested restoration
AccountabilityROPA, DPIA, decisions, control tests, supplier records and evidence linked to each system version
Data-subject rightsSearch, export, rectify, restrict and erase across source, session, vector, cache, feedback and derived stores
Automated decisionsDetect Article 22 scenarios; provide meaningful human intervention, contest and explanation where required
International transfersMap every processing location and subprocessor; assess Chapter V mechanism, SCCs and transfer risk as applicable

11.2 Build a privacy gateway​

Before a prompt reaches retrieval or inference:

  • determine whether the request is permitted for the declared purpose;
  • classify personal and confidential data;
  • remove unnecessary fields;
  • tokenise identifiers when re-identification is required later;
  • block prohibited special-category or secret data for that use case;
  • apply tenant-specific residency and retention policy;
  • record the decision without duplicating the sensitive content.

After inference, scan the response for personal data, secrets, unsupported identity claims and cross-tenant leakage. Do not rely on one probabilistic PII detector; combine managed detection, domain rules, exact identifiers, schemas and human review according to risk.

11.3 DPIA and records of processing​

GDPR Article 35 requires a DPIA before processing likely to create high risk to people’s rights and freedoms, particularly with new technologies. It must be a decision instrument, not a retrospective form. Begin during discovery, update after architecture and testing, and review when processing risk changes. Article 30 records should describe purposes, data subjects and categories, recipients, transfers, retention and security measures. GDPR Article 35

A useful AI DPIA covers:

  • data-flow diagram and controller/processor roles;
  • purposes and lawful bases for each processing operation;
  • necessity and proportionality;
  • people, vulnerable groups and expected scale;
  • model memorisation, inference and re-identification risks;
  • prompt and output leakage;
  • accuracy, bias, manipulation and automation-bias risks;
  • solely automated decision-making and human intervention;
  • transfers, subprocessors and remote support access;
  • retention, deletion and data-subject-rights execution;
  • technical and organisational measures;
  • residual risk, consultation and approval.

The EDPB’s Opinion 28/2024 emphasises that whether an AI model is anonymous requires a case-by-case assessment and addresses legitimate interest and unlawfully processed training data. Do not assume that removing names makes a model, embedding or dataset anonymous. EDPB Opinion 28/2024

11.4 Data-subject rights across derived stores​

Maintain a subject-to-artifact index where lawful and proportionate, so a verified request can locate:

  • user profile and account data;
  • conversations and attachments;
  • workflow state and memories;
  • source documents;
  • chunks and embeddings;
  • cached responses;
  • feedback and evaluation samples;
  • fine-tuning datasets and, where technically and legally relevant, model artefacts;
  • audit records subject to an applicable retention exception.

Deletion should be asynchronous but measurable: accept the request, freeze incompatible use, issue tombstones, delete from live stores and replicas, invalidate caches and indexes, handle backups under the documented restoration/deletion policy, then record completion without retaining the erased content.

11.5 AWS Regions and international transfers​

Choose an EU home Region based on service availability, latency, resilience and contractual requirements. Region choice alone does not complete a GDPR transfer assessment. Map model inference, guardrail processing, support, logging, backups, analytics, third parties and cross-Region features.

Bedrock geographic cross-Region inference can keep processing within a named geography such as the EU, but prompts and outputs may move outside the source Region to another Region within that geography, and some abuse-detection data may be stored in the destination Region. Global profiles may process outside the geography. Use geography-bound or single-Region patterns only after reviewing the current service documentation and organisational transfer policy; enforce the approved profile in IAM and SCPs. Bedrock geographic cross-Region inference

AWS provides a Data Processing Addendum and Standard Contractual Clauses for relevant processing, but the customer must still understand its roles, instructions, service choices and transfer obligations. AWS Data Processing Addendum

12. Threat model: traditional and AI-specific attacks​

Use a combined method: STRIDE for application and cloud threats, privacy threat modelling such as LINDDUN where useful, business abuse cases, OWASP GenAI risks and adversarial ML techniques. NIST’s Generative AI Profile identifies risks including confabulation, data privacy, information integrity, information security and intellectual property, and organises risk work around Govern, Map, Measure and Manage. NIST AI 600-1

ThreatExamplePreventDetect and respond
Direct prompt injectionUser asks the model to reveal system instructionsInstruction hierarchy, guardrails, task-specific prompts, least privilegeAttack scores, canary prompts, red-team regression suite
Indirect prompt injectionA retrieved PDF tells the agent to exfiltrate recordsTreat content as data, source isolation, tool policy outside modelRetrieval provenance, suspicious-instruction detection, blocked tool trace
Sensitive disclosureModel returns another tenant’s data or a secretStructural tenant isolation, DLP, minimum context, no secrets in promptsCross-tenant probes, output scans, Macie and incident workflow
Improper output handlingGenerated HTML/SQL causes injectionSchema validation, encoding, parameterised operations, sandboxingApplication security tests, WAF and runtime alerts
Supply-chain compromisePoisoned library, model, container or parserApproved registry, SBOM, signatures, pinning and supplier reviewInspector/ECR findings, provenance checks, emergency revocation
Data or model poisoningMalicious content influences future answersTrusted sources, quarantine, lineage, staging index, owner approvalDistribution and retrieval drift, source reputation, canary queries
Excessive agencyAgent sends or deletes without real approvalTool broker, Cedar/IAM policy, action tiers, human approvalAction anomaly detection, transaction limits, kill switch
System-prompt leakageUser extracts policies or hidden dataNo secrets in prompts, compartmentalisation, refusal and output filtersCanary tokens and leakage evaluation
Vector weaknessesInsecure filters or malicious nearest neighboursServer-side ACLs, index isolation, signed metadata, hybrid validationPermission-recall tests, cross-tenant test corpus
Misinformation and overrelianceFluent but incorrect advice is acceptedCitations, calibrated abstention, reviewer UI, limitationsGroundedness, correction and override rates, expert sampling
Unbounded consumptionLong loops cause denial of wallet/serviceQuotas, max tokens/steps, cost budgets, concurrency controlsToken and cost anomaly alarms, automatic circuit breaker
Model denial or dependency failureProvider throttles or model is unavailableCapacity planning, queueing, fallback and degraded modeQuota, error and latency alarms; failover runbook
Memory poisoningMalicious or incorrect fact persists across sessionsExplicit memory schema, provenance, user controls, expiryMemory-change log, anomalous write review, purge process
Tool result injectionExternal API returns adversarial textTyped results, data/instruction separation, trust labelsTool-response scanning and provenance trace

Security testing should include prompt injection, jailbreaks, multilingual and encoded attacks, sensitive-data extraction, cross-tenant retrieval, privilege escalation, tool argument manipulation, poisoned documents, denial of wallet, context overflow, model fallback, malformed structured output and approval bypass.

13. Guardrails as a layered safety architecture​

Use guardrails at four points:

  1. Before retrieval: Reject prohibited or clearly malicious requests and minimise sensitive data.
  2. After retrieval: Remove unauthorised, stale or unsafe context and isolate instructions found in documents.
  3. After generation: Check harm, PII, secrets, citations, groundedness, schema and business rules.
  4. Before action: Re-authorise, validate arguments, check value limits and obtain required approval.

Create separate, versioned guardrail profiles by use case and risk. A customer-service bot and an internal legal drafting assistant do not have identical denied topics, disclosure needs or escalation thresholds.

Guardrail evaluation should report:

  • true-positive and false-positive rates by category;
  • false-negative rate on adversarial sets;
  • performance by language and modality;
  • subgroup impacts and accessibility;
  • blocked legitimate task rate;
  • bypass success rate;
  • added latency and cost;
  • correct escalation and user-recovery rate.

Amazon Bedrock Guardrails provides configurable content, denied-topic, word, sensitive-information, prompt-attack, contextual-grounding and automated-reasoning capabilities. Automated Reasoning checks return findings rather than simply blocking, so the application must decide how to handle each result. Bedrock Guardrails, Automated Reasoning checks

14. Evaluation: the production release gate​

An AI release is the combination of application code, infrastructure, model, prompt, guardrail, retrieval configuration, corpus, tools and policy. A change to any of these can change behaviour.

14.1 Build an evaluation pyramid​

LevelPurposeExamples
Unit and contractDeterministic correctnessSchema, prompt-template variables, policy decisions, ACL filters, tool idempotency
ComponentAI component qualityRetrieval recall, reranker quality, classifier precision, guardrail errors
End-to-end offlineTask and safety qualityGroundedness, completeness, refusal, prompt injection, subgroup results
Human expertContextual acceptabilityCorrectness, usefulness, harm, explainability, reviewer effort
Shadow/canaryReal traffic without uncontrolled impactLatency, drift, disagreement, policy blocks, cost
Production monitoringContinual assuranceSLOs, complaints, incidents, override rate, business outcome

14.2 Golden datasets​

Maintain versioned datasets containing:

  • representative normal tasks;
  • difficult and ambiguous cases;
  • “unanswerable” cases where abstention is correct;
  • stale or conflicting documents;
  • all supported languages and user groups;
  • prohibited and out-of-scope requests;
  • direct and indirect prompt injections;
  • attempts to retrieve another tenant’s content;
  • malformed tool arguments;
  • historical incidents and near misses;
  • counterfactual pairs for fairness analysis.

Evaluation data must itself have a lawful basis, provenance, licence, retention policy and access control.

14.3 Release thresholds​

Example gates for Atlas, to be calibrated with domain experts:

  • permission-filter test: 100% pass on the security regression suite;
  • high-severity data-leak test: zero tolerated;
  • tool-authorisation bypass: zero tolerated;
  • citation correctness: ≥ 98% on the golden set;
  • grounded answer rate: ≥ 95% for answerable cases;
  • correct abstention: ≥ 90% for unanswerable cases;
  • harmful-response rate: below the approved risk threshold with confidence interval;
  • task completion: no material regression from the production baseline;
  • p95 latency and cost per task: within SLO and budget;
  • no subgroup below an approved minimum performance floor.

Average quality must not hide severe failures. Use hard security/safety gates, per-segment floors and statistical uncertainty.

Amazon Bedrock supports model evaluation, LLM-as-a-judge and human evaluation. LLM judges are scalable but should be calibrated against expert labels and should not be the sole judge of high-impact correctness or compliance. Amazon Bedrock evaluations, Human-based model evaluation

15. LLMOps, MLOps and secure delivery​

15.1 Version the whole AI bill of materials​

For each release, create a signed manifest containing:

  • source-code commit and build provenance;
  • infrastructure-as-code version;
  • container digest and SBOM;
  • foundation or custom model ID and version;
  • embedding and reranking models;
  • prompt templates and system instructions;
  • guardrail ID and version;
  • retrieval index and corpus snapshot;
  • tool schemas and policy bundle;
  • evaluation dataset and report IDs;
  • DPIA, risk and approval versions;
  • feature flags and rollout plan.

15.2 CI/CD stages​

  1. Developer pre-commit checks, secret scanning and unit tests.
  2. Static analysis, dependency and licence scan.
  3. Container build, SBOM, vulnerability scan and signing.
  4. IaC policy checks and CloudFormation/CDK/Terraform validation.
  5. Deploy to isolated test account.
  6. Contract, retrieval, guardrail, red-team and end-to-end evaluations.
  7. Privacy, security and AI-risk gates based on classification.
  8. Human approval by authorised roles with separation of duties.
  9. Progressive production deployment using canary or blue/green strategy.
  10. Automatic rollback or traffic stop on SLO, safety or security breach.
  11. Evidence bundle written to the governance repository.

No engineer should edit production prompts, guardrails or tool permissions directly in a console. Emergency change procedures still require identity, reason, time limit, evidence and post-change review.

16. Observability without creating a privacy liability​

16.1 Three telemetry planes​

Operational telemetry

  • request volume, latency, availability and errors;
  • token input/output, cache hit and cost;
  • retrieval latency and result count;
  • tool call count and failure;
  • queue depth, throttling and capacity.

AI quality and safety telemetry

  • groundedness and citation validation;
  • abstention, refusal and escalation;
  • guardrail category and policy outcome;
  • prompt-injection detection;
  • user correction, override and complaint;
  • subgroup and language performance;
  • drift from the approved baseline.

Audit telemetry

  • actor, tenant, purpose and authorisation;
  • model, prompt, policy, guardrail and corpus versions;
  • retrieved source identifiers;
  • tool proposal, approval and execution result;
  • release and evidence IDs.

Use CloudWatch metrics and logs, AWS X-Ray or OpenTelemetry traces, CloudTrail for AWS API activity, and centralised S3 for protected long-term evidence. Bedrock runtime publishes invocation volume, latency, token, error and logging-delivery metrics to CloudWatch. Bedrock runtime metrics

16.2 Logging rules​

  • Log metadata by default, content by exception.
  • Redact or tokenise before telemetry leaves the request boundary.
  • Use separate log groups and retention for operational, security and audit data.
  • Do not put secrets, credentials or unrestricted prompts into logs.
  • Restrict prompt/output logging to approved use cases and datasets.
  • Detect and alert when log delivery fails.
  • Protect log integrity and deletion permissions.
  • Test access, export, legal hold and deletion procedures.

Amazon Bedrock model invocation logging can record inputs and outputs to CloudWatch Logs or S3 when configured. That can be valuable for diagnosis and evaluation, but it can also duplicate sensitive data. Enable it only with an explicit data and retention design. Bedrock model invocation logging

16.3 Example service-level objectives​

ObjectiveExample targetResponse
API availability99.9% monthlyError budget and release freeze policy
p95 time to first token< 2.5 seconds for standard chatCapacity, routing and streaming review
p95 full response< 8 seconds for standard RAGRetrieval/model latency decomposition
Authorisation correctness100% regression passStop release on any failure
High-severity leakage0 accepted eventsImmediate containment and incident process
Groundedness≥ agreed use-case thresholdWarn, abstain or roll back when breached
Action approval coverage100% for tiers requiring approvalBlock tool execution if approval service is unavailable
Cost per completed taskWithin product budgetRoute, cache, compress context or redesign
Evidence completeness≥ 99.99% for consequential actionsFail closed or queue action until evidence path recovers

17. Reliability and disaster recovery​

Design graceful degradation before multi-Region complexity.

Failure modes and safe responses​

FailureSafe behaviour
Preferred model throttledRetry with jitter, queue, then use an evaluated fallback model
Guardrail unavailableFail closed for high-risk flows; offer non-AI/manual route
Retrieval unavailableDo not answer as if grounded; abstain or provide approved static help
Policy engine unavailableDeny tool actions; read-only degraded mode if authorised locally
Approval service unavailableQueue or cancel consequential actions; never auto-approve
Audit write unavailableFor consequential actions, pause or use a durable local outbox
Tool returns uncertain outcomeQuery status using idempotency key before retrying
Region outageRoute to pre-approved EU recovery Region with replicated policy and data controls
Bad prompt/model releaseRoll back manifest or disable feature flag
Poisoned corpusSwitch index alias to last approved snapshot and quarantine source

Resilience design​

  • Multi-AZ by default for production stateful services.
  • SQS buffering and dead-letter queues for asynchronous work.
  • Exponential backoff with jitter and bounded retries.
  • Circuit breakers per model and tool.
  • Bulkheads so one tenant, model or tool cannot exhaust the platform.
  • Transactional outbox for audit and side-effect events.
  • Idempotency for every retried write.
  • Point-in-time recovery and tested backups.
  • Documented RTO and RPO per data class, not one number for the platform.
  • Game days covering quota exhaustion, poisoned data, credential compromise, guardrail failure and Region loss.

For EU residency, a recovery design must use only approved Regions, services, keys and subprocessors. A technically successful failover that violates data policy is not successful.

18. Security operations and incident response​

Centralise findings in the Security Tooling account. AWS’s reference architecture places services such as Config, GuardDuty, Macie, Security Hub CSPM, Inspector, CloudTrail and Audit Manager under delegated administration to support organisation-wide visibility. AWS Security Tooling account guidance

Create AI-specific incident categories:

  • personal or confidential data exposure;
  • cross-tenant retrieval;
  • harmful or discriminatory output;
  • prohibited use;
  • prompt-injection compromise;
  • unauthorised or runaway tool action;
  • data/model poisoning;
  • material quality regression;
  • unavailable human oversight;
  • supplier/model incident;
  • evidence or logging failure;
  • abnormal consumption or denial of wallet.

The incident runbook should:

  1. contain traffic, model, prompt, corpus or tool capability;
  2. preserve appropriate forensic evidence;
  3. assess affected people, decisions and data;
  4. involve security, privacy, AI risk, legal, product and business owners;
  5. determine GDPR breach and AI Act serious-incident duties where applicable;
  6. notify within applicable timeframes;
  7. correct or reverse affected actions where possible;
  8. communicate clearly to users and customers;
  9. identify root cause and update controls, tests and documentation;
  10. obtain approval before restoration.

Maintain kill switches at multiple scopes: single session, tenant, tool, prompt version, model, corpus, use case and entire service. Test them.

19. Performance, scalability and cost engineering​

AI cost is a distributed-systems concern. A single request may trigger classification, query rewriting, retrieval, reranking, multiple reasoning turns, tool calls and response validation.

19.1 Capacity model​

Estimate:

Daily inference tokens=active users×tasks per user×calls per task×tokens per call\text{Daily inference tokens} = \text{active users} \times \text{tasks per user} \times \text{calls per task} \times \text{tokens per call}

Then model peak concurrency, context distribution, output length, retry rate, cache hit rate, model mix and tool latency. Test quotas before launch and request increases early.

19.2 Optimisation order​

  1. Remove unnecessary model calls.
  2. Use deterministic code for routing, policy and arithmetic.
  3. Reduce retrieved and historical context while preserving quality.
  4. Cache safe, non-personal and permission-stable results.
  5. Route simple tasks to smaller evaluated models.
  6. Batch offline embedding and evaluation work.
  7. Stream responses where it improves experience.
  8. Choose on-demand or provisioned capacity based on measured traffic.
  9. Limit agent steps, parallelism and retries.
  10. Optimise infrastructure only after end-to-end unit economics are visible.

Tag costs by tenant, use case, environment, model and feature. Allocate shared vector, logging and evaluation costs as well as model tokens. The product KPI should be cost per successful business outcome, not cost per token.

20. Responsible AI user experience​

Responsible architecture is visible in the interface.

Users should be able to see:

  • that they are interacting with AI where disclosure is required;
  • what the system can and cannot do;
  • sources and their dates for grounded answers;
  • when evidence is insufficient;
  • whether a response is a draft, recommendation or completed action;
  • the exact consequence of an approval;
  • how to correct, challenge, appeal or reach a human;
  • relevant use of their data and retention choices;
  • status of long-running or queued actions.

Human oversight is ineffective if the reviewer sees only an “Approve” button. Show the proposed action, evidence, uncertainty, changed fields, recipient, policy warnings and a safe alternative. Protect reviewers from automation bias through training, independent information and meaningful time to decide.

21. Governance roles and decision rights​

RoleAccountable for
Executive sponsorBusiness outcome, resources and risk appetite
AI system ownerIntended purpose, lifecycle, classification and overall control effectiveness
Product ownerUser need, acceptable behaviour and benefit realisation
Solution architectEnd-to-end design, boundaries, NFRs and trade-offs
Data owner/stewardAuthority, quality, provenance, access, retention and deletion
Model/AI engineeringModel, prompt, retrieval and evaluation implementation
SecurityThreat model, security controls, testing and incident response
Privacy/DPOLawful processing, DPIA advice, rights and transfer assessment
Legal/complianceRegulatory interpretation, contracts and conformity obligations
Human-oversight ownerReviewer competence, authority, workload and procedures
SRE/operationsSLOs, capacity, runbooks, monitoring and recovery
Internal auditIndependent assurance over control design and operation

Critical decisions need named approvers: initial use-case acceptance, risk classification, production release, high-risk conformity, supplier approval, residual-risk acceptance, emergency exception and retirement.

22. A phased implementation roadmap​

Phase 0: Discovery and control framing​

  • Create the use-case contract and business baseline.
  • Classify AI Act role/risk and screen for prohibited use.
  • Map data, jurisdictions, suppliers and sector rules.
  • Start the DPIA and AI impact assessment.
  • Define action tiers, human oversight, SLOs and risk appetite.
  • Decide build/buy and Bedrock/SageMaker/orchestrator options.

Exit: Sponsor, system owner, security, privacy and legal agree the allowed scope and evidence plan.

Phase 1: Secure cloud foundation​

  • Establish multi-account landing zone and delegated security.
  • Configure identity federation, MFA, break-glass access and SCPs.
  • Deploy VPCs, endpoints, egress control, KMS and central logs.
  • Define resource tags, configuration rules and budgets.
  • Create separate non-production and production pipelines.

Exit: Cloud security baseline and recovery controls pass review.

Phase 2: Read-only RAG minimum viable system​

  • Ingest only approved, low-risk sources.
  • Implement tenant-aware retrieval and citations.
  • Add input/output privacy and safety controls.
  • Create golden evaluation sets and hard security gates.
  • Launch to a small trained cohort without write actions.

Exit: Quality, privacy, security, latency and cost thresholds pass.

Phase 3: Production hardening​

  • Add progressive delivery, rollback, DR and load testing.
  • Operationalise evidence, incident response and user feedback.
  • Complete documentation, instructions for use and supplier reviews.
  • Test data-subject rights and deletion end to end.
  • Run independent penetration and AI red-team assessments.

Exit: Production readiness and residual-risk acceptance.

Phase 4: Bounded agentic actions​

  • Introduce one reversible, low-impact tool.
  • Deploy tool broker, Cedar policies, idempotency and approvals.
  • Test prompt/tool injection, partial failure and compensation.
  • Measure human workload and override quality.
  • Expand autonomy only when evidence supports it.

Exit: Each tool passes its action-tier controls and business-owner approval.

Phase 5: Continuous assurance and scale​

  • Post-market/production monitoring and periodic reassessment.
  • Automated control evidence plus internal audits.
  • Supplier, model and regulatory change monitoring.
  • Regular red teaming, game days and management review.
  • Retirement, data disposal and knowledge-transfer plans.

23. Production-readiness checklist​

Business and governance​

  • Intended purpose, users, exclusions and success measures are approved.
  • Provider/deployer/GPAI roles and AI Act classification are documented.
  • Prohibited-use and foreseeable-misuse assessment is complete.
  • AI system inventory and owners are current.
  • Residual risks have named acceptance authority.
  • User support, appeal, correction and retirement paths exist.

Privacy and data​

  • Data-flow map, ROPA and DPIA screening are complete.
  • Lawful basis and purpose are documented for every processing stage.
  • Data minimisation occurs before retrieval, inference and logging.
  • Region, transfer, supplier and subprocessor decisions are approved.
  • Retention and deletion propagate to chunks, vectors, cache and feedback.
  • Data-subject rights have been tested end to end.

Security​

  • Threat model includes traditional, AI, agent and supply-chain threats.
  • Tenant identity and authorisation are enforced outside the model.
  • Bedrock/model access is private and restricted to approved resources.
  • Tools use least-privilege, short-lived credentials and deterministic policy.
  • Inputs, retrieved data, outputs and tool arguments are independently validated.
  • Red-team, penetration and cross-tenant tests pass.
  • Secrets never enter prompts, datasets or unprotected logs.

Quality and responsible AI​

  • Golden datasets cover normal, edge, subgroup, adversarial and abstention cases.
  • Retrieval and generation are measured separately.
  • Security and severe-harm failures use zero-tolerance release gates.
  • Human evaluation validates automated metrics and LLM judges.
  • Human oversight is trained, authorised and able to stop/reverse.
  • Users see appropriate disclosure, sources, limitations and escalation.

Operations​

  • SLOs, error budgets, quotas and capacity tests are approved.
  • Fallback models and degraded modes have passed the same safety bar.
  • Retries are bounded and all writes are idempotent.
  • Audit/evidence failure produces a safe outcome.
  • Backups, restore, rollback and Region recovery are tested.
  • Kill switches exist at system, tenant, model, corpus and tool levels.
  • AI-specific incident response and regulatory assessment are rehearsed.

Delivery and evidence​

  • Code, infrastructure, model, prompt, guardrail, corpus, tools and policies are versioned.
  • CI/CD blocks releases without required evidence.
  • Production changes use separation of duties and progressive rollout.
  • Technical documentation and instructions for use match the deployed release.
  • Control owners, tests, frequency and remediation SLAs are assigned.
  • Management review and continual improvement operate on real metrics.

24. Common architecture mistakes​

  1. Starting with a chatbot and searching for a business case later. This produces unclear quality and risk thresholds.
  2. Calling all GenAI “low risk.” Intended purpose and downstream decision impact determine the control profile.
  3. Trusting RAG as a factuality or security guarantee. Retrieval can be wrong, poisoned or unauthorised.
  4. Putting authorisation in the prompt. The model cannot be the policy enforcement point.
  5. Giving one agent a broad enterprise role. A compromise then inherits the full blast radius.
  6. Logging everything for observability. This creates a second, poorly governed personal-data store.
  7. Using one retention period for all AI data. Source, sessions, security logs, approvals and evaluation evidence have different purposes.
  8. Treating human review as a decorative button. Oversight needs competence, information, authority and time.
  9. Evaluating only average answer quality. Rare security or rights failures can dominate risk.
  10. Promoting prompt changes without release controls. Prompt, policy and corpus changes can be as consequential as code.
  11. Using global cross-Region inference without a data-flow review. Availability optimisation can change processing locations.
  12. Assuming AWS or ISO certification certifies the customer workload. Customer design and operation remain decisive.

25. Final architecture position​

The strongest production AI architecture is not the one with the most models, agents or AWS services. It is the one that makes boundaries explicit:

  • AI is used only for a defined and valuable purpose.
  • The risk class and organisational role are known.
  • Data is minimised, authorised and traceable.
  • The model cannot grant itself access or authority.
  • Retrieval is permission-aware and source-grounded.
  • Agent actions are bounded, policy-controlled and reversible where possible.
  • Humans retain meaningful control over consequential outcomes.
  • Every release is evaluated as a complete system.
  • Failure produces a safe, understandable state.
  • Evidence is generated continuously rather than assembled in panic before an audit.

AWS provides a powerful foundation: a multi-account landing zone, strong identity and encryption services, private connectivity, managed foundation models, guardrails, scalable data systems, policy engines, monitoring and evidence tooling. The EU AI Act, GDPR, ISO/IEC 42001 and ISO/IEC 27001 add different but complementary requirements for lawful purpose, accountable governance, security, risk management, human oversight and continuous assurance.

The architectural goal is to connect them. A business requirement becomes an intended-purpose statement; the intended purpose determines classification; classification selects controls; controls become code and operating procedures; tests generate evidence; runtime telemetry detects drift; incidents and reviews drive continual improvement.

That is how an AI prototype becomes an enterprise system that can be trusted in production.

Authoritative references​

European Union and privacy​

  1. Regulation (EU) 2024/1689 — Artificial Intelligence Act
  2. European Commission — AI Act regulatory framework and current timeline
  3. European Commission — AI Act enforcement framework
  4. Regulation (EU) 2016/679 — General Data Protection Regulation
  5. EDPB Opinion 28/2024 on personal data and AI models

International standards and risk frameworks​

  1. ISO/IEC 42001:2023 — AI management systems
  2. ISO/IEC 27001:2022 — Information security management systems
  3. ISO/IEC 23894:2023 — Guidance on AI risk management
  4. NIST AI Risk Management Framework
  5. NIST AI 600-1 — Generative AI Profile
  6. OWASP Top 10 for LLM and GenAI applications

AWS architecture, security and AI​

  1. AWS Well-Architected Generative AI Lens
  2. AWS Well-Architected Agentic AI Lens
  3. AWS Security Reference Architecture
  4. AWS Shared Responsibility Model
  5. Amazon Bedrock data protection
  6. Amazon Bedrock encryption
  7. Amazon Bedrock Guardrails
  8. Amazon Bedrock geographic cross-Region inference
  9. Amazon Bedrock model invocation logging
  10. Amazon Bedrock RAG evaluation
  11. Amazon Verified Permissions
  12. AWS Audit Manager for GDPR evidence
  13. AWS Data Processing Addendum

Regulatory and service details are current to 21 August 2026 and should be revalidated before implementation or publication because law, regulatory guidance and AWS services evolve.

Discussion

Comments​

Share feedback or questions about this page. No account required.

Loading comments…