Skip to main content

Designing Production-Grade AI on Azure

· 57 min read
AI Playbook author

An end-to-end architecture for secure, scalable, observable, and governed AI under the EU AI Act, ISO/IEC 42001, ISO/IEC 27001, and GDPR

Status note, 21 August 2026: Microsoft now calls its Azure AI application and model platform Microsoft Foundry. Some older documentation and interfaces still use Azure AI Foundry. This article uses the current name while retaining familiar Azure service names such as Azure OpenAI, Azure AI Search, Azure Machine Learning, and Azure AI Content Safety.

Important: This is an engineering and governance blueprint, not legal advice or a guarantee of certification. Compliance depends on the organisation, use case, contractual roles, operating procedures, evidence, and the configuration actually deployed—not merely on choosing Azure services.

Executive summary​

A production-grade AI system is not an API call to a large language model. It is a socio-technical system containing data pipelines, identities, networks, models, prompts, retrieval, tools, human decisions, monitoring, incident response, supplier controls, and governance evidence.

The most important architectural principle is therefore:

Keep deterministic control outside the probabilistic model.

The model may interpret language, retrieve context, plan, summarise, classify, or propose an action. It must not be the sole enforcement point for authentication, authorisation, data access, transaction limits, regulatory policy, safety-critical decisions, or irreversible actions.

This reference architecture uses:

  • an Azure landing zone with separate platform and workload boundaries;
  • Microsoft Entra ID, workload identities, least privilege, and privileged identity management;
  • Azure Front Door and Web Application Firewall for a protected global edge;
  • Azure API Management as the governed AI gateway;
  • private Azure Kubernetes Service, Azure Container Apps, or Microsoft Foundry Agent Service as the application and orchestration runtime;
  • Microsoft Foundry model deployments, including Azure OpenAI models, behind private endpoints;
  • Azure AI Search for permission-filtered hybrid retrieval;
  • Azure Storage or Azure Data Lake Storage Gen2 for governed source data;
  • Cosmos DB or Azure Database for PostgreSQL for application state, plus Azure Managed Redis for short-lived state and carefully scoped caching;
  • a hardened tool broker for agent actions;
  • Azure Monitor, Application Insights, OpenTelemetry, Log Analytics, Microsoft Sentinel, Microsoft Defender for Cloud, Microsoft Purview, and immutable audit storage;
  • infrastructure as code, model and prompt versioning, evaluation gates, red-team tests, canary releases, and automated rollback.

Microsoft's own baseline Foundry architecture similarly uses private endpoints, managed identity, AI Search, controlled egress, state storage, and enterprise monitoring as the foundations of a production solution (Microsoft baseline Foundry chat architecture).

The outcome should be an AI platform that can answer five questions for every consequential output or action:

  1. Who requested it, and under which identity and role?
  2. What data, model, prompt, policy, and tools were used?
  3. Why was the result permitted, blocked, escalated, or approved?
  4. How well did the system perform against defined quality, safety, fairness, and reliability thresholds?
  5. What evidence proves that the required controls operated?

1. Start with the use case, not the model​

The reference workload in this article is a multi-tenant enterprise knowledge and action assistant. It can:

  • answer questions from authorised enterprise documents with citations;
  • summarise cases and draft responses;
  • search operational systems;
  • propose low-risk workflow steps;
  • execute approved actions through controlled tools;
  • hand off uncertain, sensitive, or high-impact cases to a qualified person.

It processes confidential business data and may process personal data. It is designed so that it can be strengthened for regulated or high-risk contexts, but the architecture does not assume that every enterprise assistant is automatically a high-risk AI system under the EU AI Act.

If the same platform is used for recruitment, employee evaluation, creditworthiness, access to essential services, biometric identification, medical safety functions, critical infrastructure, or another listed area, the classification and control obligations can change substantially. Classification must be based on the system's intended purpose, context, affected people, role in a decision, and reasonably foreseeable misuse.

Example non-functional requirements​

These figures are illustrative. Replace them with tested business requirements.

RequirementExample target
Registered users250,000
Daily active users15,000
Burst traffic50 requests/second
Concurrent streaming conversations600
Availability SLO99.9% monthly for the complete user journey
Time to first tokenp95 under 2 seconds
Grounded answer completionp95 under 8 seconds for normal requests
Retrieval permission leakageZero tolerated
Recovery time objective60 minutes
Recovery point objective5 minutes for transactional metadata
Data locationApproved EU regions or an approved EU Data Zone deployment
Public access to data servicesDisabled
High-impact actionsExplicit human approval required
Evidence retentionDefined by legal, security, and business schedules—not indefinitely

An SLO for only the model endpoint is insufficient. Users experience the entire chain: identity, edge, gateway, application, retrieval, model, tools, streaming, and persistence.


2. Ten architecture principles​

2.1 Classify before you build​

Create a documented intake gate covering business purpose, affected people, benefits, failure impact, prohibited uses, EU AI Act role and risk, GDPR applicability, data categories, security classification, human oversight, and sector rules. A model choice should follow this assessment, not precede it.

2.2 Separate the control plane, data plane, and evidence plane​

  • The control plane manages approved models, prompts, policies, identities, infrastructure, deployments, and configuration.
  • The data plane handles live prompts, retrieval, inference, tool calls, and responses.
  • The evidence plane records control operation, evaluation results, approvals, versions, incidents, access reviews, and change history.

This separation limits blast radius and makes audit evidence easier to protect and query.

2.3 Assume every natural-language input is untrusted​

User prompts, retrieved documents, websites, emails, tool responses, model outputs, and agent-to-agent messages can all contain malicious instructions. Treat them as untrusted data. Prompt injection is not solved by writing a stronger system prompt.

2.4 Enforce authorisation before retrieval and before action​

Never retrieve a document and then ask the model whether the user should see it. Apply tenant, identity, group, purpose, and document-level access rules in the retrieval query. Reauthorise every tool action against deterministic policy immediately before execution.

2.5 Minimise data at every boundary​

Send the model only the context required for the task. Redact or pseudonymise where possible. Do not put secrets in prompts. Do not log raw prompts and responses by default. Give every store a retention purpose and deletion path.

2.6 Design for bounded autonomy​

An agent should have a narrow tool set, minimum permissions, maximum loop count, time and token budgets, transaction limits, schema-validated arguments, idempotency controls, and human approval for consequential actions.

2.7 Evaluate the system, not only the base model​

A strong base model can still produce a poor application because of weak retrieval, bad prompts, excessive permissions, unsafe tools, stale data, or a broken user interface. Test the composed system end to end.

2.8 Make every change reversible​

Pin versions, use feature flags, support blue-green or canary deployment, retain the last known-good configuration, and implement a kill switch that disables models, agents, connectors, tools, tenants, or features independently.

2.9 Prefer graceful degradation over unsafe availability​

When a model, index, or tool fails, the safe response may be search-only results, a read-only mode, a queued job, or human handoff—not silent use of an untested model or an unrestricted tool.

2.10 Build evidence as a product feature​

An auditor should not need to reconstruct the system from screenshots and memory. Generate traceable evidence from version control, pipelines, policy compliance, evaluation runs, approvals, monitoring, incident systems, and access reviews.


3. Azure reference architecture​

The diagram is deliberately logical rather than tied to a single compute product. The runtime can be:

  • Microsoft Foundry Agent Service when a managed agent runtime meets the network, state, tool, and governance requirements;
  • Azure Container Apps for a simpler event-driven or microservice architecture with lower platform-operating effort;
  • private AKS when the organisation needs maximum control over networking, sidecars, policy enforcement, scaling, specialised compute, multi-service orchestration, or portability.

For a complex regulated multi-tenant platform, this article assumes a private AKS or Container Apps environment for application services and Microsoft Foundry for approved model endpoints. A managed Foundry agent can replace parts of the custom orchestrator after a documented fit-gap and risk review.

3.1 Platform foundation​

Use an Azure landing zone, not a flat subscription. Microsoft describes the platform landing zone as the central foundation for governance, security, identity, connectivity, and management, with application landing zones hosting individual workloads (Azure landing zones).

A practical structure is:

  • platform management group;
  • connectivity subscription with hub network or Virtual WAN, Azure Firewall, DNS, VPN or ExpressRoute, and DDoS controls;
  • security and management subscription with Log Analytics, Sentinel, Defender for Cloud, and central dashboards;
  • identity services and privileged administration boundary;
  • separate development, test, pre-production, and production workload subscriptions;
  • optional dedicated data and AI platform subscriptions where scale or organisational ownership justifies them;
  • sandbox isolated from production data, identities, networks, quotas, and model deployments.

Enforce approved regions, mandatory tags, diagnostic settings, private networking, encryption, managed identities, Defender plans, backup, and resource restrictions through Azure Policy initiatives. Use exemptions that are time-limited, owned, justified, and reviewed.

3.2 Network topology and trust boundaries​

The default data path should be private:

  1. Users enter through Azure Front Door Premium with TLS, Web Application Firewall, bot protection, geo controls where lawful, and rate limits.
  2. The origin accepts traffic only from the approved Front Door path. Where supported, use Private Link. Otherwise combine platform service tags or approved address ranges with strict validation of the Front Door identifier. Microsoft explicitly warns that an unrestricted origin can bypass Front Door's WAF and DDoS controls (secure Front Door origins).
  3. Azure API Management validates identity and claims, enforces consumer quotas, sanitises headers, applies request limits, assigns correlation identifiers, and routes to an internal application endpoint.
  4. Application services call Azure PaaS services through private endpoints.
  5. Private DNS zones are centrally managed and linked only to approved networks.
  6. All non-private egress follows user-defined routes through Azure Firewall or an approved secure egress service.
  7. Egress is deny-by-default and allowlisted by destination, port, identity, and purpose. DNS and firewall logs are retained as security evidence.
  8. Administrative access uses privileged workstations, Entra Privileged Identity Management, just-in-time access, Azure Bastion where needed, and no routine public management endpoints.

Do not enable public network access temporarily and forget to remove it. Make private-only configuration a policy and pipeline test.

3.3 Identity and access​

Use Microsoft Entra ID or Entra External ID for user authentication. Apply:

  • phishing-resistant multifactor authentication for privileged and sensitive roles;
  • Conditional Access based on device, risk, location, and workload sensitivity;
  • short-lived tokens;
  • group and application-role claims that map to explicit business permissions;
  • Entra PIM for time-bound privileged access;
  • quarterly or risk-based access reviews;
  • break-glass accounts that are monitored, tested, and excluded only from controls necessary for recovery.

Use managed identities and workload identity federation for service-to-service calls. Avoid client secrets and storage keys. A service identity should have access only to the specific model deployment, search index, storage container, queue, database schema, or secret it needs.

For actions performed on behalf of a person, preserve the user's identity using a supported delegated or on-behalf-of flow. Do not collapse all user actions into a powerful shared service principal. Record both the initiating human and executing workload identity.

3.4 Edge and AI gateway​

Azure API Management is the policy enforcement point between applications and model endpoints. Its AI gateway capabilities include token quotas and rate limits, and it can route among backends. Microsoft documents token-per-minute limits and quotas by consumer, together with semantic caching options (API Management AI gateway capabilities).

Use it for:

  • OAuth token validation and application identity;
  • per-user, per-tenant, per-product, and per-model request limits;
  • prompt and body-size limits;
  • token-per-minute and token-budget enforcement;
  • model allowlists and routing policy;
  • timeout, retry, circuit-breaker, and backend-health behaviour;
  • usage metering, cost attribution, and chargeback;
  • response header policy and removal of internal details;
  • consistent diagnostic events without indiscriminate content logging;
  • controlled failover only to pre-approved and pre-evaluated deployments.

Do not use the dedicated API Management AI Gateway tier as the default production dependency while it remains in preview. Standard generally available API Management capabilities are the safer baseline; preview adoption requires explicit risk acceptance and an exit plan.

3.5 Application runtime​

Split the application into services with clear responsibilities:

ServiceResponsibility
Conversation APISessions, streaming, user-visible history, feedback, deletion
Policy decision pointDeterministic access, purpose, tenant, risk, and action policy
OrchestratorExecutes a bounded workflow and manages model, retrieval, and tool calls
Retrieval serviceQuery transformation, hybrid search, reranking, ACL filters, citations
Guardrail servicePII checks, injection signals, content safety, output policy
Tool brokerTool registry, argument validation, authorisation, approval, execution
Evaluation serviceOffline and sampled online quality and safety measurements
Audit serviceAppend-only business and compliance events with integrity controls
Ingestion serviceSource validation, malware scanning, extraction, classification, indexing
Admin serviceApproved configuration, models, prompts, policies, tenants, kill switches

Use synchronous calls only on the latency-critical path. Use Azure Service Bus for ingestion, long-running tools, evaluation, notification, and retryable background work. Configure dead-letter queues, duplicate detection or idempotency, retry ceilings, poison-message handling, and operational ownership.

3.6 Model layer​

Deploy approved models through Microsoft Foundry. Use separate deployments for development, testing, pre-production, and production. Separate workloads further when data classification, capacity, safety configuration, or regulatory scope differs.

Model selection should be an evidence-based decision across:

  • task success and domain accuracy;
  • groundedness and citation quality;
  • harmful-content behaviour;
  • fairness and subgroup performance where relevant;
  • prompt-injection and jailbreak resistance;
  • tool-call accuracy;
  • latency, context window, throughput, availability, and quota;
  • data location and service processing behaviour;
  • model licence, acceptable-use restrictions, and supplier terms;
  • cost per successful business outcome, not cost per token alone.

Pin a model version where the service permits it. Treat a model version, system prompt, tool schema, retrieval configuration, content filter, and policy set as a single release unit. A model fallback is not a purely operational change: it can change accuracy, safety, latency, language coverage, and legal risk. Evaluate fallback paths in advance.

Microsoft states that prompts, completions, embeddings, and training data for models sold by Azure are not made available to other customers or model providers and are not used to train foundation models without permission or instruction (Foundry model data, privacy, and security). This is useful supplier evidence, but it does not remove the controller's duty to assess the chosen service, feature, deployment type, region, retention, abuse-monitoring process, contracts, subprocessors, and actual application logging.

3.7 RAG and knowledge architecture​

Azure AI Search is a strong default for enterprise RAG because it supports keyword, vector, hybrid retrieval, filtering, and semantic ranking. Microsoft's baseline architecture also uses AI Search as its primary grounding store (grounding store guidance).

The ingestion path should be:

For every document and chunk, keep metadata such as:

  • tenant and source system;
  • authoritative document identifier and version;
  • owner and business domain;
  • security classification and sensitivity label;
  • allowed users, groups, roles, or policy attributes;
  • purpose restriction;
  • jurisdiction and data-residency tag;
  • effective and expiry dates;
  • legal hold and retention class;
  • extraction and chunking version;
  • embedding model and version;
  • content hash and lineage identifiers;
  • deletion or supersession status.

Retrieval sequence​

  1. Authenticate the user.
  2. Resolve tenant, groups, entitlements, purpose, and risk context.
  3. Classify the query for sensitivity and allowed data domains.
  4. Normalise and safely transform the query without changing its security context.
  5. Apply tenant and document-level ACL filters at query time.
  6. Run hybrid keyword and vector search.
  7. Rerank only authorised candidates.
  8. Apply freshness, jurisdiction, source-authority, and policy filters.
  9. Assemble the smallest sufficient context.
  10. Generate an answer with verifiable citations.
  11. Check that material claims are supported by retrieved sources.
  12. Return the answer, uncertainty, citations, and safe next steps.

Never rely on namespace naming alone for tenant separation. Apply explicit tenant filters, test cross-tenant attacks, and consider separate indexes, search services, subscriptions, or encryption boundaries for high-isolation tenants.

RAG evaluation​

Measure the pipeline in layers:

  • ingestion coverage and extraction accuracy;
  • chunk correctness and metadata completeness;
  • retrieval recall@k, precision@k, mean reciprocal rank, and nDCG where appropriate;
  • permission-filter correctness;
  • context relevance;
  • groundedness and citation correctness;
  • answer completeness and task success;
  • correct refusal when evidence is missing;
  • latency and cost;
  • deletion propagation time.

A fluent answer without a supported citation is not a successful RAG answer.

3.8 Agent tools and action safety​

Tools create a much larger risk than text generation because they convert model output into real-world effects.

Every tool must be registered with:

  • an owner;
  • a business purpose;
  • input and output schemas;
  • data classification;
  • required human and workload permissions;
  • read, write, financial, legal, or safety impact class;
  • maximum transaction value or scope;
  • approval rules;
  • timeout, retry, and idempotency behaviour;
  • network destination;
  • logging and retention rules;
  • test cases and kill switch;
  • review and expiry date.

The safe execution sequence is:

  1. The model proposes a structured tool call.
  2. The broker validates the tool is approved for this agent and tenant.
  3. JSON Schema or equivalent validation rejects unknown fields and invalid types.
  4. The broker independently authorises the user, purpose, resource, and action.
  5. A policy engine checks transaction limits, segregation of duties, and risk.
  6. High-impact actions generate a human-readable preview and require explicit approval.
  7. The executing service obtains a short-lived, least-privileged credential.
  8. The operation uses an idempotency key and bounded timeout.
  9. The broker validates and sanitises the response before returning it to the model.
  10. The audit event records request, decision, approver, result, and relevant versions.

The model must never generate SQL, shell commands, URLs, or API parameters that are executed without deterministic validation. Use parameterised queries, allowlisted destinations, fixed command templates, egress restrictions, and safe parsers.

3.9 State, memory, and caching​

Use different stores for different state:

  • Cosmos DB for globally distributed conversation or agent state when its access pattern and consistency options fit;
  • Azure Database for PostgreSQL for relational business data, strong constraints, row-level security, and complex transactions;
  • Azure Managed Redis for short-lived sessions, locks, idempotency, and carefully scoped caches;
  • Blob Storage or ADLS Gen2 for documents, evaluation datasets, reports, and immutable evidence;
  • Service Bus for durable asynchronous commands and events.

Do not retain hidden chain-of-thought. Store the minimum user-visible messages, tool decisions, citations, feedback, and audit facts necessary for the declared purpose. Long-term memory should be opt-in or purpose-justified, editable, attributable, and deletable.

Semantic caching can reduce cost and latency, but it can also leak another person's response, serve stale policy, or reuse an answer under different permissions. For personalised or regulated answers, disable it by default. Where caching is approved, include at least the following in the key:

  • tenant;
  • user or permission-set hash;
  • purpose;
  • corpus and ACL version;
  • model and prompt version;
  • safety and policy version;
  • locale;
  • response mode;
  • data-classification boundary.

Use short TTLs, encryption, explicit invalidation, and a rule that cached content is rechecked against current policy before release.


4. End-to-end request flow​

The following sequence shows how a production request should operate.

  1. Client integrity: The client establishes TLS and sends a short-lived Entra token, request identifier, tenant context, locale, and user input. It does not contain model keys.
  2. Edge defence: Front Door applies WAF, bot management, request-size limits, rate limits, and DDoS protections.
  3. Gateway authentication: API Management validates issuer, audience, signature, expiry, scopes, application identity, and tenant policy.
  4. Abuse and quota controls: The gateway applies request and token budgets before expensive work begins.
  5. Context resolution: The application resolves user roles, groups, data entitlements, consent or preference state, purpose, and risk class.
  6. Input minimisation: The guardrail layer detects secrets, sensitive data, disallowed content, and injection indicators. It blocks, redacts, pseudonymises, or routes according to policy.
  7. Workflow selection: A deterministic router selects an approved workflow. The model cannot invent a new workflow or tool.
  8. Retrieval: The retrieval service applies tenant and ACL filters, searches authorised sources, reranks results, and assembles a minimal context package with citations and provenance.
  9. Prompt construction: The system separates trusted instructions from untrusted user and document content, marks source boundaries, supplies tool schemas, and sets output constraints.
  10. Model routing: The gateway chooses an approved model deployment based on task, sensitivity, region, capacity, and release configuration.
  11. Generation or planning: The model returns text or a structured proposed tool call.
  12. Output validation: Content safety, sensitive-data checks, groundedness checks, schema validation, and business rules run outside the model.
  13. Tool authorisation: If a tool is requested, the broker performs a fresh deterministic authorisation and obtains approval when required.
  14. Execution: The tool runs with a least-privileged short-lived identity, idempotency key, timeout, and allowlisted network path.
  15. Response composition: The application returns a cited answer, uncertainty or limitations, AI disclosure, action result, and escalation path.
  16. Persistence: Only purpose-approved conversation state is stored. Raw prompts are not copied into every telemetry system.
  17. Evidence: The audit service records the actor, policy decision, model/prompt/index/tool versions, approval, safety outcome, and result using pseudonymous identifiers where possible.
  18. Monitoring: Operational, quality, safety, security, privacy, and business signals feed dashboards and alerts.
  19. Feedback and correction: Users can report a problem, challenge a consequential outcome, request human review, and exercise applicable data rights.

5. System design for scale and reliability​

5.1 Capacity model​

Size each dependency separately. A useful starting point is:

Peak TPM=requests per minute×(p95 input tokens+p95 output tokens)×safety factor\text{Peak TPM} = \text{requests per minute} \times (\text{p95 input tokens} + \text{p95 output tokens}) \times \text{safety factor} In-flight concurrency≈peak requests per second×p95 request duration in seconds\text{In-flight concurrency} \approx \text{peak requests per second} \times \text{p95 request duration in seconds} Search QPS=request QPS×retrieval calls per request×agent loop factor\text{Search QPS} = \text{request QPS} \times \text{retrieval calls per request} \times \text{agent loop factor}

Use measured distributions, not averages. A small number of huge prompts can dominate capacity and cost. Enforce maximum context, output, tool iterations, uploaded-file size, and request duration.

Reserve independent capacity for:

  • interactive inference;
  • embeddings and ingestion;
  • offline evaluation;
  • red-team exercises;
  • batch processing;
  • secondary-region failover;
  • critical tenants or workflows.

One workload must not exhaust the entire model quota. Apply per-consumer budgets at API Management and deployment-level quota controls.

5.2 Latency budget​

An illustrative streaming latency budget is:

Stagep95 budget
Edge, WAF, and gateway150 ms
Authentication and policy context150 ms
Query transformation150 ms
Search and reranking500 ms
Prompt assembly and guardrails250 ms
Model time to first token1,500 ms
Streaming generationWorkload dependent
Post-generation validation300 ms

Measure user-perceived time to first meaningful content, not only server response time. Stream safely: do not release unvalidated high-impact recommendations token by token if the full response requires policy or groundedness validation.

5.3 Autoscaling​

Scale stateless application pods by CPU, memory, request concurrency, queue length, and custom latency metrics. KEDA can scale background workers from Service Bus depth. Maintain minimum warm capacity for interactive services to avoid cold starts.

Protect dependencies using bulkheads:

  • separate worker pools by workflow or tenant class;
  • separate queues for critical and non-critical work;
  • per-dependency connection pools;
  • concurrency limits for search, model, and tools;
  • bounded retries with exponential backoff and jitter;
  • circuit breakers that fail safely;
  • admission control before overload becomes collapse.

Never automatically retry a non-idempotent action unless the downstream system and idempotency design make it safe.

5.4 Multi-region design​

Use active-active stateless application services when the business SLO justifies the cost. Keep regional dependencies aligned with data-location obligations.

A typical pattern is:

  • Front Door routes to healthy regional application stacks;
  • each region has independent runtime, API capacity, model deployment, search capacity, queues, secrets, and monitoring;
  • configuration is promoted from the same signed release;
  • transactional data uses a documented replication and conflict strategy;
  • source documents and indexes follow an approved replication plan;
  • the knowledge index is reproducible from governed sources;
  • failover is tested, not assumed.

Do not use a global or cross-region model deployment merely for availability if it conflicts with the approved processing boundary. Select regional, data-zone, or global deployment only after privacy, security, resilience, capacity, and contractual review.

Define degradation levels:

ConditionSafe behaviour
Primary model throttledRoute to pre-evaluated equivalent in the approved boundary
No approved model availableSearch-only results or queue request; do not use an unapproved model
Search unavailableDisable grounded answers; provide service notice and human route
Tool unavailableReturn draft or pending action; do not claim execution
Safety service unavailableFail closed for high-risk flows; tightly limited fallback for low-risk flows
Audit path unavailableBuffer durably; block consequential actions if evidence cannot be recorded
Identity or policy unavailableFail closed

5.5 Backup and recovery​

Back up source documents, configuration, prompts, policy bundles, tool registry, evaluation datasets, transactional state, audit evidence, and infrastructure definitions according to distinct RPOs and retention rules.

Test:

  • point-in-time database restoration;
  • accidental document deletion;
  • corrupted index rebuild;
  • compromised prompt or policy rollback;
  • region loss;
  • key rotation and emergency revocation;
  • restoration without resurrecting data that was lawfully deleted;
  • recovery of evidence with integrity and access controls preserved.

An index should normally be rebuildable from the authoritative source plus versioned processing code. Do not treat the vector index as the sole copy of a business record.


6. Observability: five kinds of truth​

Traditional APM is necessary but insufficient. A production AI system needs five views.

6.1 Operational telemetry​

  • availability and error rate;
  • p50, p95, and p99 latency;
  • time to first token and tokens per second;
  • queue age and depth;
  • dependency saturation and throttling;
  • model, search, database, and tool health;
  • regional routing and failover state;
  • deployment and configuration versions.

6.2 AI quality telemetry​

  • task completion;
  • groundedness and relevance;
  • citation correctness;
  • retrieval recall and permission-filter correctness;
  • tool-call selection and argument accuracy;
  • abstention and escalation correctness;
  • drift by language, tenant, domain, and workflow;
  • user corrections and upheld complaints.

6.3 Safety and security telemetry​

  • prompt-injection and jailbreak signals;
  • sensitive-data disclosure attempts;
  • policy blocks and overrides;
  • unusual tool-call patterns;
  • cross-tenant access attempts;
  • repeated denied actions;
  • model, prompt, connector, or dependency changes;
  • suspicious egress and secret access;
  • content-safety category and severity trends.

6.4 Privacy and compliance telemetry​

  • data-subject request progress;
  • deletion propagation and failures;
  • retention-policy execution;
  • access reviews and privileged activity;
  • DPIA and risk-treatment actions;
  • approval coverage for high-impact actions;
  • model, dataset, prompt, and policy lineage;
  • evidence completeness and exceptions.

6.5 Business telemetry​

  • successful outcome rate;
  • containment versus appropriate human handoff;
  • time saved and cycle-time reduction;
  • cost per successful outcome;
  • user adoption and repeat use;
  • harm, complaint, rework, and error rates;
  • benefits by stakeholder group, including possible unequal outcomes.

Microsoft Foundry integrates evaluation, production monitoring, and OpenTelemetry-based tracing with Azure Monitor and Application Insights, including metrics such as groundedness, relevance, safety, tool accuracy, latency, errors, and token usage (Foundry observability).

Telemetry privacy rules​

  • Do not record secrets, access tokens, hidden instructions, or full retrieved documents.
  • Default to hashes, categories, counts, pseudonymous IDs, and structured decisions.
  • Put any approved content capture in a separate, highly restricted store with a shorter retention period.
  • Apply RBAC, private endpoints, encryption, immutable retention where justified, and monitored break-glass access.
  • Do not use production prompts for training or evaluation without a documented lawful basis, purpose, minimisation, and approval.
  • Avoid recording private model reasoning. Capture inputs, outputs, decisions, tool events, and user-visible explanations required for accountability.

7. Security architecture and threat model​

The latest OWASP GenAI guidance should be incorporated into threat modelling; the OWASP GenAI LLM Top 10 2026 is the current release as of this article (OWASP 2026 release). Also use OWASP API Security, ASVS, the OWASP Top 10 for Agentic Applications, MITRE ATLAS, and normal cloud and application threat modelling. AI risks do not replace familiar vulnerabilities such as broken access control, SSRF, injection, insecure deserialisation, dependency compromise, or credential theft.

ThreatExampleRequired controlsAzure implementation and evidence
Direct prompt injectionUser asks the agent to ignore policy and reveal dataTreat prompts as untrusted; deterministic policy; minimum context; output checksGuardrail service, Foundry content controls, APIM limits, adversarial test report
Indirect prompt injectionA retrieved document tells the agent to exfiltrate memorySeparate instructions from data; source trust; sanitise content; restrict tools and egressIngestion scanner, metadata trust score, Azure Firewall allowlist, retrieval trace
Sensitive-data disclosureModel exposes another tenant's documentAuthorise before retrieval; tenant isolation; DLP; no shared cacheEntra, AI Search filters or separate indexes, Purview labels, isolation tests
Excessive agencyAgent refunds money or deletes records without approvalBounded tools, transaction limits, human approval, idempotency, kill switchTool broker, policy engine, Service Bus, approval and audit events
Improper output handlingModel output becomes SQL, HTML, code, or commandSchema validation, encoding, parameterisation, sandboxing, allowlistsAPI schemas, parameterised data layer, CSP/WAF, secure code tests
Data or model poisoningMalicious source changes embeddings or evaluation resultsApproved sources, signatures and hashes, provenance, review, anomaly testsADLS immutable raw zone, Defender scans, Purview lineage, signed releases
Supply-chain compromiseMalicious package, model, container, or MCP serverSCA, SBOM, signatures, licence review, dependency pinning, provenanceDefender for DevOps/Cloud, ACR scanning, signed images, supplier register
Vector-store weaknessACL metadata omitted or stale index serves deleted dataSchema validation, ACL tests, index versioning, deletion SLAAI Search index gates, reconciliation jobs, cross-tenant test suite
Model theft or extractionHigh-volume probing reconstructs behaviour or dataRate limits, anomaly detection, output minimisation, no secret promptsAPIM quotas, Sentinel analytics, alert evidence
Denial of wallet or serviceHuge contexts and recursive tool loops consume quotaSize, token, time, loop, and spend limits; quotas; admission controlAPIM token policies, orchestrator budgets, Cost Management alerts
Hallucination and misinformationConfident unsupported answer changes a business decisionRAG, citations, groundedness threshold, abstention, reviewAI Search, Foundry evaluation, user-visible evidence, release gate
Insecure MCP or tool discoveryAgent connects to an unapproved server or tool changes silentlyCurated registry, pinned schema/version, private endpoints, egress denyAPI Center/tool registry, APIM, private Container Apps, Firewall logs
Memory poisoningA malicious instruction persists across sessionsTyped memory, user visibility, source attribution, TTL, sanitisationCosmos/PostgreSQL schema, retention policy, memory security tests
Cross-agent manipulationOne agent sends untrusted instructions to anotherAuthenticate agents, signed envelopes, schema and policy per hopManaged identities, APIM, message schema, Service Bus RBAC
Insider misuseAdministrator accesses prompts or weakens filtersPIM, segregation of duties, approval, monitoring, immutable auditEntra PIM, access reviews, Sentinel, policy-change alerts
Secret leakageKey appears in code, prompt, trace, or model outputManaged identity, Key Vault, scanning, redaction, rotationKey Vault or Managed HSM, secret scan results, rotation records
Unsafe model fallbackOutage routes to an unevaluated modelPre-approved routing matrix, equivalent policies, no silent downgradeAPIM backend policy, evaluation artefact, change approval

Defence in depth for prompt injection​

No single detector is sufficient. Combine:

  1. data and instruction separation;
  2. content provenance and source trust;
  3. minimum retrieved context;
  4. deterministic authorisation;
  5. a narrow tool allowlist;
  6. per-tool schema and policy validation;
  7. network egress restrictions;
  8. sensitive-data filtering;
  9. sandboxing for code or file processing;
  10. human approval for consequential actions;
  11. adversarial testing and monitoring;
  12. graceful refusal and escalation.

Azure AI Content Safety and Foundry guardrails can detect categories of harmful content and some advanced threats, but they are one layer, not a replacement for architecture-level controls (Microsoft Foundry guardrails).

Encryption and keys​

  • Use TLS 1.2 or later according to organisational baseline and service support.
  • Encrypt data at rest with platform-managed keys by default; use customer-managed keys where the risk assessment, contracts, or regulation require them.
  • Keep keys in Azure Key Vault or Managed HSM with soft delete, purge protection, rotation, private endpoints, logging, and separated administration.
  • Use different keys or vaults where tenant, environment, or data-classification isolation requires them.
  • Do not assume customer-managed encryption prevents authorised application processing; it primarily changes key control and revocation.
  • Document cryptographic dependencies and the recovery effect of disabling or losing a key.

Secure software and model supply chain​

Every release should produce:

  • source commit and reviewed pull request;
  • build provenance;
  • software bill of materials;
  • signed container or package;
  • dependency and licence scan;
  • SAST, secret scanning, infrastructure-as-code scanning, and container scanning;
  • API and dynamic security tests;
  • model, dataset, prompt, policy, and tool versions;
  • model supplier and licence assessment;
  • evaluation and red-team results;
  • approval and deployment record.

Use private Azure Container Registry, content trust or an approved signing framework, admission policy, and Defender for Cloud. Production should accept only approved signed artefacts from the release pipeline.


8. EU AI Act by design​

8.1 Current timeline​

The EU AI Act entered into force on 1 August 2024. Prohibited practices and AI literacy obligations applied from 2 February 2025; governance and general-purpose AI obligations applied from 2 August 2025; most other provisions and Article 50 transparency rules apply from 2 August 2026. Following the July 2026 AI Omnibus, high-risk rules for certain Annex III areas apply from 2 December 2027, while high-risk systems embedded in regulated Annex I products have an extended period until 2 August 2028 (European Commission AI Act timeline).

Do not interpret later high-risk dates as permission to postpone architecture. Data lineage, logging, testing, human oversight, technical documentation, and quality management are expensive to retrofit.

8.2 The classification gate​

For each system, record:

  1. intended purpose and excluded purposes;
  2. whether it meets the AI-system definition;
  3. whether a prohibited practice is involved;
  4. whether it is a safety component or a listed high-risk use;
  5. whether an Annex III exception applies and why;
  6. whether Article 50 transparency duties apply;
  7. whether the organisation is provider, deployer, importer, distributor, product manufacturer, authorised representative, or multiple roles;
  8. whether a third-party model is a GPAI model and what downstream duties remain;
  9. affected people and fundamental rights;
  10. intended geography and where outputs are used;
  11. whether a modification, rebranding, or change of purpose could make the organisation the provider;
  12. required conformity, registration, post-market, and incident activities.

The system owner, legal counsel, DPO, security, risk, domain expert, and responsible-AI lead should approve the classification. Reassess it when purpose, model, data, autonomy, user group, geography, or tool access changes.

8.3 Transparency obligations​

Article 50 applies from 2 August 2026. The Commission's guidance states, among other things, that people should be informed when directly interacting with an AI system and that providers of certain generative systems must support machine-readable marking of generated or manipulated content; deployers have specified disclosure duties for deepfakes and certain public-interest text (Article 50 transparency guidance).

Implement transparency in the product:

  • clearly identify the experience as AI where required;
  • state purpose, capabilities, limitations, and data use in accessible language;
  • show citations and distinguish source text from generated text;
  • show when a human has reviewed or approved a result;
  • label or mark generated content when applicable;
  • disclose uncertainty and the route to human assistance;
  • make the current model or system version available to support and audit functions;
  • provide accessible notices before—not after—the person relies on the system.

8.4 High-risk control mapping​

The consolidated AI Act requires high-risk systems to support risk management, appropriate data governance, technical documentation, automatic logs, deployer transparency, human oversight, and appropriate accuracy, robustness, and cybersecurity. It also requires provider quality management, conformity and post-market activities. The official text specifically requires effective human oversight and the ability, as appropriate, to disregard, override, reverse, intervene, or safely stop the system (consolidated AI Act, Articles 9–17).

AI Act needArchitecture responseEvidence
Risk-management systemAI risk register integrated with enterprise risk; foreseeable misuse; test thresholds; residual-risk approvalApproved risk assessment, threat model, treatments, acceptance
Data and data governanceApproved sources, provenance, representativeness tests, quality rules, lineage, ACL metadata, deletionDataset cards, Purview lineage, quality report, source approvals
Technical documentationSystem description, architecture, versions, intended purpose, limitations, metrics, controls, dependenciesVersioned technical file generated per release
Record keepingAutomatic structured events across identity, retrieval, inference, tools, approvals, and incidentsProtected audit records, schema, retention and integrity report
Transparency to deployersInstructions for use, expected performance, limitations, input requirements, oversight and maintenanceProduct guide, release notes, model/system card
Human oversightRole-based review, meaningful context, override, stop, escalation, automation-bias trainingApproval logs, training records, UI tests, override metrics
AccuracyIntended-purpose metrics, subgroup tests, acceptance thresholds, ongoing monitoringEvaluation report and signed release gate
RobustnessFault handling, redundancy, adversarial tests, drift detection, safe degradationChaos and resilience tests, SLOs, failover results
CybersecurityThreat modelling, isolation, injection defence, poisoning defence, least privilege, incident responseSecurity test report, scan results, incident exercises
Quality managementDocumented lifecycle, responsibilities, change control, supplier management, CAPAQMS/AIMS procedures, audits, management review
Conformity and registrationFormal legal and quality workflow before market or use where applicableDeclaration, assessment, registration and approval records
Post-market monitoringComplaints, performance trends, serious incidents, corrective action, recall or shutdownMonitoring plan, incident reports, CAPA and release history

For high-risk decisions, “human in the loop” must be meaningful. A rushed person who sees only the model's score and almost always clicks approve is not an effective control. The reviewer needs competence, time, authority, relevant evidence, system limitations, an independent way to check the case, and the ability to refuse or stop.

8.5 AI literacy​

Create role-specific training for:

  • end users;
  • human reviewers;
  • product owners;
  • data scientists and AI engineers;
  • security and privacy teams;
  • operations and incident responders;
  • procurement and vendor managers;
  • executives and risk committees.

Training should cover limitations, automation bias, safe prompting, data handling, challenge and escalation, tool risk, incident reporting, prohibited use, and the responsibilities of the specific system. Record attendance and assess effectiveness.


9. GDPR by design​

GDPR applies to personal-data processing regardless of whether AI is used. The European Commission summarises its seven principles as lawfulness, fairness and transparency; purpose limitation; data minimisation; storage limitation; accuracy; integrity and confidentiality; and accountability (GDPR processing principles).

9.1 Privacy architecture checklist​

Purpose and lawful basis​

  • Define a specific purpose for every input, retrieval source, memory store, evaluation dataset, telemetry event, and downstream action.
  • Identify and document an Article 6 lawful basis; identify any Article 9 condition for special-category data.
  • Do not default to consent when it is not freely given or practical to withdraw.
  • Reassess compatibility before reusing production data for evaluation, fine-tuning, analytics, or a new feature.

Controller and processor roles​

  • Map the controller, joint controller, processor, subprocessor, and independent-controller relationships.
  • Put appropriate processor terms, security obligations, deletion, assistance, audit, breach, location, and subprocessor controls into contracts.
  • Maintain a current supplier and data-flow register.

DPIA​

The European Commission states that a DPIA is required when processing is likely to result in a high risk to people's rights and freedoms, including specified cases such as systematic and extensive evaluation, large-scale sensitive-data processing, or large-scale monitoring (DPIA guidance).

Perform the DPIA before deployment and revisit it after material changes. Link it to the AI impact assessment and security threat model, but do not collapse them into one vague document: each has distinct legal and governance purposes.

Data minimisation​

  • Pseudonymise identifiers before model calls where feasible.
  • Retrieve only the minimum number and size of chunks.
  • Strip hidden metadata and unrelated document sections.
  • Avoid sending entire case files when fields or summaries suffice.
  • Use lower environments with synthetic or properly de-identified data.
  • Prevent production prompts from entering general developer logs.

Accuracy​

Allow people to correct source data and propagated embeddings or indexes. An AI disclaimer does not cure inaccurate personal data. Maintain a source-of-truth reference and reindex after correction.

Retention and deletion​

Define separate schedules for:

  • raw prompts and responses;
  • conversation history;
  • source documents;
  • extracted chunks and embeddings;
  • application and security logs;
  • evaluation datasets;
  • incident evidence;
  • backups;
  • immutable regulatory records.

Build a subject-to-store map so a verified request can find data across operational databases, object storage, search indexes, caches, analytics, feedback systems, and approved logs. Deletion should be asynchronous, idempotent, observable, and evidenced. Backups need a documented suppression or re-deletion process after restoration.

Data-subject rights​

Provide processes for access, rectification, erasure, restriction, portability where applicable, objection, and rights related to automated decision-making. The Commission lists these rights and the right to information (individual rights under GDPR).

Automated decision-making​

If the system makes or materially determines a decision about a person, assess Article 22 and related transparency, lawful basis, human intervention, explanation, and challenge requirements. Do not disguise a solely automated decision by adding a nominal reviewer.

International transfers and location​

  • Map where prompts, content, logs, support access, backups, and telemetry are processed.
  • Select Azure region and model deployment type deliberately.
  • Assess adequacy, Standard Contractual Clauses, supplementary measures, or another transfer mechanism where required.
  • Check all connectors and web-grounding features separately; a core model's boundary does not automatically cover external APIs.

Breach response​

Instrument confidentiality, integrity, and availability incidents. The Commission states that qualifying personal-data breaches must be notified to the supervisory authority without undue delay and, where feasible, within 72 hours of awareness (data-breach obligations).

The incident workflow should quickly answer what people and data were affected, which tenants, models, prompts, tools, and regions were involved, whether data left the approved boundary, how the incident was contained, and what evidence remains.

9.2 Purview's role​

Microsoft Purview can support data discovery, classification, sensitivity labels, lineage, Data Loss Prevention, audit, records, and Data Security Posture Management. Its DSPM capabilities are intended to identify sensitive-data exposure and oversharing risks across traditional and AI applications (Microsoft Purview DSPM).

Purview is not a substitute for application-level authorisation. A sensitivity label is valuable context, but the application still needs an enforceable policy that determines whether this user, for this purpose, can retrieve this content and send it to this model or tool.


10. ISO/IEC 42001: build an AI management system around the architecture​

ISO/IEC 42001:2023 specifies requirements for establishing, implementing, maintaining, and continually improving an AI management system, or AIMS (ISO/IEC 42001 overview). It governs the organisation's management of AI; it is not simply a checklist of cloud settings.

The AIMS should include:

Context and scope​

  • organisational purpose and interested parties;
  • legal, contractual, ethical, and sector requirements;
  • AIMS boundaries, business units, locations, AI systems, suppliers, and exclusions;
  • interfaces with the ISMS, privacy management, quality management, enterprise risk, and product governance.

Leadership and accountability​

  • AI policy approved by leadership;
  • named accountable executive and system owners;
  • decision rights and escalation routes;
  • AI governance committee with legal, risk, privacy, security, technical, domain, and affected-stakeholder perspectives;
  • resources and competence.

Planning and risk​

  • common AI risk method and risk appetite;
  • AI system inventory;
  • impact assessments;
  • objectives and measurable indicators;
  • treatment plans, owners, deadlines, and residual-risk acceptance;
  • change triggers requiring reassessment.

Operational controls​

  • data acquisition and quality;
  • model and supplier selection;
  • system design and verification;
  • responsible-AI testing;
  • transparency and user information;
  • human oversight;
  • third-party and customer responsibilities;
  • deployment, monitoring, incident response, and retirement.

Performance evaluation​

  • operational and responsible-AI metrics;
  • internal audit;
  • compliance monitoring;
  • management review;
  • stakeholder feedback and complaints;
  • effectiveness review of risk treatments.

Improvement​

  • nonconformity and corrective action;
  • incident learning;
  • control tuning;
  • changes to policies, training, architecture, and tests;
  • evidence that improvements are completed and effective.

Useful companion standards include ISO/IEC 23894 for AI risk management, ISO/IEC 42005 for AI system impact assessment, and ISO/IEC 42006 for bodies auditing and certifying AI management systems. Use licensed standards and qualified advisers for a formal implementation.


11. ISO/IEC 27001: integrate AI into the ISMS​

ISO/IEC 27001:2022 defines requirements for an information security management system (ISO/IEC 27001 overview). Its purpose is systematic risk management and continual improvement of confidentiality, integrity, and availability—not deployment of a fixed catalogue of technologies.

Bring the AI platform into the ISMS:

  • include prompts, embeddings, models, indexes, agent configurations, tool schemas, evaluation datasets, audit logs, and AI suppliers in the asset inventory;
  • identify asset owners and information classifications;
  • assess threats and vulnerabilities across the AI lifecycle;
  • select risk treatments and document applicability in the Statement of Applicability;
  • apply access, cryptographic, operational, supplier, development, incident, continuity, and compliance controls;
  • monitor control effectiveness;
  • audit the actual deployment and operating process;
  • include AI risks in management review and continual improvement.

Azure certifications and compliance reports can be useful supplier evidence, but they do not certify the customer's application, people, configuration, purpose, or management system.

Consider additional standards where relevant:

  • ISO/IEC 27017 for cloud-security controls;
  • ISO/IEC 27018 for protection of PII in public cloud processing;
  • ISO/IEC 27701 for privacy information management;
  • ISO 22301 for business continuity;
  • sector-specific requirements such as DORA, NIS2, financial-services rules, medical-device rules, or critical-infrastructure requirements.

12. One control system, multiple obligations​

Do not create a separate duplicated control for each framework. Create one control library with a clear objective, owner, implementation, frequency, test, evidence, exceptions, and mappings.

Unified control objectiveEU AI ActISO/IEC 42001ISO/IEC 27001GDPRAzure implementation
Maintain AI inventory and ownershipProvider/deployer accountabilityAIMS scope and operational governanceAsset inventory and ownershipRecords and accountabilityPurview catalogue, CMDB, Foundry project inventory
Assess AI and fundamental-rights riskHigh-risk risk management and impact dutiesAI risk and impact managementInformation-security riskDPIA where high riskRisk workflow linked to release gates
Govern data quality and provenanceData-governance requirementsData controlsIntegrity and supplier controlsAccuracy, minimisation, purposeADLS zones, Purview lineage, quality tests
Enforce least privilegeHuman and system controlResponsible use and resource controlsAccess controlIntegrity/confidentialityEntra, PIM, managed identity, RBAC, private endpoints
Preserve traceabilityAutomatic logs and technical documentationDocumented information and monitoringLogging and monitoringAccountability and securityOpenTelemetry, Monitor, immutable audit store
Assure intended-purpose performanceAccuracy and robustnessEvaluation and objectivesChange and testing controlsFairness and accuracyFoundry evaluation, CI gates, production monitoring
Provide meaningful human oversightHuman-oversight requirementsHuman involvement controlsSegregation and operating proceduresArticle 22 safeguards where applicableApproval UI, tool broker, override and stop controls
Manage suppliers and modelsGPAI and value-chain informationThird-party and lifecycle controlsSupplier relationshipsProcessor and transfer dutiesSupplier register, model cards, contracts, route allowlist
Respond to incidentsSerious incident and corrective actionNonconformity and improvementIncident managementBreach assessment and notificationSentinel, incident automation, kill switches, CAPA
Retain and delete appropriatelyRequired recordsDocumented-information controlRecords and secure disposalStorage limitation and rightsLifecycle policies, deletion orchestrator, backup procedures

Evidence record for each control​

Each control record should contain:

  • control ID and objective;
  • regulatory and standard mappings;
  • owner and operator;
  • in-scope systems and data;
  • implementation description;
  • automation or manual frequency;
  • test method and success threshold;
  • latest test result;
  • evidence link and retention;
  • exception, expiry, compensating control, and approver;
  • related risks, incidents, and corrective actions.

This structure turns compliance into an observable operating system rather than an annual document exercise.


13. LLMOps and secure delivery lifecycle​

Stage 0: Discover and classify​

  • Define the business outcome and baseline without AI.
  • Identify affected users and failure consequences.
  • Complete AI Act screening, privacy screening, data classification, threat model, and initial impact assessment.
  • Define success, safety, fairness, security, latency, cost, and human-oversight requirements.
  • Decide whether AI is necessary and proportionate.

Gate: accountable owner approves problem, role, risk tier, data scope, and evaluation plan.

Stage 1: Build the golden dataset​

  • Create representative, lawful, minimised test cases.
  • Include normal, difficult, multilingual, adversarial, stale-data, missing-data, and edge cases.
  • Include affected subgroups where performance differences could matter.
  • Label expected answer, sources, refusal, tool, action, and escalation.
  • Separate development, validation, and final holdout sets.

Gate: data owner and privacy or risk owner approve provenance, quality, and use.

Stage 2: Select model and design RAG​

  • Compare candidate models under the same system conditions.
  • Tune retrieval, chunking, metadata, reranking, and context assembly.
  • Establish a simple deterministic baseline.
  • Document trade-offs and model limitations.

Gate: selected release unit meets quality, safety, cost, latency, and location thresholds.

Stage 3: Implement securely​

  • Use infrastructure as code with Bicep or Terraform.
  • Use managed identities and private endpoints.
  • Implement policy enforcement and tool broker outside the model.
  • Generate SBOM and sign artefacts.
  • Add unit, integration, contract, security, tenant-isolation, and deletion tests.

Gate: architecture and security review approve the built design.

Stage 4: Evaluate and red-team​

Test:

  • task success;
  • retrieval and groundedness;
  • hallucination and appropriate refusal;
  • content safety;
  • bias and subgroup performance;
  • direct and indirect prompt injection;
  • data exfiltration;
  • tool abuse and excessive agency;
  • output injection;
  • denial of wallet;
  • cross-tenant and privilege escalation;
  • model and dependency failure;
  • disaster recovery and safe shutdown.

NIST's AI RMF and Generative AI Profile offer a useful voluntary structure for Govern, Map, Measure, and Manage activities (NIST AI RMF resources).

Gate: no unresolved critical risk; residual risks are owned, justified, time-bound, and approved.

Stage 5: Release progressively​

  • Deploy the immutable release to pre-production.
  • Run smoke, synthetic, security, quality, and performance tests.
  • Use shadow traffic only with lawful, minimised data handling.
  • Start with employees or low-risk tenants.
  • Canary a small percentage.
  • Compare live signals to the baseline.
  • Promote gradually or roll back automatically.

Gate: go-live board confirms operational, legal, privacy, security, responsible-AI, support, and business readiness.

Stage 6: Operate and improve​

  • Monitor SLOs, quality, safety, fairness, cost, and complaints.
  • Sample cases under controlled access for human review.
  • Re-run evaluation on every material change and on a schedule.
  • Review access, suppliers, models, tools, and data sources.
  • Track corrective actions and effectiveness.
  • Maintain post-market or production monitoring where applicable.

Stage 7: Retire safely​

  • Disable ingress, agents, models, tools, and credentials.
  • preserve required records and legal holds;
  • delete or return data according to contract and law;
  • remove indexes, embeddings, caches, replicas, and backups under the approved process;
  • update the inventory and user notice;
  • document residual dependencies and lessons learned.

14. CI/CD release pipeline​

The pipeline release manifest should bind together:

  • application image digest;
  • infrastructure module versions;
  • model provider, family, deployment, and version;
  • embedding model and version;
  • system and task prompts;
  • content filters and guardrail configuration;
  • retrieval index schema and corpus version;
  • chunking and enrichment code;
  • tool registry and schemas;
  • policy bundle;
  • evaluation dataset and thresholds;
  • approvals and exceptions.

Configuration in a portal that is not exported, reviewed, and versioned creates an audit and recovery gap. Prefer declarative deployment and drift detection. Where a service lacks full infrastructure-as-code coverage, export configuration and create compensating review evidence.


15. Evaluation gates and example SLOs​

Define thresholds by use case and harm, not by industry fashion.

DimensionIllustrative release thresholdProduction alert
End-to-end task successAt least 90% on approved holdout5-point decline from baseline
Citation correctnessAt least 98% for material factual claimsBelow 97% over review window
GroundednessAt least 95%Statistically significant decline
Permission-filter correctness100%Any confirmed leakage is severity 1
Correct refusalAt least 95% on unsupported requestsBelow agreed threshold
Tool selection precisionAt least 99% for write toolsAny unauthorised or wrong write
Human approval coverage100% for high-impact actionsAny bypass
Harmful-content escapeBelow risk-specific maximumCritical-category confirmed escape
p95 time to first tokenUnder 2 secondsOver 2.5 seconds for 15 minutes
End-to-end availability99.9% monthlyError-budget burn alerts
Cost per successful outcomeWithin product budget20% increase without outcome gain
Deletion completionWithin documented SLAAny overdue high-risk deletion

Do not combine all metrics into one opaque score. A high average can hide a catastrophic security or fairness failure. Use hard gates for zero-tolerance controls and separate trade-off metrics for quality, speed, and cost.

Online evaluation without creating a privacy problem​

  • Evaluate metadata and deterministic signals for every request.
  • Use automated content evaluation on a justified, minimised sample.
  • Use human review only in a restricted environment with trained reviewers.
  • Pseudonymise cases and hide unnecessary identifiers.
  • Do not let evaluators write back into production decisions without validation.
  • Track evaluator model/version because an LLM judge can drift and be biased.
  • Calibrate automated evaluation against expert human labels.

16. Cost and sustainability engineering​

Major cost drivers are:

  • model input and output tokens or provisioned throughput;
  • embeddings and re-embedding;
  • AI Search replicas, partitions, semantic ranking, and storage;
  • application compute and warm capacity;
  • Document Intelligence or OCR;
  • databases and cross-region replication;
  • monitoring and long log retention;
  • WAF, gateways, private networking, Firewall, and egress;
  • evaluation and red-team runs;
  • human review and support.

Optimise safely:

  • route simple tasks to a smaller approved model;
  • reduce irrelevant context before inference;
  • use concise system instructions and structured output;
  • cap maximum output and agent iterations;
  • batch embeddings and incremental indexing;
  • avoid re-embedding unchanged content using hashes;
  • cache public or non-personal deterministic responses only when isolation and freshness are proven;
  • separate interactive and batch capacity;
  • apply tenant and application budgets;
  • archive or aggregate high-volume telemetry after its useful window;
  • measure cost per successful outcome and cost of errors.

Do not reduce cost by removing auditability, security testing, human oversight, or data isolation without a formal risk decision.

Track energy-relevant proxies such as model size, tokens, repeated calls, agent loop count, reprocessing, and idle capacity. The smallest system that reliably achieves the intended outcome is often cheaper, faster, and easier to govern.


17. Required architecture decisions​

Record these as Architecture Decision Records.

DecisionQuestions to answer
Business purposeWhat measurable outcome justifies AI? What uses are prohibited?
Regulatory roleAre we provider, deployer, or both? Is high-risk or Article 50 in scope?
Human authorityWhich outputs are advisory, and which actions require approval?
Deployment geographyRegional, EU Data Zone, or global? What processing boundary is approved?
RuntimeFoundry Agent Service, Container Apps, or AKS, and why?
Model strategyApproved models, versions, fallback, throughput, and retirement rules?
GatewayWhich APIM tier and topology meet private networking and availability needs?
Tenant isolationShared with filters, separate index/database, separate resource, or subscription?
Data storesWhich store is authoritative for each data class and state type?
RetrievalHybrid search, reranking, ACL model, freshness, citations, deletion?
Tool policyApproved tools, delegated identity, limits, approval, idempotency, kill switch?
LoggingWhich content is never logged? What is sampled, where, by whom, and for how long?
EncryptionPlatform keys or customer-managed keys? Rotation and recovery?
ResilienceSLO, RTO, RPO, regions, capacity reserve, degradation, and failover?
EvaluationGolden sets, thresholds, subgroup tests, adversarial tests, online monitoring?
Change controlWhat triggers reclassification, DPIA update, reassessment, or conformity work?
RetirementHow are data, embeddings, agents, keys, records, and user communications handled?

18. Common anti-patterns​

“The model will follow the system prompt”​

Prompts are behavioural guidance, not an authorisation system. Enforce rules outside the model.

“RAG eliminates hallucinations”​

RAG can improve grounding, but poor retrieval, stale sources, weak chunking, prompt injection, or unsupported synthesis still produces wrong answers.

“Private endpoints make the system secure”​

Private networking reduces exposure. It does not fix excessive permissions, poisoned data, unsafe tools, weak code, compromised identities, or privacy violations.

“Azure is certified, so our application is compliant”​

Cloud attestations cover defined provider services and controls. The customer remains responsible for its purpose, configuration, data, access, integration, processes, testing, and evidence.

“We need to log every prompt for audit”​

Auditability requires relevant facts, not indiscriminate duplication of personal and confidential content. Design structured, minimised evidence and tightly controlled exception capture.

“A human approval button solves risk”​

Oversight must be informed, independent enough to challenge, resourced, measurable, and able to stop or reverse.

“Fail over to any available model”​

An untested fallback can silently change risk. Only route to approved models with validated controls and data boundaries.

“One shared vector index is always cheaper”​

It may be cheaper until a filtering defect exposes another tenant's information. Choose the isolation level based on harm, assurance, and customer commitments.

“Agents need broad permissions to be useful”​

Useful agents need well-designed tools, not administrator rights. Narrow APIs, delegated identity, scoped actions, and approvals improve both safety and product clarity.

“We can add governance after the pilot”​

If the pilot uses production data, real users, or real decisions, governance is already required. Evidence, lineage, rights handling, and safe actions are difficult to retrofit.


19. Phased implementation roadmap​

Phase 1: Governance and landing zone​

  • Establish AI inventory, intake, roles, risk method, and prohibited-use policy.
  • Define AIMS and ISMS scope and integration.
  • Deploy management groups, subscriptions, policy initiatives, central logging, security, networking, and budgets.
  • Establish Entra groups, PIM, managed identities, and separation of duties.

Phase 2: Secure platform services​

  • Deploy Front Door/WAF, API Management, private runtime, private DNS, Firewall, Key Vault, storage, databases, AI Search, and Foundry resources.
  • Disable public network paths.
  • Create deployment modules, diagnostics, backup, and recovery tests.
  • Create security and compliance dashboards.

Phase 3: Governed data and RAG​

  • Approve sources and data contracts.
  • Implement raw, processed, and index zones.
  • Add malware scan, extraction, classification, ACL metadata, lineage, quality gates, and deletion.
  • Establish retrieval and citation evaluation.

Phase 4: Orchestration and tools​

  • Implement policy-aware orchestration.
  • Add input/output guardrails and strict schemas.
  • Build the tool registry, broker, delegated identity, approvals, and kill switches.
  • Add idempotency, queueing, and safe failure behaviour.

Phase 5: LLMOps and evidence​

  • Build golden and adversarial datasets.
  • Version the complete release unit.
  • Add automated evaluation, security testing, signed artefacts, canary, rollback, and drift detection.
  • Generate technical documentation and control evidence per release.

Phase 6: Controlled rollout​

  • Start with low-risk users and read-only tools.
  • Train users and reviewers.
  • Test support, complaint, breach, incident, and human-handoff paths.
  • Expand autonomy only after evidence shows adequate quality and control effectiveness.

Phase 7: Continuous assurance​

  • Reassess risk and DPIA after material change.
  • Review model and supplier updates.
  • Perform internal audits, access reviews, red-team exercises, failover tests, and management reviews.
  • Track nonconformities, incidents, and corrective actions to closure.

20. Production readiness checklist​

Purpose and accountability​

  • Intended and prohibited purposes are documented.
  • System owner, data owner, model owner, security owner, DPO contact, and operational owner are named.
  • EU AI Act role and risk classification is approved and versioned.
  • GDPR lawful basis, privacy notice, records, and DPIA decision are complete.
  • ISO/IEC 42001 AIMS and ISO/IEC 27001 ISMS scopes include the system.
  • Sector-specific requirements are assessed.

Data​

  • Sources, licences, provenance, quality, representativeness, and authority are approved.
  • Personal and special-category data are identified and minimised.
  • Purview classification and lineage are configured where applicable.
  • Tenant and document ACLs are enforced during retrieval.
  • Retention, deletion, correction, legal hold, backup, and restoration behaviour are tested.
  • Evaluation and training use of production data has a separate approval.

Identity and network​

  • Entra authentication, MFA, Conditional Access, PIM, and access reviews are configured.
  • Services use managed identities or federation, not embedded secrets.
  • RBAC is least-privileged and environment-specific.
  • Data and AI services use private endpoints and public access is disabled.
  • Egress is deny-by-default and monitored.
  • Origin bypass of Front Door/WAF is prevented.
  • Keys rotate and Key Vault recovery controls are enabled.

Models, prompts, and retrieval​

  • Model, embedding, prompt, filters, policy, tool, and index versions are pinned in the release manifest.
  • Supplier terms, data processing, location, licence, and limitations are reviewed.
  • Fallback models and regions are pre-approved and tested.
  • Golden, subgroup, multilingual, edge, and adversarial evaluation thresholds pass.
  • Unsupported claims trigger abstention, citation warning, or escalation.
  • Retrieval permission isolation is tested with hostile cross-tenant cases.

Agents and tools​

  • Tool registry contains owner, purpose, schema, permissions, limits, and expiry.
  • Tool arguments are deterministically validated.
  • Authorisation is repeated immediately before execution.
  • Consequential actions require meaningful human approval.
  • Tool identities are least-privileged and short-lived.
  • Time, token, loop, transaction, and network limits are enforced.
  • Idempotency, rollback or compensation, and kill switches are tested.

Security​

  • Threat model covers normal application, API, cloud, LLM, RAG, model, MCP, and agent risks.
  • SAST, SCA, secret, IaC, container, DAST, and API tests pass.
  • SBOM, signed artefacts, provenance, and admission controls are present.
  • Direct and indirect prompt injection and exfiltration tests pass.
  • Cross-agent, memory, vector, and denial-of-wallet tests pass.
  • Defender and Sentinel detections have owners and response playbooks.

Reliability and operations​

  • End-to-end SLOs and error budgets are approved.
  • Capacity includes peak distributions and regional failover reserve.
  • Timeouts, retries, circuit breakers, bulkheads, dead-letter queues, and backpressure are configured.
  • Safe degradation is defined for each dependency.
  • Backup, restore, index rebuild, region failover, key incident, and kill switch are exercised.
  • On-call, vendor escalation, and incident roles are staffed.

Transparency and human oversight​

  • Users are informed that they interact with AI where required.
  • Generated content is labelled or marked where applicable.
  • Citations, limitations, uncertainty, and human-contact routes are visible.
  • Reviewers can understand, challenge, override, reverse, and stop.
  • AI literacy and reviewer training are completed and measured.
  • Complaints, corrections, appeals, and data-rights requests are operational.

Monitoring and evidence​

  • Operational, quality, safety, privacy, security, business, and cost dashboards exist.
  • Telemetry is minimised, protected, and retention-controlled.
  • Release documentation is generated from versioned sources.
  • Control tests, approvals, exceptions, and evidence are linked.
  • Production evaluation is calibrated and privacy-approved.
  • Post-market or production monitoring and corrective-action workflows are active.

Conclusion​

The most reliable way to build enterprise AI on Azure is to treat the model as one replaceable component inside a governed system.

The production boundary is not the model endpoint. It includes identity, data, retrieval, prompts, tools, human decisions, networks, operations, suppliers, and evidence. The architecture succeeds when it can remain useful under uncertainty while preventing the model from becoming the authority for access, policy, or irreversible action.

On Azure, the practical pattern is:

  1. establish a governed landing zone;
  2. keep data and AI traffic private;
  3. centralise model access through a policy-aware gateway;
  4. authorise before retrieval and before every action;
  5. use RAG with provenance, security trimming, citations, and deletion;
  6. bound agent autonomy with a hardened tool broker;
  7. version and evaluate the whole release unit;
  8. monitor operational health, quality, safety, privacy, security, cost, and outcomes;
  9. design human oversight and transparency into the user experience;
  10. generate audit evidence continuously;
  11. connect the technical controls to the AIMS, ISMS, GDPR programme, and AI Act obligations;
  12. prefer safe degradation and reversible change over uncontrolled automation.

That is the difference between an impressive AI demonstration and a trustworthy production AI system.


Primary references​

Discussion

Comments​

Share feedback or questions about this page. No account required.

Loading comments…