Enterprise AI Solution Engineering: An End-to-End Best-Practice Playbook
Enterprise AI solution engineering is not the practice of connecting a user interface to a large language model and calling the result production-ready. It is the discipline of converting a real business problem into an AI-enabled operating capability that is valuable, secure, reliable, measurable, governable and sustainable.
That requires much more than model knowledge. An AI solution engineer must work across business strategy, user experience, process design, data and knowledge architecture, AI engineering, cloud platforms, integration, cybersecurity, privacy, responsible AI, software delivery, evaluation, operations, FinOps and organizational change.
The central question is therefore not:
Which model should we use?
It is:
What business capability are we improving, what evidence will demonstrate success, what is the safest and simplest architecture that can deliver it, and how will the organization operate it responsibly at scale?
This article provides a complete reference for answering that question.
Version note: This article reflects public guidance available on 21 August 2026. Laws, standards, provider services and AI capabilities evolve quickly. Regulatory interpretations should be confirmed with qualified legal, privacy, risk and compliance specialists for the relevant jurisdiction and use case.
Executive summary
Strong AI solution engineering follows ten principles:
- Start with the business outcome, not the model.
- Establish a measurable baseline before promising value.
- Use the least autonomous solution that can solve the problem.
- Treat data, identity and integration as first-class architecture domains.
- Design for non-determinism, failure and change.
- Separate component, system, operational and business evaluation.
- Build security, privacy and responsible AI into every lifecycle stage.
- Operate prompts, models, agents, retrieval and policies as versioned production assets.
- Optimize cost per successful business outcome, not merely cost per token.
- Scale through reusable platforms, standards, multidisciplinary teams and clear accountability.
A production AI solution should not pass its release gate until the organization has evidence for five questions:
- Value: Does the solution improve a measurable business or user outcome?
- Quality: Does it complete the intended task accurately and consistently enough?
- Safety: Does it remain within defined behavioral, security, privacy and compliance boundaries?
- Operability: Can teams observe, support, recover, change and retire it?
- Accountability: Are ownership, decision rights, human oversight and residual risks explicit?
1. What AI solution engineering actually covers
AI solution engineering connects strategy to production delivery. It overlaps several disciplines but is not identical to any one of them.
| Discipline | Primary concern | AI solution engineering contribution |
|---|---|---|
| Business strategy | Competitive, operational and financial outcomes | Converts strategic priorities into a feasible AI portfolio and roadmap |
| Product management | User needs, adoption and product value | Defines the AI-enabled experience, success measures and feedback loop |
| Enterprise architecture | Organizational capabilities, standards and target state | Aligns the solution to business, data, application and technology architecture |
| Solution architecture | End-to-end system design | Owns architecture decisions, integrations, qualities and trade-offs |
| AI engineering | Models, prompts, retrieval, agents and evaluation | Implements and validates the intelligent behavior |
| Data engineering | Data pipelines, quality, lineage and access | Makes governed, reliable data and knowledge available to the AI system |
| Platform engineering | Developer experience and reusable infrastructure | Provides secure paths to build, release, observe and scale AI workloads |
| Cybersecurity and privacy | Confidentiality, integrity, availability and rights | Defines threats, controls, assurance evidence and incident response |
| Responsible AI | Fairness, transparency, accountability and human impact | Converts principles and policy into lifecycle controls and evidence |
| Delivery leadership | Scope, teams, dependencies and releases | Moves the initiative through decisions, gates and operational readiness |
| Change management | Process, role and behavioral adoption | Ensures the solution becomes a used capability rather than an unused deployment |
| FinOps | Cost visibility, accountability and optimization | Connects AI consumption to unit economics and business value |
The best AI solution engineers are therefore T-shaped. They have deep expertise in AI and architecture, but sufficient breadth to lead decisions across the full enterprise system.
2. The five transformations every AI initiative must achieve
Many unsuccessful initiatives complete only the technical transformation: a prototype exists and can generate plausible output. Production value requires five transformations.
2.1 Problem transformation
An ambiguous request such as “we need an AI assistant” must become:
- A defined user and business problem
- A bounded process or decision
- A baseline
- A measurable target
- Explicit exclusions
- Named owners
- Known constraints and risks
2.2 Information transformation
Documents and data must become governed, accessible and usable knowledge through:
- Source ownership
- Data contracts
- Quality controls
- Permissions
- Parsing and normalization
- Metadata and taxonomy
- Retrieval or feature pipelines
- Lineage, retention and deletion
2.3 Intelligence transformation
Model capability must become dependable system behavior through:
- Task decomposition
- Prompt and policy design
- Retrieval
- Tools and workflows
- Structured outputs
- Guardrails
- Evaluation
- Human oversight
2.4 Operational transformation
A demonstration must become a supportable service through:
- Version control
- Automated delivery
- Infrastructure as code
- Observability
- SLOs
- Incident response
- Capacity management
- FinOps
- Change and rollback
2.5 Organizational transformation
The system must become part of how work is performed through:
- Process redesign
- Role clarity
- Training
- Trust and transparency
- Adoption measurement
- Benefits realization
- Governance and continuous improvement
If any transformation is missing, the initiative may still produce a working prototype but will struggle to deliver durable enterprise value.
3. The end-to-end AI solution engineering lifecycle
AI delivery should be iterative, but it should not be unstructured. A practical lifecycle uses explicit evidence and decision gates.
The phases are not a one-way waterfall. New evaluation evidence, risk findings, user behavior or provider changes may send the team back to discovery or architecture. The purpose of the lifecycle is to make decisions and evidence visible.
4. Phase 1: discover the real business problem
4.1 Do not accept the proposed solution as the problem statement
A client may say:
- “We need an enterprise chatbot.”
- “We want a multi-agent platform.”
- “We need to fine-tune a model.”
- “Our competitors have a copilot.”
These are proposed solutions or motivations, not validated requirements.
A stronger problem statement is:
Operations analysts spend a median of 32 minutes locating and reconciling policy evidence across six repositories. Fifteen percent of cases are escalated because the applicable policy cannot be established confidently. The organization wants to reduce median research time by 40% while preserving source-level access controls and requiring a human decision for every regulated outcome.
This formulation identifies the process, users, baseline, target, risk boundary and control expectation.
4.2 Run structured discovery
Interview business owners, users, operations, enterprise architecture, data owners, security, privacy, legal, risk, compliance, platform teams and support teams. Different participants see different failure modes.
Use the following discovery questions.
Outcome
- What business capability or user outcome should improve?
- Why is it important now?
- Which strategic objective does it support?
- Who owns the outcome and budget?
- What would happen if the organization did nothing?
Current process
- What triggers the process?
- Which people and systems participate?
- Where are delays, rework, handoffs and errors?
- Which decisions require judgement?
- What exceptions occur?
- What evidence is retained?
Baseline
- Task volume
- Cycle time
- Cost per case
- Error or rework rate
- Escalation rate
- Customer or employee satisfaction
- Compliance incidents
- Revenue conversion or loss
Users
- Primary and secondary personas
- Knowledge and training level
- Accessibility needs
- Trust requirements
- Expected frequency of use
- Consequences of incorrect advice
Data and knowledge
- Authoritative sources
- Owners and stewards
- Formats and update frequency
- Quality and completeness
- Rights and permitted purposes
- Personal, confidential or regulated data
- Regional storage or processing restrictions
Technology
- Systems of record
- APIs and events
- Identity provider
- Network and landing-zone constraints
- Existing AI, data and integration platforms
- Observability and service-management standards
Risk
- Could the system affect rights, employment, credit, health, safety or access to services?
- What harm could follow an incorrect answer or action?
- Can the error be detected, contested and reversed?
- Which decisions must remain with a qualified person?
- Which laws, standards, policies and contracts apply?
4.3 Map the current process before designing the future process
Use a simple process model such as BPMN or SIPOC to identify:
- Inputs and outputs
- Decision points
- Rework loops
- Manual reconciliation
- Waiting time
- Control points
- System boundaries
- Ownership transfers
AI should improve the process, not merely add another interface on top of it.
4.4 Produce a discovery evidence pack
The minimum output is:
- One-page problem definition
- Current-state process
- Stakeholder map
- User journey
- Baseline metrics
- Data-source inventory
- Constraints and assumptions
- Initial risk classification
- Candidate outcome metrics
- Decision on whether further exploration is justified
Microsoft's current Cloud Adoption Framework similarly begins AI adoption with motivations, business outcomes, planning, readiness and governance rather than service selection. See Microsoft AI strategy guidance and the Cloud Adoption Framework overview.
5. Phase 2: assess AI suitability and prioritize the use case
5.1 Determine whether AI is necessary
Consider the least complex viable approach:
- Process or policy change
- Search or improved information architecture
- Business rules
- Robotic or deterministic automation
- Traditional analytics or predictive ML
- Generative AI
- Tool-using agent
- Multi-agent system
Complexity must be earned by requirements. A deterministic workflow is usually more testable, predictable and auditable than an autonomous agent. Use agentic behavior only where dynamic reasoning or tool selection creates material value.
5.2 Score the opportunity
Use a weighted score, but preserve the underlying evidence. A high numeric score cannot repair a prohibited or uncontrollable use case.
| Dimension | Example questions |
|---|---|
| Business value | Does it reduce material cost, risk or cycle time, or improve revenue or experience? |
| User desirability | Is the problem frequent and painful? Will users trust and adopt the proposed change? |
| Technical feasibility | Are model capability, data, integrations and infrastructure sufficient? |
| Data readiness | Are authoritative sources available, current, permitted and governable? |
| Risk controllability | Can important failures be detected, constrained, reviewed and reversed? |
| Strategic alignment | Does it support a priority capability or reusable platform direction? |
| Time to evidence | Can the riskiest assumptions be tested quickly? |
| Scalability | Can the pattern, data or platform be reused across additional use cases? |
An illustrative score is:
Where each factor is scored from 1 to 5 and represents value, desirability, feasibility, risk controllability, strategic fit, time to evidence and adoption readiness. Weightings should reflect the organization, not be copied blindly.
5.3 Define hypotheses
Write testable statements:
- Value hypothesis: If analysts receive cited, access-controlled policy answers inside the case-management workflow, median research time will fall by at least 25% during the pilot.
- Quality hypothesis: The solution can achieve at least 90% citation precision on the approved evaluation set.
- Adoption hypothesis: At least 70% of trained pilot users will use it weekly by week six.
- Risk hypothesis: Unauthorized document disclosure can be prevented through source-permission filtering, identity propagation and adversarial testing.
- Cost hypothesis: The solution can maintain an agreed cost per successfully resolved research task at projected volume.
Every prototype should be designed to prove or disprove hypotheses, not to maximize visual impact.
5.4 Build, buy, customize or avoid
Assess:
- Strategic differentiation
- Time to value
- Required control and customization
- Data sensitivity
- Integration complexity
- Model and region availability
- Vendor lock-in
- Exit and portability
- Operational maturity
- Total cost of ownership
- Contractual rights over prompts, data, outputs and telemetry
- Provider security, assurance and incident obligations
The decision is rarely “build everything” or “buy everything.” A common enterprise pattern is to buy foundation capability, configure or build the orchestration and controls, and retain ownership of domain data, evaluation assets, policies and integration logic.
6. Phase 3: define requirements and architecture drivers
6.1 Separate functional and nonfunctional requirements
Functional requirements describe what the system does. Nonfunctional requirements describe how well it must do it and under what constraints.
AI projects often under-specify nonfunctional requirements because demonstrations focus on visible answers. In production, qualities such as security, reliability, traceability, cost and maintainability frequently determine whether the system is acceptable.
6.2 Create measurable nonfunctional requirements
| Quality | Weak requirement | Stronger example |
|---|---|---|
| Availability | The assistant must be highly available | The service will achieve 99.9% monthly availability excluding approved maintenance |
| Latency | Responses must be fast | Under agreed load, 95% of requests will begin streaming within 1.5 seconds and complete within 8 seconds |
| Retrieval quality | Search must be relevant | The approved evaluation set will achieve Recall@10 of at least 0.92 |
| Grounding | Answers must not hallucinate | Policy answers must cite supporting evidence; unsupported material claims must remain below the approved threshold |
| Security | Data must be secure | Retrieval must enforce tenant, user and source-document authorization before context reaches the model |
| Privacy | Comply with GDPR | Personal data use will have a documented purpose and lawful basis, minimization, retention, rights and DPIA controls where required |
| Auditability | Actions must be auditable | Every consequential action will record actor, identity, evidence, tool, parameters, approval and result in tamper-resistant logs |
| Recovery | Recover quickly | Critical service will meet RTO and RPO targets validated through scheduled exercises |
| Cost | Keep cost low | Monthly cost and cost per successful task will remain within approved budget and unit-economic thresholds |
| Portability | Avoid lock-in | Model and retrieval interfaces will isolate provider-specific behavior where a realistic exit requirement exists |
The numbers above are illustrative. Targets must be based on user research, risk, projected load, business value and technical testing.
6.3 Capture architecture drivers
Architecture drivers normally include:
- Business criticality
- Number of users and tenants
- Peak concurrency
- Geographic distribution
- Data classification and residency
- Integration protocols
- Required human oversight
- Model capability
- Latency and availability
- Cost envelope
- Regulatory classification
- Provider and procurement constraints
- Existing enterprise standards
- Team capability and operating model
6.4 Use architecture decision records
An Architecture Decision Record should contain:
- Decision title and status
- Context
- Decision drivers
- Options considered
- Evidence
- Trade-offs
- Decision
- Positive and negative consequences
- Assumptions
- Risks and mitigations
- Owner and date
- Revisit triggers
Typical AI ADRs include:
- RAG versus fine-tuning
- Workflow versus agent
- Single-agent versus multi-agent
- Managed model API versus self-hosting
- One provider versus multi-provider resilience
- Vector-only versus hybrid or graph retrieval
- Shared platform versus solution-specific services
- Serverless versus Kubernetes
- Synchronous versus asynchronous processing
- Short-term conversation state versus durable memory
- Full prompt storage versus privacy-preserving telemetry
An ADR is valuable because AI models, prices, laws and provider capabilities change rapidly. It records why a decision was reasonable at the time and what evidence should trigger reconsideration.
7. Reference architecture for production enterprise AI
The following logical architecture is deliberately vendor-neutral. It can be implemented with managed cloud services, Kubernetes, serverless components or a hybrid architecture.
7.1 Experience layer
Responsibilities:
- Web, mobile, collaboration and contact-centre channels
- AI disclosure
- Citations and evidence
- Confidence-aware interaction
- Feedback and correction
- Approval and escalation
- Accessibility
- Conversation and task status
- Degraded-mode communication
The experience should make limitations and accountability understandable. Users must know when they are interacting with AI, which sources support a recommendation, what requires verification and how to challenge or escalate an outcome.
7.2 Edge and API layer
Responsibilities:
- WAF and DDoS protection
- API gateway
- Authentication
- Request validation
- Rate limits and quotas
- Tenant resolution
- Regional routing
- API versioning
- Abuse and anomaly signals
Do not rely on the model to validate access or protect the API. Standard application and cloud security controls still apply.
7.3 Identity and policy layer
Responsibilities:
- Enterprise identity federation
- Workload identities
- Role- and attribute-based authorization
- Tenant boundaries
- Purpose and consent signals
- Tool permissions
- Model and data access policy
- Step-up authentication for sensitive actions
Identity must travel with the request. A RAG application that retrieves documents without enforcing the caller's permissions can become a high-speed data-exfiltration system.
7.4 Orchestration layer
Responsibilities:
- Deterministic workflows
- Agent planning and execution
- State management
- Prompt assembly
- Tool selection
- Retry and timeout policies
- Idempotency
- Step and budget limits
- Human approval
- Compensation or rollback
Keep business rules outside natural-language prompts where deterministic enforcement is possible. Prompts guide model behavior; policy engines, code and authorization systems enforce hard constraints.
7.5 Model gateway
Responsibilities:
- Provider abstraction where justified
- Model routing
- Request normalization
- Token budgets
- Safety policies
- Caching
- Provider quotas
- Circuit breaking
- Fallback
- Usage and cost attribution
- Model-version inventory
Avoid creating a lowest-common-denominator abstraction that hides capabilities the solution needs. Isolate provider-specific code at deliberate boundaries while allowing explicit use of differentiated features.
7.6 Knowledge and retrieval layer
Responsibilities:
- Source ingestion
- Parsing and normalization
- Chunking
- Metadata enrichment
- Embeddings
- Vector, keyword and graph indexes
- Query transformation
- Permission filtering
- Reranking
- Context assembly
- Citation verification
- Freshness and deletion propagation
7.7 Tool and integration layer
Responsibilities:
- Approved tool registry
- Typed schemas
- Fine-grained credentials
- Input validation
- Transaction controls
- Idempotency keys
- Execution receipts
- Sandbox boundaries
- Timeouts and rate limits
- Side-effect classification
Separate read tools, recommendation tools and consequential action tools. The approval requirement should increase with consequence and reversibility.
7.8 Data layer
Responsibilities:
- Systems of record
- Operational databases
- Object stores
- Search and vector stores
- Graph stores
- Feature stores where relevant
- Conversation and workflow state
- Audit evidence
- Metadata, lineage and catalog
- Retention and legal hold
7.9 Evaluation and observability layer
Responsibilities:
- Offline evaluation datasets
- Experiment tracking
- Prompt, model and retriever versioning
- Distributed traces
- Quality and safety evaluation
- Token and cost metrics
- Business events
- User feedback
- Drift and anomaly detection
- Audit logging
OpenTelemetry's GenAI work uses traces, metrics and events to improve visibility into model requests, token usage, tool calls and retrieval behavior. Content collection must remain opt-in and governed because prompts, outputs and tool data can contain sensitive information. See OpenTelemetry GenAI observability.
7.10 Platform and operations layer
Responsibilities:
- Landing zone
- Network segmentation and private connectivity
- Secrets and key management
- Infrastructure as code
- CI/CD and policy as code
- Artifact registries
- Feature flags
- Environment promotion
- Backup and recovery
- SLOs and incident management
- FinOps and capacity management
Azure, AWS and Google Cloud all publish production AI or RAG architecture guidance. Their implementations differ, but they converge on governed data, secure identity, modular application services, evaluation, operational monitoring and well-architected quality attributes. See Azure Well-Architected AI workloads, the AWS Generative AI Lens and Google Cloud RAG reference architectures.
8. Data and knowledge engineering best practices
8.1 Treat source authority as a product decision
For every source, identify:
- Business owner
- Technical owner
- Authoritative status
- Classification
- Permitted purpose
- Update frequency
- Quality expectations
- Retention
- Deletion process
- Access-control model
- Geographic restrictions
Do not ingest everything simply because it is technically accessible. Larger corpora can increase cost, ambiguity, conflict and disclosure risk.
8.2 Establish data contracts
A data contract should define:
- Schema or content expectations
- Required metadata
- Ownership
- Freshness
- Quality checks
- Access-policy fields
- Change notification
- Failure behavior
- Retention and deletion
- Service-level expectations
8.3 Design ingestion for repeatability
The ingestion pipeline should be:
- Idempotent
- Incremental where possible
- Observable
- Versioned
- Retryable
- Permission-aware
- Able to quarantine invalid content
- Able to propagate updates and deletions
- Able to reconstruct an index from authoritative sources
8.4 Preserve lineage
For every retrieved unit, retain enough metadata to answer:
- Which source produced this content?
- Which source version was used?
- When was it ingested?
- Which parser and chunking version transformed it?
- Which embedding model and index version were used?
- Who was allowed to access it?
- Which answer or action used it?
8.5 Handle conflicting sources
Define precedence rules based on authority, jurisdiction, effective date and policy status. The model should not be expected to infer organizational authority reliably from contradictory text.
9. Model strategy and model selection
9.1 Select models using a task-specific scorecard
Compare models on:
- Task quality
- Structured-output reliability
- Tool-use accuracy
- Context requirements
- Latency
- Throughput and quotas
- Input and output cost
- Regional availability
- Data-use and retention terms
- Safety features
- Explainability requirements
- Customization options
- Operational support
- Portability and exit
Do not select a model because it is first on a general benchmark. Enterprise performance depends on the specific task, prompts, context, tools, language, risk and latency envelope.
9.2 Use model cascades deliberately
A practical system may use:
- A small model for classification or routing
- Embedding and reranking models for retrieval
- A stronger model for complex reasoning
- A specialist model for vision or document understanding
- A deterministic rules engine for mandatory policy
Every additional model increases evaluation, security, cost and operational scope. Add models only when the measured benefit justifies the complexity.
9.3 Separate model lifecycle from application lifecycle
Models can change independently from the application. Maintain:
- Approved-model inventory
- Model and provider assessment
- Deployment and region records
- Evaluation results
- Known limitations
- Version compatibility
- Rollback path
- Deprecation and migration plan
Model replacement should run through regression evaluation before production promotion.
10. RAG engineering best practices
RAG is a system, not a vector-database feature.
10.1 Engineer the entire retrieval lifecycle
- Acquire authoritative sources.
- Parse structure, tables, images and metadata.
- Normalize content without destroying meaning.
- Chunk using document and task semantics.
- Add source, version, date, jurisdiction and access metadata.
- Generate embeddings and lexical indexes.
- Transform or decompose the query when necessary.
- Apply tenant and user authorization.
- Retrieve candidates.
- Rerank using task-relevant signals.
- Assemble context within token and evidence budgets.
- Generate a bounded answer.
- Verify citations and apply abstention rules.
- Record feedback and evaluation evidence.
10.2 Choose retrieval patterns based on failure analysis
| Pattern | Useful when | Main trade-off |
|---|---|---|
| Vector retrieval | Semantic similarity dominates | Can miss exact identifiers and rare terms |
| Lexical retrieval | Exact names, codes or legal wording matter | Weaker on paraphrases |
| Hybrid retrieval | Both semantic and exact matching matter | More tuning and infrastructure |
| Metadata filtering | Jurisdiction, date, product or permissions matter | Requires reliable metadata |
| Reranking | Initial retrieval has good recall but poor order | Adds latency and cost |
| Query decomposition | Questions contain several sub-problems | More calls and orchestration |
| GraphRAG | Relationships and connected evidence matter | Higher modelling and operational complexity |
| Parent-child retrieval | Small chunks retrieve well but require broader context | More index and context logic |
10.3 Diagnose failures stage by stage
When an answer is wrong, ask:
- Did the correct source exist?
- Was it current and authorized?
- Was it parsed correctly?
- Did the chunk preserve the required meaning?
- Was the query transformed appropriately?
- Was the evidence retrieved?
- Was it removed by filters or reranking?
- Was it truncated during context assembly?
- Did the model ignore or contradict it?
- Did the citation point to the correct passage?
This avoids the common mistake of trying to solve every RAG problem by changing the prompt.
10.4 Evaluate retrieval separately from generation
Retrieval metrics can include:
- Recall@k
- Precision@k
- Mean reciprocal rank
- NDCG
- Permission-filter accuracy
- Source freshness
- Citation recall
Generation metrics can include:
- Correctness
- Groundedness
- Citation precision
- Completeness
- Relevance
- Abstention quality
- Style and policy compliance
Major cloud platforms now provide explicit RAG evaluation capabilities, reinforcing the need to evaluate retrieval and generation rather than relying on subjective demonstrations. See Amazon Bedrock RAG evaluation, Microsoft Foundry observability and evaluation and Google GenAI evaluation.
11. Agentic AI engineering best practices
An AI agent combines model reasoning with tools, memory, goals and the ability to affect an environment. This creates value where the task genuinely requires dynamic decisions, but it also creates additional failure paths.
AWS's 2026 Agentic AI Lens describes the shift from proving that agents can be built to running them reliably, securely and cost-effectively at scale. See the AWS Agentic AI Lens.
11.1 Decide whether an agent is justified
Use a deterministic workflow when:
- Steps and rules are stable
- Auditability requires a fixed path
- Variation is limited
- Consequences are high
- Deterministic automation can achieve the target
Consider an agent when:
- The task requires flexible sequencing
- Several tools may be selected in different orders
- The environment or information changes during execution
- The agent must interpret unstructured input before choosing a path
- The value of adaptation outweighs the additional control burden
Consider multiple agents only when role separation, scale or specialization provides measured value that cannot be obtained cleanly through one orchestrator and deterministic services.
11.2 Define an agent contract
Every production agent should have a machine-enforceable and human-readable contract:
- Purpose
- Intended users
- Allowed goals
- Prohibited goals
- Authorized tools
- Tool scopes
- Data boundaries
- Memory rules
- Maximum steps
- Token, time and financial budgets
- Approval points
- Escalation triggers
- Termination conditions
- Logging requirements
- Failure and rollback behavior
11.3 Classify tools by consequence
| Tool class | Example | Recommended control |
|---|---|---|
| Read-only public | Retrieve public product information | Validation, rate limits and logging |
| Read-only internal | Search internal policies | Identity propagation and source authorization |
| Reversible write | Create a draft case note | Scoped permission, user preview and audit |
| Consequential write | Change an account or submit a transaction | Strong authentication, policy checks and approval |
| Irreversible or high-impact | Make a payment, terminate access or send a legal notice | Human decision, separation of duties and strict limits |
The model must never receive broader credentials than the user and task require. Prefer short-lived workload credentials, fine-grained scopes and transaction-specific authorization.
11.4 Control loops and cascading failure
Implement:
- Maximum step count
- Maximum repeated tool calls
- Timeouts
- Retry budgets
- Idempotency keys
- Duplicate-action detection
- Cost and token budgets
- Circuit breakers
- Dead-letter handling
- Human escalation
- Global kill switch
An agent that retries a consequential action without idempotency can convert a transient failure into a duplicated business transaction.
11.5 Treat memory as governed data
Separate:
- Short-lived conversation state
- Workflow state
- User preferences
- Long-term semantic memory
- Audit evidence
For each, define purpose, accuracy, write authority, retention, deletion, user visibility and access policy. Do not allow untrusted content to become durable memory without validation; memory poisoning can influence future behavior long after the original interaction.
11.6 Threat-model the agent as a socio-technical system
OWASP's agentic guidance identifies risks across goals, tools, identity, memory, inter-agent communication, human trust and cascading failures. Use it alongside conventional application threat modelling rather than as a replacement. See OWASP Top 10 for Agentic Applications 2026 and the Securing Agentic Applications Guide.
12. Security engineering for AI systems
AI does not remove traditional security requirements. It adds new assets, trust boundaries and failure modes.
12.1 Identify assets
Protect:
- User and organizational data
- System prompts and policies
- Model and provider credentials
- Vector and graph indexes
- Fine-tuning datasets
- Evaluation datasets
- Agent memory
- Tool credentials
- Model artifacts and adapters
- Audit evidence
- Safety filters and policy configuration
- Business logic and proprietary knowledge
12.2 Map trust boundaries
Common boundaries include:
- User device to edge
- Edge to application
- Application to model provider
- Orchestrator to tools
- Ingestion to knowledge store
- One tenant to another
- One region to another
- Human approver to execution service
- Development to production
- Organization to external vendor
12.3 Address AI-specific threats
| Threat | Example control set |
|---|---|
| Prompt injection | Treat retrieved and user content as untrusted; isolate instructions; restrict tools; validate outputs; test adversarially |
| Sensitive-data disclosure | Minimize data; enforce source authorization; redact; apply DLP; control telemetry and retention |
| Excessive agency | Narrow tools and scopes; action limits; approvals; idempotency; kill switches |
| Insecure output handling | Schema validation; escaping; parameterized APIs; never execute raw model output |
| Model denial of service | Rate limits; quotas; request-size limits; timeouts; caching; anomaly detection |
| Supply-chain compromise | Approved artifacts; provenance; signed images; dependency scanning; SBOM; controlled model sources |
| Data or memory poisoning | Source validation; lineage; quarantines; write authorization; anomaly and integrity checks |
| Model theft or extraction | Network controls; authentication; quotas; monitoring; access restrictions |
| Cross-tenant leakage | Hard tenant partitioning; policy enforcement; security testing; tenant-aware caches and indexes |
| Tool abuse | Typed inputs; allowlisted operations; scoped credentials; policy checks; transaction receipts |
| Unsafe model update | Version pinning; regression evaluation; staged rollout; rollback |
12.4 Apply defence in depth
Do not depend on a single content filter. Combine:
- Identity and authorization
- Network controls
- Secure software development
- Data minimization
- Input and output validation
- Tool restrictions
- Policy engines
- Model guardrails
- Human oversight
- Monitoring and incident response
NIST SP 800-218A extends the Secure Software Development Framework with AI-model-specific practices. It is useful for integrating AI security into normal development and supply-chain processes. See NIST SP 800-218A.
12.5 Secure the delivery pipeline
The AI delivery pipeline should include:
- Protected branches and peer review
- Secret scanning
- Static and dynamic analysis
- Dependency and container scanning
- Infrastructure-as-code scanning
- SBOM generation
- Artifact signing and provenance
- Prompt and policy versioning
- Dataset and evaluation integrity controls
- Environment separation
- Deployment approval
- Automated regression and adversarial evaluation
- Controlled rollback
13. Privacy and data protection by design
Privacy must be addressed when defining the purpose and architecture, not after the system has already ingested personal information.
The UK ICO describes data protection by design and default as embedding appropriate technical and organizational measures during design and throughout the lifecycle. See ICO data protection by design guidance and ICO guidance on AI and data protection.
13.1 Define the processing purpose
Document:
- The specific purpose
- Categories of personal data
- Data subjects
- Sources
- Lawful basis
- Expected outputs or decisions
- Recipients
- International transfers
- Retention
- Individual rights
- Whether automated decision-making provisions may apply
13.2 Minimize data through architecture
Use:
- Pre-ingestion filtering
- Redaction or pseudonymization
- Purpose-specific indexes
- Attribute-level access
- Short-lived context
- Selective telemetry
- Regional processing
- Deletion propagation
- Provider configurations that match approved data-use terms
13.3 Conduct a DPIA when required
A Data Protection Impact Assessment should be performed where processing is likely to create high risk to individuals. It should inform design choices rather than merely document them after implementation.
Assess:
- Necessity and proportionality
- Rights and reasonable expectations
- Accuracy and fairness
- Special-category or vulnerable-person data
- Systematic monitoring or profiling
- Explainability
- Human review and contestability
- Data-sharing and provider risks
- Residual risk and approval
13.4 Design for rights and deletion
The system must be able to locate, access, correct or delete relevant personal data across:
- Source systems
- Ingestion stores
- Vector or graph indexes
- Caches
- Conversation history
- Agent memory
- Evaluation datasets
- Logs and backups, subject to lawful retention requirements
Deletion from the source without deletion from derived indexes is not a complete lifecycle design.
14. Responsible AI, governance and compliance
14.1 Governance is a system of decisions and evidence
Governance should answer:
- Which AI systems exist?
- Who owns them?
- What are they permitted to do?
- Which risks and obligations apply?
- What evidence supports approval?
- Who accepted residual risk?
- How is performance monitored?
- What changes require reassessment?
- How are incidents handled?
- How is the system retired?
14.2 Use complementary frameworks
No single framework covers everything.
| Framework or obligation | Primary contribution |
|---|---|
| NIST AI RMF | AI risk outcomes structured around Govern, Map, Measure and Manage |
| NIST GenAI Profile | GenAI-specific risks and risk-management actions |
| ISO/IEC 42001 | Organization-wide AI management system and continual improvement |
| ISO/IEC 27001 | Information-security management system and risk-based controls |
| OWASP GenAI guidance | Application and agentic threat patterns and mitigations |
| GDPR and UK GDPR | Lawful, fair and transparent personal-data processing and individual rights |
| EU AI Act | Risk-based obligations for AI providers, deployers and other operators in scope |
| Cloud well-architected frameworks | Reliability, security, operational, performance, cost and sustainability practices |
NIST's AI RMF is voluntary and organizes risk work around Govern, Map, Measure and Manage. Its GenAI Profile extends the framework for generative AI. See NIST AI RMF and the NIST Generative AI Profile.
ISO/IEC 42001 specifies requirements for establishing, implementing, maintaining and continually improving an AI management system. ISO/IEC 27001 provides the corresponding information-security management foundation. See ISO/IEC 42001 and ISO/IEC 27001.
14.3 Operationalize governance with lifecycle gates
Intake gate
Required evidence:
- Named business and technical owners
- Intended purpose and users
- Initial use-case classification
- Data categories
- Jurisdictions
- Initial risk screening
Design gate
Required evidence:
- Requirements and architecture
- Data-flow and trust-boundary diagrams
- Threat model
- Privacy assessment
- AI impact and risk assessment
- Human-oversight design
- Vendor and model assessment
Validation gate
Required evidence:
- Quality evaluation
- Safety and security evaluation
- Bias and fairness analysis where relevant
- User testing
- Control testing
- Known limitations
- Residual-risk decision
Release gate
Required evidence:
- Production-readiness review
- Monitoring and alerting
- Incident and rollback procedures
- Support ownership
- User communication and training
- Formal approval
Change gate
Trigger reassessment when:
- Model or provider changes
- Intended purpose expands
- New data is added
- Agent tools or permissions change
- A new region or user population is introduced
- Evaluation performance materially changes
- A serious incident or new threat emerges
- Law or policy changes
14.4 Maintain an AI system record
Each production AI system should have:
- System name and owner
- Intended and prohibited uses
- Users and affected people
- Architecture and data flows
- Model and provider inventory
- Data sources and purposes
- Risk classification
- Evaluation results
- Human oversight
- Controls and evidence
- Known limitations
- Incidents and corrective actions
- Approvals
- Review and retirement dates
14.5 Understand the EU AI Act timeline
The EU AI Act entered into force on 1 August 2024. According to the European Commission, it became broadly applicable on 2 August 2026, subject to exceptions and extended dates. Prohibited-practice and AI-literacy obligations applied from February 2025; governance and general-purpose AI obligations applied from August 2025; certain high-risk-system dates have been extended. Confirm the current scope and role-specific obligations for every EU-facing use case. See the European Commission AI Act overview and enforcement framework.
Do not reduce AI Act work to a checklist. Determine:
- Whether the organization is a provider, deployer, importer, distributor or another operator
- Whether the use case is prohibited, high-risk, transparency-regulated or otherwise in scope
- Which system components and vendors carry which obligations
- Required technical documentation, records, transparency, human oversight, accuracy, robustness and cybersecurity
- Post-market monitoring and incident responsibilities
This article is engineering guidance, not legal advice.
15. Evaluation: the heart of production AI quality
Traditional software is largely assessed against deterministic expected behavior. AI systems require statistical and scenario-based evidence because outputs can vary and quality is multidimensional.
15.1 Build an evaluation hierarchy
Level 1: component evaluation
Evaluate:
- Parsing
- Chunking
- Retrieval
- Reranking
- Classification
- Tool selection
- Schema adherence
- Safety filters
Level 2: end-to-end system evaluation
Evaluate:
- Task completion
- Correctness
- Groundedness
- Citation accuracy
- Completeness
- Policy compliance
- Safe refusal
- Escalation
- Tool execution
Level 3: operational evaluation
Evaluate:
- Availability
- Latency
- Throughput
- Error rates
- Recovery
- Cost
- Drift
- Abuse and security signals
Level 4: business and human evaluation
Evaluate:
- Adoption
- Time saved
- Rework
- Decision quality
- Satisfaction
- Override rate
- Complaints
- Loss or incident reduction
- Realized financial value
15.2 Create a representative golden dataset
Include:
- Typical cases
- Difficult cases
- Ambiguous cases
- Missing-information cases
- Conflicting-source cases
- Multilingual cases where relevant
- Long-context cases
- Unauthorized-access attempts
- Prompt injections
- Harmful or prohibited requests
- Cases requiring refusal or escalation
- Historical incidents
Each item should include the input, expected behavior, required evidence, acceptable variation, risk level and scoring method.
15.3 Combine evaluation methods
Use:
- Deterministic assertions
- Retrieval metrics
- Human expert review
- User testing
- Model-based evaluators with calibration
- Pairwise comparisons
- Adversarial testing
- Red teaming
- Production feedback
Model-based evaluation can scale review, but it is not an unquestionable ground truth. Calibrate evaluators against human experts, version the evaluator and retain disagreement analysis.
15.4 Define release thresholds by risk
Not every metric has equal importance. A customer-service tone score and an unauthorized disclosure test should not be averaged into one number.
Use blocking thresholds for critical controls:
- No cross-tenant disclosure in the approved security suite
- Consequential tools cannot execute without required authorization
- Required citations are present and valid
- High-risk cases escalate correctly
- Schema and transaction validation pass
Use weighted thresholds for quality dimensions where controlled variation is acceptable.
15.5 Test changes through regression evaluation
Run evaluation when changing:
- Model version
- System prompt
- Tool definition
- Retrieval strategy
- Chunking
- Embedding or reranker
- Safety policy
- Data source
- Orchestration framework
- Provider
Treat prompt, retrieval and agent changes like code changes: version, review, test, release gradually and roll back when necessary.
16. LLMOps, MLOps and GenAIOps
Production AI requires coordinated lifecycle management for code, data, prompts, models, retrieval, agents, policies and evaluation.
16.1 Version the complete behavior stack
Record:
- Application version
- Workflow or agent version
- System and task prompt versions
- Policy version
- Model and provider version
- Embedding and reranker versions
- Index and corpus version
- Tool-schema versions
- Evaluation-dataset version
- Configuration and feature flags
Without this, teams cannot reproduce incidents or compare releases reliably.
16.2 Use environment promotion
Maintain controlled development, test, staging and production environments with:
- Separate credentials and data
- Approved synthetic or masked test data
- Automated evaluation
- Security gates
- Infrastructure as code
- Change approvals appropriate to risk
- Deployment records
16.3 Use progressive delivery
Release through:
- Shadow mode
- Internal users
- Limited pilot
- Tenant or cohort canary
- Feature flags
- Gradual traffic expansion
- Automatic rollback conditions
For high-impact use cases, begin in recommendation-only mode before granting action authority.
16.4 Preserve rollback capability
Rollback may require reverting:
- Model
- Prompt
- Retriever
- Index
- Tool permission
- Feature flag
- Application version
Because data and indexes evolve, a code rollback alone may not reconstruct the prior behavior.
17. Observability, SRE and incident management
17.1 Observe the complete request
A trace should connect:
- User or service request
- Identity and tenant context
- Policy decisions
- Retrieval query and selected documents
- Model calls
- Token usage
- Tool selection and arguments
- Tool result
- Human approval
- Final response or action
- Feedback
- Cost and outcome events
Sensitive content should be redacted, sampled or excluded according to policy. Observability must not become an uncontrolled duplicate data store.
17.2 Monitor four categories
System health
- Availability
- Error rate
- Queue depth
- Resource saturation
- Provider quota
- Dependency health
- RTO and RPO evidence
Performance
- Time to first token or chunk
- End-to-end latency
- Retrieval latency
- Tool latency
- p50, p95 and p99
- Timeout and retry rates
AI quality and safety
- Task success
- Groundedness
- Citation validity
- Refusal and escalation
- Tool-call accuracy
- Safety violations
- Prompt-injection signals
- Drift
Business and cost
- Active users
- Adoption
- Successful tasks
- Time saved
- Cost per successful task
- Cost by model, tenant and feature
- Rework and override
- Benefits realized
17.3 Define SLOs and error budgets
Set SLOs for user-relevant outcomes, not only infrastructure:
- Availability
- Latency
- Successful task completion
- Retrieval freshness
- Citation validity
- Consequential-action integrity
Use error budgets to balance release velocity with reliability. A system that exhausts its reliability or safety budget should prioritize stabilization.
17.4 Prepare AI-specific incident playbooks
Create runbooks for:
- Sensitive-data disclosure
- Cross-tenant retrieval
- Harmful or discriminatory output
- Prompt-injection exploitation
- Unauthorized tool action
- Model or provider outage
- Sudden quality regression
- Cost spike or agent loop
- Poisoned data or memory
- Model deprecation
The incident process should preserve evidence, contain impact, disable affected behavior, notify required stakeholders, remediate root cause and feed findings into evaluation.
18. FinOps and AI unit economics
AI costs are usage-dependent and can grow through long contexts, repeated calls, reranking, multimodal inputs, agent loops and GPU underutilization. Financial design therefore belongs in architecture.
The FinOps Foundation recommends tracking AI-specific usage, allocating cost and aligning optimization with business value. See FinOps for AI and AI workload cost estimation.
18.1 Calculate total cost of ownership
Include:
- Discovery and design
- Application engineering
- Data preparation
- Model inference or hosting
- Embeddings and reranking
- Storage and networking
- Evaluation
- Observability
- Security and compliance
- Platform engineering
- Support and incident management
- Training and change
- Vendor and licensing cost
- Migration and exit
18.2 Measure useful units
Track:
- Cost per request
- Cost per active user
- Cost per document processed
- Cost per resolved case
- Cost per successful agent task
- Cost per approved business action
- Cost of failed and abandoned executions
- Cost by tenant, feature and model
Cost per successful outcome is more useful than cost per token because token efficiency alone says nothing about value or quality.
18.3 Build a simple economic model
An illustrative productivity calculation is:
Where:
- (V) is annual task volume
- (M) is minutes saved per task
- (C) is loaded hourly cost
- (A) is adoption
- (R) is the proportion of saved capacity that can actually be realized
Then:
Avoid presenting all time saved as cash savings. Capacity has financial value only when it is redeployed, avoids hiring, increases throughput, improves service or reduces loss.
18.4 Optimize without damaging quality
Techniques include:
- Prompt and context reduction
- Semantic, prefix and response caching
- Smaller models for routing and classification
- Model cascades
- Batching
- Asynchronous processing
- Retrieval filtering before expensive reranking
- Token and step budgets
- Quotas
- Right-sized GPU or endpoint capacity
- Autoscaling
- Request deduplication
Evaluate every optimization against quality, safety, latency and business outcomes.
19. Reliability, resilience and scale
19.1 Design for dependency failure
For each dependency, document:
- Failure modes
- User and business impact
- Detection
- Timeout
- Retry policy
- Circuit breaker
- Fallback
- Manual recovery
- Data-consistency implications
- RTO and RPO
19.2 Use graceful degradation
Possible degraded modes include:
- Search-only experience when generation is unavailable
- Read-only mode when action tools are unavailable
- Smaller approved model for non-critical tasks
- Queueing long-running work
- Human handoff
- Clearly communicated temporary limitations
Fallback must preserve safety. A cheaper or less capable model should not automatically inherit a high-risk task without equivalent validation.
19.3 Plan multi-region and data residency deliberately
Consider:
- User location
- Data classification
- Model availability by region
- Cross-region replication
- Key management
- Provider processing locations
- Recovery strategy
- Consistency
- Cost
- Regulatory and contractual restrictions
Multi-region design is not automatically active-active. Choose active-passive, active-active or regional isolation based on requirements and test the operating procedure.
19.4 Engineer multi-tenancy
For each layer, define the tenant boundary:
- Identity
- Application state
- Database
- Object storage
- Vector index
- Cache
- Prompt and configuration
- Encryption keys
- Model quota
- Logs and cost allocation
Test isolation with adversarial cross-tenant scenarios. Tenant identifiers in a prompt are not a security boundary.
20. Human oversight and user experience
20.1 Match oversight to consequence
Human oversight can include:
- Review before output is shown
- Approval before an action
- Review of exceptions
- Sampling after operation
- Real-time monitoring
- Ability to intervene or stop
- Appeal and correction
The appropriate design depends on impact, detectability, reversibility and user competence.
20.2 Avoid automation bias
Design the interface so that users can:
- See source evidence
- Understand uncertainty and limitations
- Inspect proposed actions
- Edit or reject recommendations
- Record reasons
- Escalate
Do not display invented numeric confidence simply because it appears authoritative. Use confidence only where it is meaningfully calibrated and understandable.
20.3 Design safe failure
The system should know when to:
- Ask a clarifying question
- State that evidence is insufficient
- Refuse
- Recommend human review
- Stop an agent workflow
- Preserve partial work
- Explain how the user can continue
Abstention and escalation are product capabilities, not defects.
21. Delivery leadership and multidisciplinary operating model
21.1 Establish clear accountability
Typical roles include:
- Executive sponsor
- Business owner
- Product owner
- AI solution lead
- Solution architect
- AI engineer
- Data engineer
- Application engineer
- Platform engineer
- Security architect
- Privacy and legal adviser
- Responsible AI or model-risk lead
- UX and service designer
- Change lead
- SRE or operations owner
- FinOps partner
One person may hold several roles in a small initiative, but the accountabilities must still exist.
21.2 Use RACI for work and RAPID for decisions
RACI clarifies who is Responsible, Accountable, Consulted and Informed. RAPID is useful for material decisions by defining who Recommends, Agrees, Performs, provides Input and Decides.
Important decision rights include:
- Use-case approval
- Architecture acceptance
- Model and provider approval
- Data-use approval
- Residual-risk acceptance
- Production release
- Emergency shutdown
- Scope expansion
- Retirement
21.3 Maintain a practical engagement cadence
- Weekly outcome and delivery review
- Weekly architecture and integration forum
- Weekly RAID and dependency review
- Fortnightly user demonstration
- Fortnightly evaluation and risk review
- Monthly value and FinOps review
- Formal assurance before production
- Operational review after release
The aim is not meeting volume. Each forum must have clear decisions, evidence and owners.
21.4 Lead through constraints and principles
Technical leaders should:
- Clarify the outcome
- Establish architecture principles
- Make trade-offs explicit
- Invite challenge early
- Separate facts from assumptions
- Delegate outcomes
- Protect quality gates
- Remove blockers
- Escalate material risks
- Record decisions
- Develop team capability
Leadership is demonstrated when the whole team makes better decisions, not when one architect writes every component.
22. Organizational adoption and benefits realization
22.1 Design the new operating process
Specify:
- Which tasks change
- Which decisions remain human
- New roles and skills
- Escalation
- Quality assurance
- Performance management
- Policy changes
- Support
22.2 Treat adoption as a measurable outcome
Track:
- Eligible users
- Activated users
- Weekly and monthly active users
- Repeat use
- Task completion
- Abandonment
- Trust and satisfaction
- Overrides
- Training completion
- Support demand
High login volume is not proof of value. Link adoption to process and outcome measures.
22.3 Manage the productivity dip
Users may initially become slower while learning a new process. Plan:
- Role-specific training
- Practice scenarios
- Champions
- Office hours
- Feedback channels
- Updated procedures
- Manager support
- A realistic measurement period
DORA's 2025 research emphasizes that AI tends to amplify the surrounding delivery system. Platforms, clear strategy, quality practices and feedback loops matter if individual productivity is to become organizational performance. See the 2025 DORA State of AI-Assisted Software Development.
22.4 Validate realized benefits
Compare the post-release result with the baseline:
- Did cycle time actually fall?
- Was work shifted elsewhere?
- Did rework increase?
- Was capacity redeployed?
- Did quality or customer experience improve?
- Were new risks or controls introduced?
- Did total cost match the model?
Continue, scale, redesign or retire based on evidence.
23. End-to-end enterprise case study
23.1 Scenario
A multinational financial-services organization employs 8,000 operations analysts who investigate customer cases. Analysts search policies, product rules, regulatory guidance and historical case notes across six repositories.
Current-state findings from discovery:
- Median research time is 28 minutes per case.
- Policy content is duplicated and occasionally inconsistent.
- Access rights differ by region and business unit.
- Analysts must document the evidence behind every recommendation.
- Final regulated decisions must remain with qualified employees.
- Data must remain within approved regions.
- The organization wants measurable productivity improvement without increasing conduct or privacy risk.
The proposed capability is a policy and case-resolution copilot, not an autonomous decision-maker.
23.2 Problem definition
Reduce the time required to locate and reconcile authoritative policy evidence while improving consistency and traceability. The system may retrieve, summarize and recommend next steps but may not make or execute regulated customer decisions.
23.3 Success hypotheses
- Reduce median research time by at least 25% during a controlled pilot.
- Achieve the approved citation-precision and retrieval-recall thresholds.
- Preserve source-document authorization with no cross-role or cross-tenant disclosure in the release suite.
- Ensure 100% of regulated outcomes are confirmed by an authorized analyst.
- Maintain cost per successful research task within the approved unit-economic threshold.
23.4 Options considered
Option A: improve enterprise search
Benefits:
- Lower complexity
- Predictable behavior
- Faster implementation
Limitations:
- Analysts must still reconcile evidence manually
- Limited process integration
Option B: access-controlled RAG copilot
Benefits:
- Cited synthesis
- Process integration
- Strong human control
- Moderate complexity
Limitations:
- Requires high-quality permissions, metadata and evaluation
Option C: autonomous case-resolution agent
Benefits:
- Greater theoretical automation
Limitations:
- High conduct, operational and assurance burden
- Weak alignment with the requirement for a qualified human decision
23.5 Recommendation
Select Option B. Begin with recommendation-only operation. Use deterministic workflow rules for case stages and approvals. Consider narrowly scoped actions later only after quality, control and adoption evidence justify expansion.
23.6 Logical architecture
- Analyst signs in through enterprise identity.
- The case system provides case context and regional attributes.
- API management validates the request and applies rate limits.
- Policy service resolves tenant, role, region and case purpose.
- Orchestrator classifies the request and constructs a retrieval plan.
- Retrieval applies source authorization before hybrid search.
- A reranker selects authoritative and current passages.
- Model gateway invokes the approved model with a bounded prompt.
- Output validator checks schema, citations, policy and sensitive content.
- The user receives a draft recommendation, evidence and limitations.
- Analyst edits, accepts or rejects the recommendation.
- The case system records the human decision and supporting evidence.
- Evaluation, telemetry, audit and cost events are emitted.
23.7 Key architecture decisions
- RAG rather than fine-tuning because source knowledge changes frequently and evidence must be cited.
- Hybrid retrieval because policies contain semantic concepts and exact regulatory identifiers.
- Permission filtering before context construction.
- Recommendation-only mode for the initial release.
- Stateless generation path with workflow state stored separately.
- Model gateway for quotas, approved-model routing, cost attribution and controlled fallback.
- Regional deployment aligned to data and model availability.
- No full prompt storage by default; telemetry is minimized and selectively enabled for approved debugging.
23.8 Evaluation design
The golden set contains:
- Common policy questions
- Region-specific cases
- Conflicting-policy versions
- Missing-evidence cases
- Historical incidents
- Unauthorized document requests
- Prompt-injection attempts in case notes and documents
- Cases requiring escalation
Metrics:
- Recall@k
- Citation precision and recall
- Answer correctness
- Groundedness
- Abstention quality
- Authorization accuracy
- Analyst acceptance and override
- Research time
- p95 latency
- Cost per successful research task
Critical security and approval tests are release blockers rather than averaged quality scores.
23.9 Governance evidence
- Intended-use statement
- System inventory record
- Data-flow diagram
- DPIA
- AI impact and risk assessment
- Threat model
- Provider assessment
- Human-oversight design
- Evaluation report
- Control matrix
- Incident and rollback plans
- Training and user communication
- Residual-risk approval
23.10 Delivery plan
Weeks 1-3: discovery and baseline
- Process observation
- Data and permission assessment
- Risk classification
- Baseline measurement
- Pilot group selection
Weeks 4-7: thin-slice prototype
- One region
- Two authoritative sources
- Read-only integration
- Initial evaluation set
- User testing
Weeks 8-12: controlled pilot
- Expanded source coverage
- Security and adversarial testing
- Production observability
- Training
- Shadow and recommendation-only release
Weeks 13-16: evidence and decision
- Compare outcomes with baseline
- Review quality, control, adoption and cost
- Decide whether to scale, redesign or stop
23.11 Scale decision
The solution should scale only if:
- Business outcome improves materially
- Critical controls pass
- Users adopt the workflow
- Operations can support it
- Unit economics remain viable
- Residual risk is approved
This case demonstrates the central AI solution-engineering discipline: selecting the right degree of intelligence and autonomy, then building the surrounding enterprise system required to make it dependable.
24. Delivery artefacts by lifecycle phase
| Phase | Core artefacts |
|---|---|
| Discovery | Problem statement, stakeholder map, current process, baseline, user journey |
| Prioritization | AI suitability assessment, use-case score, hypotheses, initial business case |
| Requirements | Functional requirements, NFRs, constraints, acceptance criteria |
| Architecture | HLD, integration and data-flow diagrams, ADRs, deployment view |
| Data | Source inventory, data contracts, lineage, quality and access design |
| AI design | Model scorecard, prompt and policy design, RAG or agent specification |
| Assurance | Threat model, DPIA, AI impact assessment, risk and control matrix |
| Evaluation | Golden dataset, metrics, thresholds, test and red-team reports |
| Delivery | Roadmap, backlog, RACI, RAID, release and rollback plan |
| Operations | SLOs, dashboards, alerts, runbooks, incident playbooks |
| Commercial | TCO, unit economics, benefits model, vendor assessment |
| Adoption | Training, process design, communication, adoption dashboard |
| Scale or retirement | Benefits review, technical debt, scale decision, exit plan |
25. Production-readiness checklist
Business and product
- Business owner and technical owner are named.
- Intended use and prohibited use are documented.
- Baseline and target outcomes exist.
- User journey includes failure, challenge and escalation.
- Adoption and benefits measurement are funded.
Architecture
- Functional and nonfunctional requirements are measurable.
- Data, integration, deployment and trust boundaries are documented.
- Material decisions have ADRs.
- Failure modes and degraded operation are defined.
- Regional, tenancy and portability requirements are addressed.
Data and knowledge
- Sources are authoritative and owned.
- Rights, purposes and classifications are recorded.
- Access controls propagate to retrieval.
- Quality, lineage, freshness and deletion are tested.
- Conflicting-source precedence is defined.
AI quality
- A representative golden dataset exists.
- Retrieval and generation are evaluated separately.
- Agent tools and trajectories are evaluated where relevant.
- Critical thresholds block release.
- Regression evaluation runs on behavior-changing updates.
Security and privacy
- Threat model covers conventional, LLM and agentic threats.
- Least privilege and workload identity are implemented.
- Prompt injection, data disclosure and tool abuse are tested.
- DPIA and privacy controls are complete where required.
- Telemetry follows minimization, access and retention policy.
Responsible AI and compliance
- AI risk classification is approved.
- Applicable laws, standards, policies and contractual obligations are mapped.
- Human oversight is effective and tested.
- Transparency and user communication are implemented.
- Residual risks have named acceptance.
Delivery and operations
- Code, prompts, policies, models, tools and indexes are versioned.
- CI/CD includes security and evaluation gates.
- Observability covers system, AI, business and cost signals.
- SLOs, alerts, runbooks, rollback and kill switches are tested.
- Support ownership and incident routes are clear.
FinOps and value
- Cost is allocated by use case, model, tenant or feature as required.
- Cost per successful outcome is measured.
- Quotas and budget alerts exist.
- TCO includes governance, evaluation, operations and change.
- Benefits are reviewed against the baseline after release.
26. Common anti-patterns
26.1 Model-first architecture
Symptom: The model is selected before the problem, data, risk and NFRs are understood.
Correction: Begin with outcome, process, baseline, data and architecture drivers.
26.2 Demo-driven validation
Symptom: The same small set of polished prompts is used repeatedly.
Correction: Build representative, adversarial and failure-oriented evaluation datasets.
26.3 Autonomous-by-default design
Symptom: Agents are used where a deterministic workflow would suffice.
Correction: Use the least autonomy needed and expand only with evidence.
26.4 Prompt-only control
Symptom: Security or business policy is written into the system prompt but not enforced elsewhere.
Correction: Enforce hard boundaries through identity, code, schemas, policies, tool scopes and approvals.
26.5 RAG without permission propagation
Symptom: All indexed content is equally retrievable.
Correction: Apply source-level authorization before evidence reaches the model and test cross-boundary cases.
26.6 One aggregate quality score
Symptom: Critical safety failures are hidden by strong average relevance.
Correction: Use separate dimensions and blocking thresholds for critical controls.
26.7 Full telemetry without privacy design
Symptom: Prompts, outputs and retrieved content are stored indefinitely for debugging.
Correction: Minimize, redact, sample, restrict and retain telemetry according to purpose.
26.8 Provider abstraction without a real requirement
Symptom: Significant engineering effort creates a generic layer that suppresses useful provider features.
Correction: Isolate realistic exit risks and provider-specific code without forcing false equivalence.
26.9 Cost per token as the main KPI
Symptom: The cheapest model is chosen even when task failure increases.
Correction: Optimize cost per successful, safe business outcome.
26.10 Deployment mistaken for adoption
Symptom: Success is declared when the application is released.
Correction: Measure process use, task outcomes, behavior change and realized value.
26.11 Architect as bottleneck
Symptom: Every decision and design depends on one senior engineer.
Correction: Establish principles, decision rights, reusable patterns and coaching so teams can act safely.
27. A capability model for AI solution engineers
Use five levels:
- Awareness: Can explain concepts and recognize common patterns.
- Contribution: Can produce part of the work with guidance.
- Independence: Can own an end-to-end workstream and defend decisions.
- Leadership: Can direct multidisciplinary delivery and manage trade-offs.
- Practice building: Can establish standards, mentor leaders and scale capability across engagements.
Assess the following capabilities separately:
- Business problem framing
- AI strategy and use-case prioritization
- Enterprise and solution architecture
- Data and knowledge architecture
- GenAI and model engineering
- RAG engineering
- Agentic AI
- Cloud and platform architecture
- Security and privacy
- Responsible AI and governance
- Evaluation and observability
- SRE and operational readiness
- FinOps and commercial modelling
- Delivery and stakeholder leadership
- Adoption and benefits realization
A senior AI solution engineer should be at level 4 in problem framing, architecture, assurance, delivery and stakeholder leadership while maintaining deep level-4 technical capability in at least one major production AI stack. Level 5 is demonstrated by making multiple teams and engagements successful, not by personally knowing every new framework.
28. A 90-day practice plan
Days 1-30: decision quality
- Lead or simulate three discovery interviews.
- Produce one current-state process and measurable baseline.
- Write one AI suitability assessment.
- Create three architecture options and a recommendation.
- Write five ADRs.
- Present the same recommendation to an engineer, CIO and CFO audience.
Days 31-60: production assurance
- Build a thin-slice RAG or controlled-agent solution.
- Create a golden dataset.
- Evaluate retrieval, generation, safety and tool behavior.
- Run a threat-modelling workshop.
- Produce a DPIA-style privacy assessment and AI risk register.
- Add end-to-end traces, cost allocation, SLOs and alerts.
Days 61-90: delivery leadership
- Run a production-readiness review.
- Facilitate a failure or incident exercise.
- Create the business case and unit economics.
- Produce a 90-day adoption and scale roadmap.
- Delegate one technical workstream using clear outcomes and constraints.
- Publish a reusable reference architecture or checklist for others.
29. Recommended project repository
A well-organized engagement or portfolio repository can use:
ai-solution/
├── 01-opportunity/
│ ├── problem-statement.md
│ ├── stakeholder-map.md
│ ├── current-process.md
│ └── baseline-and-outcomes.md
├── 02-strategy/
│ ├── ai-suitability.md
│ ├── options-analysis.md
│ ├── business-case.md
│ └── roadmap.md
├── 03-requirements/
│ ├── functional-requirements.md
│ ├── nonfunctional-requirements.md
│ └── acceptance-criteria.md
├── 04-architecture/
│ ├── high-level-design.md
│ ├── integration-and-data-flows.md
│ ├── deployment-view.md
│ └── decisions/
├── 05-ai-engineering/
│ ├── model-scorecard.md
│ ├── prompt-and-policy-design.md
│ ├── rag-design.md
│ └── agent-contract.md
├── 06-assurance/
│ ├── threat-model.md
│ ├── privacy-impact.md
│ ├── ai-risk-assessment.md
│ └── control-matrix.md
├── 07-evaluation/
│ ├── evaluation-plan.md
│ ├── datasets/
│ ├── release-thresholds.md
│ └── reports/
├── 08-delivery/
│ ├── raci.md
│ ├── raid.md
│ ├── release-plan.md
│ └── rollback-plan.md
├── 09-operations/
│ ├── slos.md
│ ├── monitoring.md
│ ├── runbooks/
│ └── incidents/
├── 10-finops/
│ ├── tco.md
│ ├── unit-economics.md
│ └── cost-controls.md
└── 11-adoption/
├── training.md
├── communications.md
└── benefits-realisation.md
Conclusion
Enterprise AI solution engineering is the discipline of turning uncertain model capability into a controlled organizational capability.
The strongest solutions do not begin with an LLM, RAG framework or agent platform. They begin with a measurable problem and an accountable owner. They use the simplest architecture that can achieve the outcome. They treat data and identity as foundational. They evaluate complete tasks rather than impressive responses. They enforce security and policy outside the prompt. They make human responsibility explicit. They operate models, prompts, retrieval, agents and policies as versioned production assets. They monitor quality, reliability, safety, cost and business value together.
Most importantly, they remain open to the evidence. A responsible AI solution engineer must be willing to scale a successful capability, redesign a weak one or stop an initiative that cannot produce sufficient value or control.
That is the difference between building an AI demonstration and engineering an enterprise AI solution.
Primary references
- NIST AI Risk Management Framework
- NIST AI 600-1: Generative AI Profile
- NIST SP 800-218A: Secure Software Development Practices for Generative AI
- ISO/IEC 42001:2023 AI management systems
- ISO/IEC 27001:2022 information security management systems
- European Commission: AI Act regulatory framework
- European Commission: AI Act enforcement framework
- ICO guidance on AI and data protection
- OWASP Top 10 for Agentic Applications 2026
- OWASP Securing Agentic Applications Guide
- Azure Well-Architected Framework for AI workloads
- AWS Well-Architected Generative AI Lens
- AWS Well-Architected Agentic AI Lens
- Google Cloud RAG reference architectures
- OpenTelemetry for Generative AI
- FinOps for AI
- 2025 DORA State of AI-Assisted Software Development
Discussion
Comments
Share feedback or questions about this page. No account required.
Loading comments…