End-to-End AI Solution Engineering Playbook: Delivery, Change, Adoption and Operations for Banking Customer Service
A approved, evaluated AI system still fails if employees distrust it, managers keep old metrics, operations lack ownership or benefits never convert to value. Delivery means establishing a reliable AI-enabled service people use correctly—not merely deploying a model.
This article is Part VII of the Banking Customer-Service AI playbook. It follows Part I through Part VI.
1. Purpose of this phase
Central question:
How should MonGo deliver, release, adopt, operate and continuously improve the AI service without losing control of customer outcomes, risk or value?
Not: deploy the AI system.
Rather: establish a reliable AI-enabled service that people use correctly, that operations can support and that continues to produce measurable value.
2. Delivery model overview
| Framework | Role |
|---|---|
| Product Operating Model | Persistent ownership of outcomes |
| Dual-Track Agile | Discovery and delivery in parallel |
| Stage-Gate | Investment, risk and release decisions |
| DevOps / DevSecOps | Controlled software delivery |
| MLOps / LLMOps / AgentOps | Models, prompts, retrieval, agents |
| ITIL 4 | Service-management practices |
| SRE | Reliability objectives and ops engineering |
| ADKAR / Kotter | Individual and organisational change |
| Benefits Realisation | Convert deployment into measurable value |
Agile does not bypass governance; Stage-Gate does not eliminate experimentation.
3. Product Operating Model
AI remains a product after the programme ends—continuous ownership, evaluation, prompt/model updates, knowledge maintenance, feedback, monitoring, cost and risk reassessment.
Product: MonGo Intelligent Customer Service — agent knowledge, summarisation, payment assistant, intent classification, escalation, controlled actions, evaluation/monitoring.
Vision: help customers and employees resolve routine banking needs quickly, accurately and safely while maintaining meaningful human support.
Outcomes owned: successful safe resolution, FCR, CSAT, employee adoption, grounded answer rate, escalation quality, cost per safe resolution, recontact, AHT.
Team: product director, CS product owner, AI PM, designer, architect, AI/software/data/evaluation engineers, knowledge manager, security/risk reps, change lead, ops lead, FinOps—controls integrated early, not only full-time.
Funding: outcomes, platform, improvement, stability, evaluation, controls—not a fixed feature list. Example annual objective: raise successful safe resolution 68%→80%, then test knowledge, notifications, routing, process simplification, self-service or agent assist as needed.
4–6. Agile, Scrum, Kanban
Agile principles: small increments; test assumptions early; include users; prioritised backlog; outcomes over feature count; controls in process; release on thresholds; adapt from production evidence. Start with pending/declined/reversed payments and human-support—not 150 intents.
Scrum: two-week sprints. Sprint goal example: improve retrieval and citation of current pending-payment policy. Definition of Done: code review, unit/integration/security, evaluation regression, version records, docs, monitoring, accessibility, PO acceptance, ops support updated—not “works on my laptop.” Reviews show behaviour, eval, cost/latency, feedback, risks, benefit indicators. Retrospectives that find outdated policy drive knowledge-owner review, expiry checks and KM in refinement.
Kanban: Proposed → Analysing → Ready → Building → Evaluating → Risk review → Ready for release → Released → Measuring → Closed. WIP: max 3 in evaluation, 2 in risk review, 1 major model change in production validation. When eight items queue at evaluation, “engineering productivity” is false—automate eval, add capacity, reuse test data, cut concurrent features.
7. Dual-Track Agile
Discovery: interviews, observation, journeys, assumptions, prototypes, spikes, risk, eval design — should we build it?
Delivery: software, retrieval, tools, tests, deploy, monitoring, runbooks — how do we release the validated solution?
Example: customers prefer answer + source + human option over answer-only → deliver those three. Discovery stays ahead of delivery, not months ahead; unvalidated ideas do not enter the engineering backlog.
8. Stage-Gate
Stages: Concept → Discovery → Experiment → Pilot → Production → Scale → Optimisation → Retirement.
| Gate | Evidence highlights |
|---|---|
| Concept | Problem, purpose, sponsor, value, risk, scope, prohibitions |
| Discovery | Research, process, data readiness, prioritisation, business case, assumptions |
| Experiment | Test data, architecture, eval plan, security, RAI, suppliers |
| Pilot | Offline eval, red team, oversight, support, measures, rollback |
| Production | PRR, risk acceptance, DPIA, Assurance Case, monitoring, IR, training, customer comms |
| Scale | Pilot benefits, stability, outcomes, risk appetite, updated TCO and legal assessment |
Stage-Gate decides commitment; Agile learns and delivers inside each stage.
9–11. Lean Startup, MVP, PoC/prototype/pilot
Build–Measure–Learn: customers use the assistant but still call for trust → add timestamp, source, confirmation, human option—not more generation.
MVP includes: pending card payments, authenticated mobile, live status, approved knowledge, grounded explanation, escalation, monitoring, feedback. Excludes: voice, languages, account changes, fraud decisions, complaints, broad advice, multi-agent. Minimum ≠ no security/monitoring/escalation/evaluation/privacy.
| Artefact | Question |
|---|---|
| PoC | Can the technology work? (e.g. RAG retrieves correct policy) |
| Prototype | How might the experience work? |
| MVP | Can a limited real product create value? |
| Pilot | Can it work with limited population and operating model? (e.g. 5% mobile) |
12–17. DevOps, DevSecOps, CI/CD, progressive delivery, flags
Pipeline: branch → unit → integration → security → AI evaluation → signed artefact → test → shadow → approval → progressive release.
Security as backlog: short-lived customer-bound card-action tokens; cross-customer access test must fail; every use logged; revocation supported.
CI: code, API contracts, prompts, retrieval, model regression, security, cost, structured output. Chunking change that drops retrieval recall 93%→81% is blocked even if software tests pass.
Release package: app, infra, model, prompt, retrieval, knowledge index, guardrails, eval report, security evidence, rollback config.
Rollout: internal → shadow → 1%/5%/20%/full employees → 1%/5%/20%/50%/full eligible customers. Auto-rollback if human-request detection falls 99.4%→96.8% at 5%.
Flags: customer_payment_assistant, agent_response_generation, conversation_summary, temporary_card_freeze, new_model_version, new_retrieval_pipeline.
18–19. Change Impact Assessment
| Stakeholder | Impact |
|---|---|
| Agents | High — review AI info, validate sources, edit drafts, complex cases, report issues |
| Team leaders | High — adoption, overrides, AI quality trends, coaching, incidents |
| Knowledge managers | High — retrieval-ready content, metadata, authority, failure patterns, expiry |
| Risk/compliance | Medium–High — impact, KRIs, model changes, challenge evidence, incidents |
| Customers | Medium — conversational explanations, context, handoff, AI disclosure |
| Technology ops | High |
Heatmap guides communication, training, support, leadership and adoption monitoring.
20. McKinsey 7S
| S | Alignment | Misalignment risk | Action |
|---|---|---|---|
| Strategy | Trusted hybrid service | Cost-only frame | Pair CX/risk/employee metrics with cost |
| Structure | Hub platform + federated product | Tech owns system, CS owns nothing | Persistent product ownership |
| Systems | Knowledge, eval, monitoring, IR, benefits | Prompts as ordinary content edits | AI-specific change/eval controls |
| Shared values | Protection, evidence, transparency, human control | Containment over outcomes | Successful safe resolution primary |
| Skills | AI PM, eval, retrieval, RAI, KM, HAI | Generic awareness only | Role-based pathways |
| Style | Open uncertainty; willing to stop | “AI replaces work” without honesty | Transparent role-impact messaging |
| Staff | Persistent team + owners | Consultant-only delivery | Partners accelerate; internal ownership |
21. Kotter’s Eight Steps
- Urgency — demand, 11-min waits, fragmented knowledge, cost, competition—not “AI is fashionable.”
- Coalition — CS, CDO, CRO, tech, employees, KM, ops.
- Vision — easiest trusted path for routine needs with meaningful human support.
- Volunteer army — champions, early adopters, KM community, ops working group.
- Remove barriers — job-fear, knowledge, training, old metrics, UI, managers.
- Short-term wins — search 95→31s; fewer policy errors; higher FCR; better agent satisfaction.
- Sustain — payment → card → customer self-service → controlled actions via gates.
- Institute — roles, onboarding, procedures, dashboards, funding, risk reviews, knowledge processes.
22–24. ADKAR, training needs, agent curriculum
Awareness: first release reduces repetitive search/docs; agents remain accountable and can reject/escalate.
Desire: address job security, surveillance, autonomy, trust, transition load—co-design, feedback, non-punitive metrics.
Knowledge / Ability: sources, override, escalate, report, data rules—plus practice environment, coaching, floor support.
Reinforcement: coaching, feedback, success stories, refreshers, remove duplicate old tools.
Role matrix: agents (sources/override/escalation/data); leaders (coaching/quality/incidents); KM (metadata/quality/retrieval); engineers (eval/security/models); PO (outcomes/gates/benefits); risk (lifecycle/evidence); ops (monitor/rollback/IR); sponsor (appetite/value/accountability).
Agent modules: what it is → reviewing answers → when not to use → data protection → escalation → practice (payment/card/ambiguous/unsafe suggestions).
25–27. TAM, COM-B, nudges
TAM survey: find info faster 84%; consistency 78%; easy 74%; fits workflow 68%; trust routine 71%; trust complex 39% → accept for routine only.
COM-B: Capability (training/coaching) · Opportunity (integrate in desktop, latency, manager support, remove duplicate tools) · Motivation (honest comms, involvement, clear responsibility, demonstrate value).
Useful nudges: auto sources; outdated highlight; visible “Speak to a person”; high-risk warnings; freeze confirmation; low-confidence escalation; data-minimisation reminders. Avoid: hiding human support; preselecting AI acceptance; urgency pushing actions; punishing overrides; hiding uncertainty for containment.
28–31. Champions, CoPs, communications, customer adoption
Champions: experienced and new agents, locations, accessibility, sceptics, high performers, leaders—not only enthusiasts. CoPs: AI Product, Engineering, Responsible AI, Knowledge.
Comms: no capability exaggeration; separate decisions from exploration; role-specific; feedback channels. Matrix: executives monthly; agents weekly in rollout; managers fortnightly; risk monthly; customers every interaction; tech continuous.
Customer trust: AI disclosure; live info vs guidance; sources/timestamps; admit unavailability; easy human support; action confirmation. Prefer “immediate explanation of payment status—speak to a person anytime” over “ask our chatbot.”
32–34. Adoption measurement and Kirkpatrick
Distinguish access → trial → regular → effective → sustained → correct use. Volume ≠ productive adoption.
Appropriate reliance: accept correct guidance; correct wrong guidance; escalate high-risk; avoid inappropriate dependence—not maximum acceptance.
| Metric | Month 1 | Month 3 | Target |
|---|---|---|---|
| Active agent use | 58% | 82% | 75% |
| Suggestion acceptance | 61% | 72% | No fixed max |
| Appropriate override | 71% | 89% | 90% |
| Training completion | 96% | 99% | 98% |
| Customer completion | 52% | 68% | 70% |
| CSAT | 74% | 81% | 80% |
| Human-request success | 98.7% | 99.5% | 99%+ |
Kirkpatrick: reaction → learning (scenarios) → behaviour (override/escalation/sources/reporting) → results (errors, safe resolution, complaints, adoption, privacy incidents).
35–42. ITIL 4 practices
Service Value Chain: Plan · Improve · Engage · Design and transition · Obtain/build · Deliver and support—value co-created by customers, employees, tech, KM, risk, suppliers.
Incident: disable generation if citations fail; keep search-only—do not run full AI on fluency alone.
Major incident: commander + tech/business/risk/comms/remediation/supplier/scribe for widespread wrong info, harm, unauthorised action, data exposure, handoff failure, supplier compromise.
Problem: repeated outdated policy → expiry, ownership, single publication, index reconciliation—not one-off article deletes.
Change types: standard (spelling) · normal (prompt—eval + controlled release) · emergency (disable model—fast authority + retrospective).
Knowledge lifecycle: create → review → approve → publish → retrieve → monitor → update → retire.
Configuration: apps, models, prompts, indexes, APIs, data products, cloud, suppliers, monitoring, owners.
SLOs: customer avail 99.9%; agent 99.95%; handoff 99%; latency 95% <4s; payment API 99.5%; traces 99.99%; critical error <0.5%.
43–47. SRE, observability, continuity, DR
Error budgets: ~43 min/month at 99.9% availability; AI-quality budget e.g. critical incorrect ≤0.5%—if burned, stop features, cut traffic, require human review, remediate. Automate toil; keep human judgement for high-risk.
Trace example: ask pending → intent → API pending → PAY-014 → Model B → validator → human request → transfer → CRM summary.
Ops dashboard: reliability · AI quality · security · cost · business outcomes.
Continuity: secondary model, deterministic messages, static guidance, human routing, callback, alternate region, reduced function. Provider outage → deterministic payment messages + human for complex + keyword search for agents.
DR (illustrative): customer assistant RTO 2h / RPO 15m; agent 1h / 15m; knowledge index 4h / 1h; evaluation 24h / 24h; audit 4h / near zero.
48–54. Benefits, continuous improvement, PIR
AHT −1.2 min needs assistant and adoption, current sources, simplified workflow, removed duplicate tools, manager support, capacity reallocation—owned by Contact Centre Director.
Underperformance (only −0.4 min): old search still used, slow UI, unchanged workflow, incomplete knowledge—fix those; do not blame the model first.
PDCA / A3: e.g. abandonment 32%→18% via shorter intro, intent buttons, API latency, visible human support.
Six-month PIR:
| Measure | Baseline | Result | Target |
|---|---|---|---|
| Knowledge-search time | 95 sec | 31 sec | 35 sec |
| Handling time | 8.4 min | 7.4 min | 7.2 min |
| FCR | 68% | 76% | 78% |
| CSAT | 71% | 80% | 80% |
| Agent adoption | 0% | 82% | 75% |
| Critical policy error | 2.8% | 0.4% | <0.5% |
| Safe customer resolution | — | 72% | 70% |
Expand: card-status, callback, temporary freeze with confirmation. Do not yet: complaint decisions, financial difficulty, fraud outcomes, autonomous account changes.
55–57. Operating review and roadmap
Monthly AI service review: outcomes, quality, reliability, risk/incidents, adoption, costs, benefits, suppliers, proposed changes, scale/restrict.
Delivery gates: delivery readiness → pilot readiness → production readiness → adoption readiness → scale readiness.
Roadmap: Mobilise → Foundations → Employee pilot → Employee scale → Customer pilot → Customer scale → Controlled actions → Continuous optimisation (cost, latency, knowledge, eval, risk, retire weak capabilities).
58. Phase outputs
Product Operating Model and team; Agile/Scrum/Kanban; Dual-Track; Stage-Gate; MVP/pilot; DevOps/DevSecOps; CI/CD; progressive delivery; feature flags; Change Impact; 7S; Kotter; ADKAR; TNA and curricula; TAM/COM-B; champions and CoPs; communications; customer adoption; adoption dashboard; ITIL practices; SRE/observability; BCP/DR; benefits realisation; continuous improvement; PIR.
59. Core conclusion
Not: build the model, train users, hand to operations.
Rather a continuous lifecycle:
Discover → design → build → evaluate → approve → release → adopt → operate → measure → improve → restrict or retire.
Not: “Was the AI system launched?”
Rather:
Are customers and employees using it appropriately, is it operating reliably, are risks controlled, and is the organisation converting technical capability into measurable service value?
Next phase: Portfolio Scaling, Enterprise Transformation and Continuous Value—AI portfolio management, Lean Portfolio Management, Three Horizons, enterprise AI operating model, CoE/C4E, platform strategy, reusable patterns, capability and workforce planning, vendor ecosystem, enterprise FinOps, benefits portfolio, performance management, scaling readiness, continuous assurance and system retirement.
Discussion
Comments
Share feedback or questions about this page. No account required.
Loading comments…