Skip to main content

End-to-End AI Solution Engineering Playbook: Delivery, Change, Adoption and Operations for Banking Customer Service

· 13 min read
AI Playbook author

A approved, evaluated AI system still fails if employees distrust it, managers keep old metrics, operations lack ownership or benefits never convert to value. Delivery means establishing a reliable AI-enabled service people use correctly—not merely deploying a model.

This article is Part VII of the Banking Customer-Service AI playbook. It follows Part I through Part VI.


1. Purpose of this phase

Central question:

How should MonGo deliver, release, adopt, operate and continuously improve the AI service without losing control of customer outcomes, risk or value?

Not: deploy the AI system.
Rather: establish a reliable AI-enabled service that people use correctly, that operations can support and that continues to produce measurable value.


2. Delivery model overview

FrameworkRole
Product Operating ModelPersistent ownership of outcomes
Dual-Track AgileDiscovery and delivery in parallel
Stage-GateInvestment, risk and release decisions
DevOps / DevSecOpsControlled software delivery
MLOps / LLMOps / AgentOpsModels, prompts, retrieval, agents
ITIL 4Service-management practices
SREReliability objectives and ops engineering
ADKAR / KotterIndividual and organisational change
Benefits RealisationConvert deployment into measurable value

Agile does not bypass governance; Stage-Gate does not eliminate experimentation.


3. Product Operating Model

AI remains a product after the programme ends—continuous ownership, evaluation, prompt/model updates, knowledge maintenance, feedback, monitoring, cost and risk reassessment.

Product: MonGo Intelligent Customer Service — agent knowledge, summarisation, payment assistant, intent classification, escalation, controlled actions, evaluation/monitoring.

Vision: help customers and employees resolve routine banking needs quickly, accurately and safely while maintaining meaningful human support.

Outcomes owned: successful safe resolution, FCR, CSAT, employee adoption, grounded answer rate, escalation quality, cost per safe resolution, recontact, AHT.

Team: product director, CS product owner, AI PM, designer, architect, AI/software/data/evaluation engineers, knowledge manager, security/risk reps, change lead, ops lead, FinOps—controls integrated early, not only full-time.

Funding: outcomes, platform, improvement, stability, evaluation, controls—not a fixed feature list. Example annual objective: raise successful safe resolution 68%→80%, then test knowledge, notifications, routing, process simplification, self-service or agent assist as needed.


4–6. Agile, Scrum, Kanban

Agile principles: small increments; test assumptions early; include users; prioritised backlog; outcomes over feature count; controls in process; release on thresholds; adapt from production evidence. Start with pending/declined/reversed payments and human-support—not 150 intents.

Scrum: two-week sprints. Sprint goal example: improve retrieval and citation of current pending-payment policy. Definition of Done: code review, unit/integration/security, evaluation regression, version records, docs, monitoring, accessibility, PO acceptance, ops support updated—not “works on my laptop.” Reviews show behaviour, eval, cost/latency, feedback, risks, benefit indicators. Retrospectives that find outdated policy drive knowledge-owner review, expiry checks and KM in refinement.

Kanban: Proposed → Analysing → Ready → Building → Evaluating → Risk review → Ready for release → Released → Measuring → Closed. WIP: max 3 in evaluation, 2 in risk review, 1 major model change in production validation. When eight items queue at evaluation, “engineering productivity” is false—automate eval, add capacity, reuse test data, cut concurrent features.


7. Dual-Track Agile

Discovery: interviews, observation, journeys, assumptions, prototypes, spikes, risk, eval design — should we build it?
Delivery: software, retrieval, tools, tests, deploy, monitoring, runbooks — how do we release the validated solution?

Example: customers prefer answer + source + human option over answer-only → deliver those three. Discovery stays ahead of delivery, not months ahead; unvalidated ideas do not enter the engineering backlog.


8. Stage-Gate

Stages: Concept → Discovery → Experiment → Pilot → Production → Scale → Optimisation → Retirement.

GateEvidence highlights
ConceptProblem, purpose, sponsor, value, risk, scope, prohibitions
DiscoveryResearch, process, data readiness, prioritisation, business case, assumptions
ExperimentTest data, architecture, eval plan, security, RAI, suppliers
PilotOffline eval, red team, oversight, support, measures, rollback
ProductionPRR, risk acceptance, DPIA, Assurance Case, monitoring, IR, training, customer comms
ScalePilot benefits, stability, outcomes, risk appetite, updated TCO and legal assessment

Stage-Gate decides commitment; Agile learns and delivers inside each stage.


9–11. Lean Startup, MVP, PoC/prototype/pilot

Build–Measure–Learn: customers use the assistant but still call for trust → add timestamp, source, confirmation, human option—not more generation.

MVP includes: pending card payments, authenticated mobile, live status, approved knowledge, grounded explanation, escalation, monitoring, feedback. Excludes: voice, languages, account changes, fraud decisions, complaints, broad advice, multi-agent. Minimum ≠ no security/monitoring/escalation/evaluation/privacy.

ArtefactQuestion
PoCCan the technology work? (e.g. RAG retrieves correct policy)
PrototypeHow might the experience work?
MVPCan a limited real product create value?
PilotCan it work with limited population and operating model? (e.g. 5% mobile)

12–17. DevOps, DevSecOps, CI/CD, progressive delivery, flags

Pipeline: branch → unit → integration → security → AI evaluation → signed artefact → test → shadow → approval → progressive release.

Security as backlog: short-lived customer-bound card-action tokens; cross-customer access test must fail; every use logged; revocation supported.

CI: code, API contracts, prompts, retrieval, model regression, security, cost, structured output. Chunking change that drops retrieval recall 93%→81% is blocked even if software tests pass.

Release package: app, infra, model, prompt, retrieval, knowledge index, guardrails, eval report, security evidence, rollback config.

Rollout: internal → shadow → 1%/5%/20%/full employees → 1%/5%/20%/50%/full eligible customers. Auto-rollback if human-request detection falls 99.4%→96.8% at 5%.

Flags: customer_payment_assistant, agent_response_generation, conversation_summary, temporary_card_freeze, new_model_version, new_retrieval_pipeline.


18–19. Change Impact Assessment

StakeholderImpact
AgentsHigh — review AI info, validate sources, edit drafts, complex cases, report issues
Team leadersHigh — adoption, overrides, AI quality trends, coaching, incidents
Knowledge managersHigh — retrieval-ready content, metadata, authority, failure patterns, expiry
Risk/complianceMedium–High — impact, KRIs, model changes, challenge evidence, incidents
CustomersMedium — conversational explanations, context, handoff, AI disclosure
Technology opsHigh

Heatmap guides communication, training, support, leadership and adoption monitoring.


20. McKinsey 7S

SAlignmentMisalignment riskAction
StrategyTrusted hybrid serviceCost-only framePair CX/risk/employee metrics with cost
StructureHub platform + federated productTech owns system, CS owns nothingPersistent product ownership
SystemsKnowledge, eval, monitoring, IR, benefitsPrompts as ordinary content editsAI-specific change/eval controls
Shared valuesProtection, evidence, transparency, human controlContainment over outcomesSuccessful safe resolution primary
SkillsAI PM, eval, retrieval, RAI, KM, HAIGeneric awareness onlyRole-based pathways
StyleOpen uncertainty; willing to stop“AI replaces work” without honestyTransparent role-impact messaging
StaffPersistent team + ownersConsultant-only deliveryPartners accelerate; internal ownership

21. Kotter’s Eight Steps

  1. Urgency — demand, 11-min waits, fragmented knowledge, cost, competition—not “AI is fashionable.”
  2. Coalition — CS, CDO, CRO, tech, employees, KM, ops.
  3. Vision — easiest trusted path for routine needs with meaningful human support.
  4. Volunteer army — champions, early adopters, KM community, ops working group.
  5. Remove barriers — job-fear, knowledge, training, old metrics, UI, managers.
  6. Short-term wins — search 95→31s; fewer policy errors; higher FCR; better agent satisfaction.
  7. Sustain — payment → card → customer self-service → controlled actions via gates.
  8. Institute — roles, onboarding, procedures, dashboards, funding, risk reviews, knowledge processes.

22–24. ADKAR, training needs, agent curriculum

Awareness: first release reduces repetitive search/docs; agents remain accountable and can reject/escalate.
Desire: address job security, surveillance, autonomy, trust, transition load—co-design, feedback, non-punitive metrics.
Knowledge / Ability: sources, override, escalate, report, data rules—plus practice environment, coaching, floor support.
Reinforcement: coaching, feedback, success stories, refreshers, remove duplicate old tools.

Role matrix: agents (sources/override/escalation/data); leaders (coaching/quality/incidents); KM (metadata/quality/retrieval); engineers (eval/security/models); PO (outcomes/gates/benefits); risk (lifecycle/evidence); ops (monitor/rollback/IR); sponsor (appetite/value/accountability).

Agent modules: what it is → reviewing answers → when not to use → data protection → escalation → practice (payment/card/ambiguous/unsafe suggestions).


25–27. TAM, COM-B, nudges

TAM survey: find info faster 84%; consistency 78%; easy 74%; fits workflow 68%; trust routine 71%; trust complex 39% → accept for routine only.

COM-B: Capability (training/coaching) · Opportunity (integrate in desktop, latency, manager support, remove duplicate tools) · Motivation (honest comms, involvement, clear responsibility, demonstrate value).

Useful nudges: auto sources; outdated highlight; visible “Speak to a person”; high-risk warnings; freeze confirmation; low-confidence escalation; data-minimisation reminders. Avoid: hiding human support; preselecting AI acceptance; urgency pushing actions; punishing overrides; hiding uncertainty for containment.


28–31. Champions, CoPs, communications, customer adoption

Champions: experienced and new agents, locations, accessibility, sceptics, high performers, leaders—not only enthusiasts. CoPs: AI Product, Engineering, Responsible AI, Knowledge.

Comms: no capability exaggeration; separate decisions from exploration; role-specific; feedback channels. Matrix: executives monthly; agents weekly in rollout; managers fortnightly; risk monthly; customers every interaction; tech continuous.

Customer trust: AI disclosure; live info vs guidance; sources/timestamps; admit unavailability; easy human support; action confirmation. Prefer “immediate explanation of payment status—speak to a person anytime” over “ask our chatbot.”


32–34. Adoption measurement and Kirkpatrick

Distinguish access → trial → regular → effective → sustained → correct use. Volume ≠ productive adoption.

Appropriate reliance: accept correct guidance; correct wrong guidance; escalate high-risk; avoid inappropriate dependence—not maximum acceptance.

MetricMonth 1Month 3Target
Active agent use58%82%75%
Suggestion acceptance61%72%No fixed max
Appropriate override71%89%90%
Training completion96%99%98%
Customer completion52%68%70%
CSAT74%81%80%
Human-request success98.7%99.5%99%+

Kirkpatrick: reaction → learning (scenarios) → behaviour (override/escalation/sources/reporting) → results (errors, safe resolution, complaints, adoption, privacy incidents).


35–42. ITIL 4 practices

Service Value Chain: Plan · Improve · Engage · Design and transition · Obtain/build · Deliver and support—value co-created by customers, employees, tech, KM, risk, suppliers.

Incident: disable generation if citations fail; keep search-only—do not run full AI on fluency alone.
Major incident: commander + tech/business/risk/comms/remediation/supplier/scribe for widespread wrong info, harm, unauthorised action, data exposure, handoff failure, supplier compromise.
Problem: repeated outdated policy → expiry, ownership, single publication, index reconciliation—not one-off article deletes.
Change types: standard (spelling) · normal (prompt—eval + controlled release) · emergency (disable model—fast authority + retrospective).
Knowledge lifecycle: create → review → approve → publish → retrieve → monitor → update → retire.
Configuration: apps, models, prompts, indexes, APIs, data products, cloud, suppliers, monitoring, owners.
SLOs: customer avail 99.9%; agent 99.95%; handoff 99%; latency 95% <4s; payment API 99.5%; traces 99.99%; critical error <0.5%.


43–47. SRE, observability, continuity, DR

Error budgets: ~43 min/month at 99.9% availability; AI-quality budget e.g. critical incorrect ≤0.5%—if burned, stop features, cut traffic, require human review, remediate. Automate toil; keep human judgement for high-risk.

Trace example: ask pending → intent → API pending → PAY-014 → Model B → validator → human request → transfer → CRM summary.

Ops dashboard: reliability · AI quality · security · cost · business outcomes.

Continuity: secondary model, deterministic messages, static guidance, human routing, callback, alternate region, reduced function. Provider outage → deterministic payment messages + human for complex + keyword search for agents.

DR (illustrative): customer assistant RTO 2h / RPO 15m; agent 1h / 15m; knowledge index 4h / 1h; evaluation 24h / 24h; audit 4h / near zero.


48–54. Benefits, continuous improvement, PIR

AHT −1.2 min needs assistant and adoption, current sources, simplified workflow, removed duplicate tools, manager support, capacity reallocation—owned by Contact Centre Director.

Underperformance (only −0.4 min): old search still used, slow UI, unchanged workflow, incomplete knowledge—fix those; do not blame the model first.

PDCA / A3: e.g. abandonment 32%→18% via shorter intro, intent buttons, API latency, visible human support.

Six-month PIR:

MeasureBaselineResultTarget
Knowledge-search time95 sec31 sec35 sec
Handling time8.4 min7.4 min7.2 min
FCR68%76%78%
CSAT71%80%80%
Agent adoption0%82%75%
Critical policy error2.8%0.4%<0.5%
Safe customer resolution72%70%

Expand: card-status, callback, temporary freeze with confirmation. Do not yet: complaint decisions, financial difficulty, fraud outcomes, autonomous account changes.


55–57. Operating review and roadmap

Monthly AI service review: outcomes, quality, reliability, risk/incidents, adoption, costs, benefits, suppliers, proposed changes, scale/restrict.

Delivery gates: delivery readiness → pilot readiness → production readiness → adoption readiness → scale readiness.

Roadmap: Mobilise → Foundations → Employee pilot → Employee scale → Customer pilot → Customer scale → Controlled actions → Continuous optimisation (cost, latency, knowledge, eval, risk, retire weak capabilities).


58. Phase outputs

Product Operating Model and team; Agile/Scrum/Kanban; Dual-Track; Stage-Gate; MVP/pilot; DevOps/DevSecOps; CI/CD; progressive delivery; feature flags; Change Impact; 7S; Kotter; ADKAR; TNA and curricula; TAM/COM-B; champions and CoPs; communications; customer adoption; adoption dashboard; ITIL practices; SRE/observability; BCP/DR; benefits realisation; continuous improvement; PIR.


59. Core conclusion

Not: build the model, train users, hand to operations.

Rather a continuous lifecycle:

Discover → design → build → evaluate → approve → release → adopt → operate → measure → improve → restrict or retire.

Not: “Was the AI system launched?”

Rather:

Are customers and employees using it appropriately, is it operating reliably, are risks controlled, and is the organisation converting technical capability into measurable service value?

Next phase: Portfolio Scaling, Enterprise Transformation and Continuous Value—AI portfolio management, Lean Portfolio Management, Three Horizons, enterprise AI operating model, CoE/C4E, platform strategy, reusable patterns, capability and workforce planning, vendor ecosystem, enterprise FinOps, benefits portfolio, performance management, scaling readiness, continuous assurance and system retirement.

Discussion

Comments

Share feedback or questions about this page. No account required.

Loading comments…