Skip to main content

The Integrated 8D AI Solution Engineering Framework: Banking Customer-Service Final Playbook

· 15 min read
AI Playbook author

The 8D AI Solution Engineering Framework turns an unclear AI ambition into a valuable, secure, governed and operational service. For MonGo Bank, it transforms “build a chatbot to cut cost” into a trusted hybrid customer-service capability—and maps every framework from Parts I–VIII into one controlled learning cycle.


1. Purpose of the 8D framework

#StagePrimary questionCore output
1DefineWhat are we solving and who decides?Engagement mandate
2DiscoverWhat do users, processes and evidence tell us?Evidence-based problem definition
3DiagnoseWhat causes the problem and what prevents success?Root-cause and readiness assessment
4DesignWhat business, operating and technical solution should exist?Target solution design
5De-riskWhat could cause harm, failure or non-compliance?Controlled risk position
6DemonstrateDoes the solution work under realistic conditions?Evaluation and pilot evidence
7DecideShould we fund, scale, redesign, restrict or stop?Investment and release decision
8DeliverCan the organisation operate, adopt and improve it?Sustainable production service

Activities overlap: discovery reveals risks; evaluation changes architecture; commercial analysis changes scope; production monitoring restarts diagnosis; regulation forces redesign. Operate as a controlled learning cycle.


2. Integrated MonGo scenario

Original request: build a generative AI chatbot to reduce customer-service cost.

8D objective: create a trusted hybrid AI service that helps customers and employees resolve routine banking requests quickly, while preserving human control for sensitive, regulated and high-impact situations.

Includes: agent knowledge assist; source-grounded drafts; summarisation; intelligent routing; authenticated payment/card status; human escalation; selected reversible actions after confirmation.

Excludes: autonomous lending; final complaint adjudication; fraud liability; unrestricted financial advice; irreversible account changes; cross-customer data access.

Series map:

PartFocusPrimary 8D stages
I Strategy & DiscoveryAmbition, frameworksDefine, Discover
II Readiness & PrioritisationMaturity, scoringDiagnose, Decide
III Commercial CaseTCO, NPV, benefitsDecide
IV Architecture & Ops ModelTOM, C4, LLMOpsDesign
V Engineering & EvaluationRAG, pilots, PRRDemonstrate
VI Responsible AI & SecurityRMF, OWASP, DPIADe-risk
VII Delivery & OperationsAdoption, ITIL, SREDeliver
VIII Portfolio & ScalingLPM, CoE, FinOpsDeliver (scale)

Part I — DEFINE

Purpose: problem, outcome, scope, ownership, governance, assumptions, risks, success criteria—before uncontrolled experimentation.

Frameworks: Project Charter, AI North Star, MECE, RACI, RAPID, RAID, Stakeholder Influence–Interest, OKRs, Problem Statement Canvas, Intended-Purpose Statement.

MonGo: from “build a chatbot” to fragmented knowledge, disconnected processes, poor status visibility and weak continuity. Outcomes: successful safe resolution, lower effort, employee productivity, human support, consistency.

Artefacts: charter, problem statement, North Star, scope/exclusions, outcomes, RACI/RAPID/RAID, stakeholder map, initial risk class, intended/prohibited use, high-level roadmap.

Gate: problem without assumed solution; named business owner; explicit exclusions; measurable success; high-risk uses identified; decision rights clear; sponsor accepts mandate.


Part II — DISCOVER

Purpose: replace assumptions with evidence—customers, employees, processes, data, systems, pain points, behaviour, constraints.

Frameworks: Design Thinking, Double Diamond, VoC, JTBD, Journey Mapping, Service Blueprinting, SIPOC, VSM, BPMN, Process Mining, Assumption Mapping, Opportunity Solution Tree.

MonGo evidence: dislike of repetition; want human access; confusing payment terms; distrust generic chatbots; 95s search; 18% transfers; 12% recontact. Core problem is not “no chatbot.”

Artefacts: VoC, personas, JTBD, journeys, blueprints, SIPOC, VSM, process mining, pain-point registers, assumption map, OST, evidence repository.

Gate: representative users; observed/analysed processes; evidenced pain points; needs separated from technology; documented assumptions; first target journey defined.


Part III — DIAGNOSE

Purpose: root causes, readiness gaps, maturity, constraints, dependencies, risks, barriers to value.

Frameworks: Five Whys, Fishbone, Problem Tree, Pareto; Enterprise AI Readiness, capability/data/LLMOps/RAI/cyber/change/ops/agentic maturity; SWOT/TOWS, PESTLE, Porter, Value Chain, Capability-Based Planning.

Root cause example: inconsistent pending-payment answers → weak knowledge ownership—not “add GenAI” alone. Fix ownership, consolidation, metadata, review dates, versioning, retrieval controls.

Readiness: knowledge/eval/LLMOps/governance/AI security/product ownership lag; general cyber strong. Ready for controlled agent-assist pilot; not broad autonomous CS.

Artefacts: root-cause analysis, readiness heatmap, data/knowledge assessments, constraint register, change/ops readiness, risk/dependency map, capability gaps, remediation roadmap.

Gate: evidenced root causes; AI vs non-AI causes separated; critical gaps funded/planned; high-risk assumptions have tests; clarity on what not to automate.


Part IV — DESIGN

Purpose: future CX, process, operating model, data/app/tech architecture, AI design, governance, human oversight—a coherent target solution.

Frameworks: Strategy Choice Cascade, Playing to Win, Three Horizons, Value-Driver Tree, BMC, OMC, Wardley; TOM, AI ops hub-and-spoke, CoE/C4E, Three Lines; TOGAF, ArchiMate, C4, DDD, Event Storming, API-first, EDA, CAF, Well-Architected, Zero Trust, Data Mesh/Fabric, lakehouse/medallion, ADRs; CRISP-DM, ML/Agent lifecycles, RAG, model routing, HITL, MLOps/LLMOps/AgentOps.

Design principle: deterministic systems for banking facts; LMs for interpretation, explanation and summarisation. Status from API → policy from approved knowledge → LM explains → validator checks → human available.

Flow: authenticate → orchestrate → classify → policy rules → retrieve → tools → model gateway → structured response → validate → respond/escalate → log → telemetry.

Gate: coherent journey; named business/service owners; AI only where appropriate; SoR responsibilities explicit; meaningful oversight; fallbacks; security boundaries; trade-offs understood; build–buy–partner justified.


Part V — DE-RISK

Purpose: Responsible AI, regulatory, privacy, security, supplier, operational, customer-harm and financial risk—early and continuous.

Frameworks: NIST AI RMF, ISO 42001/23894/42005, EU AI Act, AIA, control library, cards, oversight, Assurance Case; NIST CSF, ISO 27001, Zero Trust, STRIDE, ATLAS, OWASP LLM, Secure AI SDLC; Privacy by Design, DPIA; TPRM, AI BOM, concentration risk, exit plans.

Examples: wrong pending-payment explanation → SoR API, normalisation, grounding, validation, citations, escalation, monitoring. Prompt-injection cross-customer access → session identity, ownership checks, scoped tokens, deny-by-default, redaction, immutable logs.

Gate: intended/prohibited uses; material harms assessed; legal class considered; DPIA/security complete where required; effective oversight; residual risk owned; kill switch and IR; supportable Assurance Case.


Part VI — DEMONSTRATE

Purpose: prove the complete system under realistic conditions—not demo fluency.

Frameworks: PoC, prototype, MVP, pilot; golden datasets; LLM/RAG/agent evaluation; offline/online/human eval; A/B, shadow, champion–challenger; adversarial/red team; regression; independent validation; PRR.

Golden set: standard, ambiguous, misspell, multi-intent, tool failure, complaint, vulnerability, fraud, human request, injection, conflicting knowledge.

Pilot: AHT 8.3→7.4 min; search 93→31s; FCR 69%→76%; policy error 2.8%→1.4%; CSAT 72%→78%; adoption 82%—plus 14% overrides feeding the backlog.

Gate: representative datasets; safety thresholds; escalation quality; security tests; user acceptance; ops support tested; pilot supports business case; limitations documented; rollback works.


Part VII — DECIDE

Purpose: proceed / proceed with conditions / redesign / restrict / pause / scale / stop / retire—beyond model scores.

Frameworks: DVF, Value–Feasibility–Risk, weighted scoring, portfolio matrix, WSJF, RICE, Cost of Delay, risk-adjusted value; Business Case, Five Case, TCO, ROI, NPV, IRR, payback, break-even, sensitivity, scenarios, Monte Carlo, real options, benefits dependency/realisation; RAPID, Decision Log, Stage-Gate, Investment/Design/Risk committees.

Commercial headline: ~$47M benefit / $23.1M TCO / 103.5% ROI / 2.2-year payback / 78% P(positive NPV)—still contingent on groundedness, critical errors, complaints, escalation, adoption, resilience and knowledge quality.

Proceed now: agent knowledge, summarisation, citations, intent. After enablers: payment/card status, routing. Research: vulnerability, fraud assist, agentic orchestration. Do not automate: complaint outcomes, fraud liability, creditworthiness, account closure, unrestricted advice.

Gate: measurable value; full lifecycle cost; cash vs capacity; explicit residual risk; named benefit owners; ops ownership; funded dependencies; stop criteria; documented authority.


Part VIII — DELIVER

Purpose: reliable, adopted, continuously improving service—product delivery, change, training, release, ops, reliability, benefits, scaling, retirement.

Frameworks: Product Operating Model, Agile/Scrum/Kanban, Dual-Track, Stage-Gate, Lean Startup, DevOps/DevSecOps, CI/CD, progressive delivery, flags; Change Impact, 7S, Kotter, ADKAR, TNA, TAM, COM-B, champions, CoPs, Kirkpatrick; ITIL 4, SRE, SLIs/SLOs, error budgets, IR/problem/change/KM, observability, BCP/DR, FinOps; Benefits Realisation, LPM, Three Horizons, enterprise AI ops, CoE/C4E, platform, Scaling Readiness, Continuous Assurance, retirement.

Rollout: Foundations → Employee pilot (100 agents) → Employee scale → Customer pilot (5%) → Customer scale → Controlled actions (freeze/callback/case with confirmation).

Scale gate: persistent ownership; effective (not merely high) adoption; stable SLOs; tested IR; realised benefits; acceptable CX; sustainable cost; risk in appetite; continuous assurance running.


12. VALUE quality gate

Every major deliverable: Valuable · Actionable · Logical · Understandable · Executable.

Fails: “Deploy the largest LM to improve CS.”
Passes: “Use deterministic payment APIs, hybrid retrieval and a medium LM for explanations; reserve larger models for complex employee summarisation—lower cost while staying above evaluation thresholds.”


13. Integrated stage gates

GateDecisionRequired evidence
1Approve discoveryCharter, sponsor, scope, initial risk
2Approve diagnosisUser/process evidence, assumptions
3Approve designReadiness, target journey, architecture
4Approve experimentData, controls, eval plan, supplier approval
5Approve pilotOffline eval, threat model, oversight, ops
6Approve productionAssurance Case, DPIA, monitoring, support, risk acceptance
7Approve scalePilot benefits, stable risk, cost, adoption
8Continue or retireStrategic value, quality, risk, cost, ownership

14. Framework selection guide

Decision needStart with
Problem unclearCharter, MECE, Design Thinking, Double Diamond, JTBD, Five Whys
Weak CXVoC, Journey, Blueprint, TAM, COM-B
Operational inefficiencySIPOC, VSM, BPMN, Process Mining, Fishbone, Value-Driver Tree
Organisation not readyReadiness, maturity, data, LLMOps, change, ops readiness
Prioritise use casesDVF, Value–Feasibility–Risk, weighted scoring, portfolio matrix, WSJF
Justify investmentFive Case, TCO, ROI, NPV, payback, sensitivity, benefits network
Design architectureTOGAF, ArchiMate, C4, DDD, Event Storming, API-first, Well-Architected, Zero Trust
Engineer GenAIRAG, prompts, routing, LLMOps, golden datasets, LLM/RAG eval, HITL
Agent takes actionsAgent Lifecycle, AgentOps, tool registry, least privilege, trajectory eval, kill switch
Governance evidenceAI RMF, ISO 42001/23894/42005, AIA, DPIA, System Card, Assurance Case
SecurityCSF, STRIDE, ATLAS, OWASP LLM, Zero Trust, Secure AI SDLC, AI BOM
Weak adoptionADKAR, Kotter, 7S, TAM, COM-B, Change Impact, champions, TNA
ScaleLPM, Three Horizons, AI ops model, CoE/C4E, platform, FinOps, Scaling Readiness, Continuous Assurance

15–16. RACI and RAPID

ActivitySponsorBusiness ownerProduct ownerAI engRisk/securityOpsFinance
Approve problem/scopeARCCCIC
Conduct discoveryIARCCCI
Assess readinessIARRRRC
Design solutionIARRCCI
Approve architectureICCRA/CCI
Approve risk positionICCCA/RCI
Build and evaluateICARCCI
Approve pilot / productionARRCC/RC/RC
Operate serviceIARCCRI
Realise benefitsIA/RRCICC
Approve scale / retireARR/CCCC/RR/C

Card-freeze production: Recommend AI Product Director · Agree Payments Risk, Cyber, Service Ops · Perform AI Engineering + Card Services · Input CS, Legal, Research · Decide Digital Banking Executive Sponsor.


17–18. Workplan and consulting workstreams

PhaseDurationFocus
Define2–4 weeksCharter, scope, measures, risk, governance
Discover4–8 weeksResearch, process, journeys, data, assumptions
Diagnose4–6 weeksRoot cause, readiness, gaps, prioritisation
Design6–12 weeksJourney, TOM, architecture, AI, oversight, vendors
De-riskParallelRAI, legal, privacy, security, assurance
Demonstrate8–16 weeksPoC, eval, red team, employee/customer pilots
Decide2–4 weeks / gateBusiness case, risk, investment, production, scale
DeliverContinuousRelease, training, ops, benefits, improve, retire

Seven workstreams: Strategy & value · Customer & service design · Data & knowledge · Architecture & engineering · RAI/security/privacy · Operating model & change · Delivery & operations.


19. Executive deliverables

AI strategy · Opportunity portfolio · Transformation roadmap · Business Case (Five Case + TCO/ROI/NPV) · Target Operating Model · Architecture decision pack · Responsible AI and assurance pack · Scale decision paper.


20–22. ConsultAI Lab catalogue and selection rules

Prioritise canvases that produce a decision or artefact. Full catalogues by stage: Define (Charter, North Star, RACI, RAID…) · Discover (JTBD, Journey, SIPOC, Five Whys…) · Diagnose (Readiness, maturity heatmaps…) · Design (Cascade, C4, RAG Design, Tool & Permission…) · De-risk (AI RMF, STRIDE, OWASP, cards, Assurance Case…) · Demonstrate (Golden Dataset Planner, eval scorecards, Pilot Design, PRR…) · Decide (DVF, TCO/ROI/NPV, Decision Paper, Stage-Gate…) · Deliver (Roadmap, ADKAR, SLOs, Adoption/Benefits dashboards, Retirement Checklist…).

Guided notes (not full canvases) for TOGAF, ArchiMate, DDD, ITIL, SRE, FinOps, LLMOps, AgentOps, ISO standards, ATLAS, TAM, 7S, Kotter, LPM, sales frameworks.

Example sequences:

  • Unclear opportunity: Problem Statement → MECE → JTBD → Five Whys → AI Canvas → DVF
  • Customer-service AI: Journey → Blueprint → SIPOC → VSM → Readiness → RAG Design → Oversight → Value–Feasibility–Risk → Business Case → PRR
  • Agentic AI: Purpose → Agentic Readiness → Tools/Permissions → STRIDE → AI RMF → Oversight → Agent Eval → Assurance Case → PRR
  • AI strategy: North Star → Cascade → Value Chain → Capability Map → Three Horizons → Portfolio → Operating Model → Benefits Roadmap

23. Minimum artefact map

StageMinimum artefacts
DefineCharter, North Star, scope, RACI, RAID
DiscoverCustomer evidence, journey, blueprint, process map
DiagnoseRoot causes, readiness, capability gaps, assumptions
DesignFuture journey, operating model, architecture, AI design
De-riskRisk assessment, DPIA, threat model, controls, assurance
DemonstrateGolden dataset, evaluation, red team, pilot evidence
DecideBusiness Case, scorecard, risk acceptance, decision paper
DeliverRoadmap, training, ops, adoption, benefits, monitoring

24–26. Narrative, principles, definition

MonGo needed better knowledge ownership, process redesign, secure integration, controlled generation, escalation, evaluation, monitoring, adoption and benefits ownership—not a chatbot. The result is a hybrid customer-service operating capability.

  1. Begin with the outcome, not the model.
  2. Do not automate a broken process without diagnosing it.
  3. Separate deterministic facts from generative explanations.
  4. Treat data and knowledge quality as product capabilities.
  5. Test the riskiest assumptions first.
  6. Use risk to determine governance intensity.
  7. Evaluate the complete system, not only the model.
  8. Design meaningful human oversight.
  9. Include full lifecycle cost.
  10. Make benefits someone’s responsibility.
  11. Release progressively and preserve rollback.
  12. Treat AI as a persistent product and service.
  13. Build reusable platforms without centralising all business ownership.
  14. Measure customer, business, technical and risk outcomes together.
  15. Stop or retire AI when value, ownership or risk is no longer acceptable.

Definition: End-to-end AI solution engineering is the disciplined conversion of a business or customer problem into a valuable, feasible, secure, responsible and operable AI-enabled service through evidence-based discovery, architecture, engineering, governance, evaluation, commercial decision-making, adoption and continuous improvement.

The AI solution engineer connects strategy, customers, processes, data, models, architecture, security, regulation, commercial value, operating models, people, delivery and operations. The goal is not the most sophisticated system—it is the most appropriate system the organisation can trust, fund, operate, govern, improve and stop when necessary.

Discussion

Comments

Share feedback or questions about this page. No account required.

Loading comments…