Skip to main content

End-to-End AI Solution Engineering Playbook: Responsible AI, Governance, Security and Privacy for Banking Customer Service

· 15 min read
AI Playbook author

Before production, MonGo must show the hybrid customer-service AI is lawful, fair, secure, accountable, controllable and fit for purpose. The goal is not zero risk—it is proportionate controls, measured residual risk and evidence for every decision to continue, restrict or stop.

This article is Part VI of the Banking Customer-Service AI playbook. It follows Part I through Part V.


1. Purpose of this phase

Central question:

Can the bank demonstrate that the AI system is lawful, fair, secure, accountable, controllable and appropriate for its intended purpose?

Understand risks → decide what is acceptable → implement proportionate controls → measure them → escalate or stop when risk exceeds tolerance → produce evidence of what was decided and why.


2. Governance and assurance stack

FrameworkPrimary role
NIST AI RMFOrganise AI risk-management outcomes
NIST Generative AI ProfileGenerative-AI-specific risks
ISO/IEC 42001AI management system
ISO/IEC 23894AI risk-management process
ISO/IEC 42005AI-system impact assessment
EU AI ActLegal classification and obligations where applicable
UK DP law / ICO guidancePersonal-data processing for UK deployment
NIST CSF 2.0Cybersecurity outcomes
STRIDESystem-level threats
MITRE ATLASAdversarial behaviour against AI
OWASP Top 10 for LLMsCommon LLM-application vulnerabilities
DPIAHigh-risk personal-data processing
Model / Data / System CardsComponent and limitation documentation
AI Assurance CaseTrust claims linked to evidence

3. Governance principles

  1. Purpose limitation — payment assistant may explain status, retrieve guidance, show next steps, transfer to a person; not fraud liability, borrowing recommendations, creditworthiness, account closure or investment advice.
  2. Proportionality — controls scale with harm, reversibility, scale, autonomy, data sensitivity and detectability.
  3. Accountability — named owners for outcome, technology, data, knowledge, risk acceptance, operations, benefits and suppliers; the model provider does not own MonGo’s service outcome.
  4. Traceability — reconstruct query, system version, model, prompt, documents, tools, controls, human intervention and outcome.
  5. Human agency — AI disclosure, human support, reject recommendations, correct information, challenge outcomes.
  6. Continuous governance — reassess on model/prompt/tool/scope/user-group/regulatory change, serious incident or material performance drop.

4. NIST AI RMF

GOVERN

Policies, accountability, risk appetite, roles, culture, training, suppliers, documentation, escalation.

Risk appetite: limited residual risk from minor conversational errors in low-impact informational interactions; no appetite for unauthorised disclosure, fabricated transaction status, denial of human support or autonomous high-impact financial decisions.

Outputs: AI policy, operating model, inventory, risk appetite, RACI/RAPID, control library, training and supplier requirements, exception process.

MAP

Users (authenticated retail customers, agents, managers, knowledge managers); affected people including vulnerable and non-digital customers; benefits (speed, effort, consistency, availability); harms (wrong financial info, delayed fraud escalation, PII exposure, exclusion, overreliance, unfair quality, lost human access); dependencies (model/cloud providers, banking APIs, auth, knowledge, contact centre, human capacity).

Conclusion: sociotechnical system—not a model alone.

MEASURE

Validity/reliability, safety, security/resilience, accountability/transparency, explainability, privacy, fairness.

Example: human-request detection recall ≥99%; pilot 99.4% → acceptable for controlled deployment with monitoring.

MANAGE

Avoid, reduce, transfer, accept, monitor, restrict, stop.

Fraud misclassified as payment status: inherent High → controls (fraud classifier, keyword/semantic detection, low confidence, deterministic escalation, human review, monitoring) → residual Medium. Accept for pilot only if fraud escalation recall >99%, no autonomous fraud decision, immediate specialist access, weekly false-negative review.


5. NIST Generative AI Profile

RiskControl focus
ConfabulationLive status, grounding, no unsupported causal claims, validate vs tools
Information integrityAuthoritative sources, versioning, expiry, citations, owner approval
Data privacyMinimisation, ACL, masking, restricted logging, retention, contracts
Human–AI configurationSources, uncertainty, training, overrides, escalation, quality sampling
Information securityInput filtering, allowlists, tool isolation, detection, red team, least privilege
Value-chain / supplierDue diligence, change notice, fallbacks, contracts, exit, AI BOM

6. ISO/IEC 42001 AI Management System

Context (interested parties: customers, employees, regulators, board, shareholders, suppliers, consumer reps) → Leadership (policy: purpose, owner, risk class, evaluation, monitoring, retirement before production) → Planning (100% inventory; high-impact independent review; critical-incident timelines; versioned models/prompts; periodic reassessment) → Support (agent training on reliance, sources, escalation, vulnerability, reporting) → Operation (approval, data, design, eval, deploy, monitor, suppliers, IR) → Performance evaluation (audits, management review, KPI/KRI, control testing) → Improvement (contain → correct → root cause → update controls → retest → evidence → share learning).


7. ISO/IEC 23894 — risk register example

AI-R-017: incorrect pending-payment explanation. Causes: wrong tool result, outdated policy, retrieval failure, confabulation, mapping error, prompt change. Impacts: duplicate payment, financial loss, complaint, rework, reputation. Inherent High → controls (SoR API, status normalisation, approved knowledge, grounding, citations, monitoring, escalation) → residual Medium. Owner: Head of Payments Operations. Monitor incorrect-status rate, recontact, complaints, policy errors.


8. ISO/IEC 42005 impact assessment

System definition → positive/negative impacts → impact distribution (who benefits vs who bears risk; subgroup failure; non-digital disadvantage; vulnerable support) → mitigation (human access, accessible design, subgroup eval, alternatives, transparency, complaints) → residual impact and acceptance rationale.


9. EU AI Act risk classification

Verify operative legal timetable before deployment (implementation timing can change).

Prohibited: no emotion-based financial manipulation, social scoring, unnecessary sensitive-attribute inference, exploitation of vulnerability to sell products.

High-risk boundary: “Explain why a loan application is still pending” (service scope) ≠ “Determine whether the customer qualifies for credit” (separate credit system). Keep CS assistant off credit decisioning.

Transparency: clear disclosure that the user is speaking with a digital assistant, what it can do, and that a person is available anytime—never falsely present as a human.

GPAI: identify provider vs deployer obligations, documentation, limitations, updates, downstream responsibilities.

Record entity, geography, purpose, users, decisions, classification, rationale, reviewer, date, reassessment triggers.


10. AI System Inventory

Minimum fields: ID, name, description, purpose, prohibited uses, business/technical/data owners, risk and regulatory class, providers/models/data, users, geography, autonomy, oversight, production status, last/next review, retirement.

NS-AI-CS-001 Payment Status Assistant: explain selected statuses to authenticated retail customers; information/recommendation only; Medium risk; transparency obligations apply; not high-risk for approved informational scope; Owner: Director of Digital Customer Service; review quarterly and after material change.


11. Algorithmic Impact Assessment

Decision significance Moderate · data sensitivity High · autonomy Low–Moderate · scale High · vulnerability High · reversibility Moderate → overall Medium–High. Requires formal approval, DPIA, independent evaluation, human escalation, continuous monitoring, complaint route, quarterly reassessment.


12. Responsible AI Control Library

Accountability · transparency · fairness · privacy · safety · security · traceability controls mapped to lifecycle stages (named owners, disclosure, subgroup testing, minimisation/DPIA, prohibited-use/kill switch, threat modelling/injection tests, versioning and logging).


13–15. Model, Data and System Cards

Model Card (Model B): intended for explanations, summaries, draft responses; prohibited for credit/fraud liability/investment/account closure/autonomous complaints; correctness 95.1%, groundedness 96.4%, critical error 0.3%, human-request 99.4%; limitations if incomplete context / multi-intent / needs tools for live facts.

Data Card (Golden Dataset v2.3): 1,750 cases; English-only limitation; not for training without further review; Owner: AI Evaluation Lead.

System Card: overview through retirement criteria—architecture, data flows, models, retrieval, tools, oversight, eval, security/privacy, limitations, incidents, change history.


16. Human Oversight Framework

HITL · HOTL · human-in-command. Effectiveness requires information, authority, time, competence, usable UI, freedom from automation pressure, escalation route.

Weak oversight: large Send, hidden sources, no uncertainty, handling-time pressure. Redesign: prominent sources, uncertainty, high-risk acknowledgement, no penalty for appropriate overrides, measure overreliance, train challenge behaviour.


17. Three Lines Model

LineRoleMonGo teams
FirstOwn service, operate controls, day-to-day risk, incidents, knowledgeCS, AI Product, AI Engineering, Service Ops, Knowledge
SecondPolicy, frameworks, challenge, monitoring, regulatory interpretationEnterprise Risk, Compliance, Privacy, Cyber Governance, Model Risk
ThirdIndependent assuranceInternal Audit

Second/third lines do not own the system; first line cannot outsource accountability by seeking risk approval for every operational decision.


18. NIST CSF 2.0

Govern: AI cyber policy; all production model access via gateway.
Identify: models, prompts, vector stores, tools, APIs, suppliers, attack surfaces—e.g. Payment Knowledge Vector Index (poisoning, unauthorised access, deletion, embedding leakage, stale content).
Protect: identity, least privilege, encryption, tool restrictions, DLP; short-lived scoped tokens for selected-customer status only.
Detect: injection, abnormal tools, token spikes, cross-account queries, policy bypass, retrieval anomalies, output violations.
Respond: triage, containment, communications, investigation, remediation, regulatory escalation, evidence.
Recover: known-good versions, model switch, index rebuild, customer remediation, lessons, control updates.


19. STRIDE

ThreatBanking exampleControls
SpoofingStolen session for payment infoStrong auth, session binding, token expiry, reauth
TamperingModified prompts/knowledge/tools/logsSigned artefacts, ACL, integrity, versioning, immutable logs
RepudiationDispute card-freeze confirmationTimestamped confirm, correlation ID, protected audit
Information disclosureCross-customer data / prompt leakMinimisation, filtering, session isolation, redaction
Denial of serviceExpensive loopsRate/token/step limits, budgets, circuit breakers
Elevation of privilegeRead-only assistant invokes payment-changeSeparate tool identities, deny-by-default, scoped tokens

20. MITRE ATLAS

Attack chain: recon → malicious document → weak publication → RAG index → retrieval → model follows malicious text. Mitigate with allowlists, content scanning, instruction/data separation, publication approval, retrieval filtering, output validation, red teaming.


21. OWASP Top 10 for LLM Applications

IDBanking control focus
LLM01 Prompt InjectionSeparate data/instructions; restrict tools; validate output; auth outside model
LLM02 Sensitive DisclosureCustomer ACL; session isolation; redaction; synthetic data in dev
LLM03 Supply ChainDue diligence; signed artefacts; inventories; approved repos
LLM04 PoisoningSource approval; lineage; owners; anomaly detection; rebuild
LLM05 Improper OutputUntrusted output; schema; encoding; parameterised APIs
LLM06 Excessive AgencyMinimum tools; narrow permissions; confirmation; action limits
LLM07 Prompt LeakageNo secrets in prompts; enforce security outside model
LLM08 Vector/EmbeddingMetadata filters; access-aware retrieval; tenant isolation
LLM09 MisinformationGrounding; deterministic facts; citations; escalation
LLM10 Unbounded ConsumptionToken/tool/timeout/budget/rate limits; circuit breakers; cost alerts

22–24. Privacy by Design and DPIA

Minimisation: merchant display name, amount, date, normalised status—not full history, unrelated txns, risk scores or excess identity. Purpose limitation: support data not auto-reused for marketing, surveillance, training or credit. Storage limitation: distinct retention for conversations, security logs, traces, eval samples, audit. Defaults: no provider training; no cross-session memory; minimal context; restricted logging.

DPIA: describe processing → purpose → necessity → proportionality → risks → controls → residual-risk approval (DPO + risk authority). Design decision: model receives pseudonymous session ID, not full customer identity.


25. Fairness Assessment

GroupSuccessful safe resolution
Overall74%
Digitally confident81%
Low digital confidence62%
Screen-reader users66%
Informal English68%

Average hides disparity. Remediate UI, accessibility, clearer choices, evaluation data, visible human support, assisted testing—do not full-scale until the gap shrinks.


26. Transparency and explainability

Customers: AI use, capabilities, limits, supporting info, human route, complaint path. Agents: sources, dates, confidence, tool results, escalation reason, limitations. Governance: logic, controls, eval, releases, incidents, risk acceptance. Explain sources, process, limitations, actions and oversight—not every foundation-model parameter.


27–28. Third-party risk and AI BOM

Assess security, privacy, resilience, model governance, data use, subcontractors, geography, updates, auditability, incident notice, exit. Ask: training on prompts? retention? regions? change notice? investigation support? model withdrawal? subprocessors? export logs/config?

AI BOM: app, orchestration, prompt package, foundation model, classifier, embeddings, reranker, knowledge index, guardrails, payment API, monitoring, containers, OSS, cloud—so compromised or withdrawn components map to affected systems and customers.


29. Secure AI Development Lifecycle

Requirements (prohibitions, data class, appetite, auth, audit) → Design (threat model, Zero Trust, tools, fallback, privacy, suppliers) → Build (secure code, deps, prompt versioning, signed artefacts) → Test (pen test, injection, leakage, adversarial, permissions) → Deploy (approved envs, SoD, config verify, rollback, evidence) → Operate (threat monitor, vulns, IR, periodic red team, reassessment).


30–32. AI incident management

Categories: quality, safety, privacy, security, fairness, operational, supplier.

SeverityExamplesResponse
1 CriticalCross-customer disclosure; unauthorised financial action; widespread wrong status; tool compromiseImmediate containment; executive escalation; shutdown if needed; customer/regulatory assessment
2 MajorUnsafe-answer spike; escalation down; major supplier outage; serious a11y failureRestrict capability; major-incident process
3 ModerateLocal retrieval error; bad citation; latencyNormal IR; watch for spread

Worked incident: 1,200 customers told payment “cancelled” when pending. Contain mapping → deterministic/human route → preserve logs → identify customers → notify leaders → correct communications → assess harm. Root: bad backend status mapping + missing contract tests + incomplete eval + failed change notice. Fix: contract tests, schema-change approval, regression cases, unknown-status fallback, dependency monitoring.


33–34. Kill switch and decommissioning

ActionAuthority
Disable one prompt versionAI Operations Lead
Disable one toolSecurity or Service Owner
Restrict customer trafficProduct Owner
Stop customer assistantMajor Incident Commander
Permanently retireExecutive AI Governance Committee

Decommission: approve → notify → stop new use → preserve records → revoke credentials → remove tools/endpoints → archive/delete data → terminate suppliers → update inventory → verify no hidden dependency → communicate replacements → lessons learned.


35–36. Assurance Case and evidence pack

Top claim: payment-status assistant is sufficiently trustworthy for limited authenticated customer use.

Claims with evidence: approved scope · reliable information · protected data · effective human oversight · operable and controllable.

Evidence pack: charter, purpose, inventory, risk class, impact assessment, DPIA, System/Model/Data Cards, architecture and data flows, threat model, suppliers, evaluation, red team, oversight design, control testing, ORR, incident plan, benefits baseline, risk acceptance, release approval.


37. Responsible AI stage gates

GateRequired
ConceptPurpose, users, preliminary risk, prohibitions, sponsor
DesignImpact assessment, architecture, data, oversight, threat model, legal class
ExperimentApproved test data, privacy/security controls, eval plan, supplier assessment
PilotThresholds, DPIA, transparency, escalation, red team, operational ownership
ProductionAssurance Case, control testing, IR readiness, monitoring, risk acceptance, rollback, support
ScaleStable performance, no unresolved material harm, subgroup outcomes, benefits, risk appetite, updated legal assessment

38. Integrated control example

Cross-customer retrieval attack maps across AI RMF, ISO 42001, STRIDE (spoofing/disclosure/EoP), ATLAS, OWASP (injection/disclosure/agency), DPIA and technical controls: session-bound identity, ownership verification, scoped tokens, deny-by-default, redaction, anomaly detection, immutable logs.


39. Governance dashboard

Risk KRIs: critical policy error, unsafe answers, unauthorised tools, PII exposure, failed escalations, complaints, subgroup gap, unresolved incidents.
Control indicators: registered systems, reviews, versioned prompts, assessed suppliers, trained staff, remediated vulns, kill-switch tests.
Performance: successful safe resolution, groundedness, CSAT, adoption, cost per safe resolution, availability.


40. Final governance decisions

DecisionScope
Permitted for productionAgent knowledge assist; source-grounded drafts; summaries; authenticated payment-status explanations; deterministic card-status guidance; human escalation
Explicit confirmation onlyTemporary card freeze; callback scheduling; service-case creation
Human decision requiredComplaint outcome; fraud liability; financial-difficulty resolution; disputed payment; account closure; compensation
ProhibitedCreditworthiness; autonomous lending; investment advice; social scoring; emotion-based manipulation; cross-customer retrieval; unrestricted tools

41. Phase outputs

Governance operating model; AI RMF and GenAI profile; ISO 42001/23894/42005; EU AI Act classification; inventory; AIA; control library; Model/Data/System Cards; oversight framework; Three Lines; CSF; STRIDE; ATLAS; OWASP LLM; Privacy by Design; DPIA; fairness; supplier assessment; AI BOM; Secure AI SDLC; incident process; kill switch; decommissioning; Assurance Case; production evidence pack.


42. Core conclusion

Responsible AI is not a principles poster. It is the ability to prove—with current evidence—who owns the system, what it may do, who may be harmed, which controls work, how people get human support, how incidents are handled and when the system is restricted or stopped.

Not: “Do we have a responsible AI policy?”

Rather:

Can MonGo prove that this specific AI system remains within its approved purpose, risk appetite, legal obligations and customer-protection standards?

Next phase: Delivery, Change, Adoption and Operations—Agile, Dual-Track Agile, Stage-Gate, Product Operating Model, DevOps/DevSecOps, Continuous Delivery, ITIL 4, SRE, ADKAR, Kotter, McKinsey 7S, change impact, training needs, TAM, COM-B, champion networks, adoption measurement, benefits realisation and continuous improvement.

Discussion

Comments

Share feedback or questions about this page. No account required.

Loading comments…