Skip to main content

Consulting and Problem Solving

Executive view

Focus on **Why this matters**, the scenario narratives, and the **DeliverableBox** outputs. In reviews, ask for the one-sentence problem without AI, the top three hypotheses, and what evidence would change the recommendation.

Technical view

Build the issue tree, hypothesis register, and assumption log before architecture. Treat interviews and workshops as evidence collection for hypotheses—not as requirements gathering for a predetermined solution.

Why this matters

AI engagements fail early when the problem is never structured. Teams skip from executive enthusiasm to vendor selection, then discover—weeks into delivery—that stakeholders disagree on scope, success metrics contradict each other, and the real blockers are process and policy, not inference latency. Consulting problem-solving is how you prevent that waste: it converts ambiguity into an explicit decision frame that everyone can challenge constructively.

For an AI Solution Engineer, this capability is not "acting like a strategy consultant." It is the minimum discipline to know what you are building and why it should work. Partners and clients reward teams that can say: "The problem is X; we believe Y because evidence Z; if assumption A fails, we stop." That sentence is the output of issue trees, MECE decomposition, hypothesis registers, and synthesis—not of model benchmarks.

Problem structuring also protects technical staff from endless rework. When root cause is unclear, every architecture choice becomes reversible opinion. Clear problem trees anchor non-functional requirements: if the hypothesis is "agents lack authoritative policy access," retrieval and ACL design follow. If the hypothesis is "escalation paths are undefined," workflow and human-in-the-loop design precede any model tuning. Without the tree, teams optimise tokens while complaints persist.

Finally, recommendation quality is a trust asset. Executives encounter many AI proposals; few include falsifiable hypotheses and explicit trade-offs. When you present a MECE option set with risks and a phased path, you differentiate from vendor-led demos. Risk and compliance functions engage differently when you invite them to knock down assumptions early—not when they gatekeep at the end. This topic pairs directly with Opportunity Discovery for engagement delivery and with the 8D Framework for stage-gated evidence.

Learn

Problem structuring

Definition. Problem structuring is the disciplined act of defining what is being decided, for whom, by when, with what boundaries—and expressing it in outcome language independent of a chosen solution or technology.

Engagement use. Open every engagement with a draft problem statement and refine it after first interviews. A strong statement names: affected process, affected population, current pain (with baseline if possible), desired outcome, constraints, and explicit exclusions. Test it by asking: could someone propose a non-AI solution to the same statement? If not, the problem is still tool-shaped.

Pitfalls.

  • Accepting the client's first framing ("we need GenAI") as the problem.
  • Combining multiple decisions into one problem (cost reduction and revenue growth and NPS in one sentence).
  • Omitting constraints that will appear later (regulatory, union, timeline, data residency).
  • Using jargon the economic buyer does not use in board materials.

Worked example. Weak: "Implement an AI chatbot for customer service." Structured: "Reduce average handle time on Tier-1 policy lookup contacts in UK retail banking servicing by 20% within two servicing cycles, without increasing mis-selling or complaint rate, with human agents retaining authority on account changes and vulnerability cases." This statement supports multiple solution paths—including better search, workflow fix, or staffing—while keeping AI in contention if evidence supports it.

Issue trees and MECE decomposition

Definition. An issue tree breaks a problem into sub-questions that are Mutually Exclusive, Collectively Exhaustive (MECE)—no overlaps, no major gaps—so analysis can be assigned, prioritised, and synthesised. Nodes are questions or drivers, not activities or org chart boxes.

Engagement use. Build a tree top-down after the problem statement, typically two to four levels deep for discovery. Assign hypotheses and evidence sources to leaf nodes. Use the tree to run workshops: "Today we are testing this branch only." MECE prevents double-counting benefits and stops pet solutions from hiding unrelated workstreams.

Pitfalls.

  • MECE violations: overlapping nodes ("data quality" and "missing data" without clear boundary).
  • Activity trees disguised as issue trees ("build RAG," "train model"—these are solutions, not questions).
  • Trees too shallow (everything under "technology") or too deep before any evidence (analysis paralysis).
  • Ignoring external branches (vendor market, regulatory change) when they dominate feasibility.

Worked example. Problem: "Complaints rising in insurance FNOL." Level-1 MECE tree: (1) Are customers complaining about the same failure modes? (2) Is process latency or error driving dissatisfaction? (3) Are communication expectations mis-set? (4) Is there fraud or abuse inflating cases? Under (2): handoffs, policy ambiguity, adjuster capacity, digital channel gaps—each a testable branch with different interventions. "Add chatbot" appears only if a branch supports it.

Hypothesis-driven consulting

Definition. Hypothesis-driven consulting states upfront what you believe is true, what would falsify it, and what evidence you will collect—then updates beliefs as data arrives. It inverts "gather all requirements then decide" into "decide what to test, then gather efficiently."

Engagement use. Maintain a hypothesis register: ID, hypothesis statement, confidence, evidence needed, owner, status (open/ supported/ rejected). Link each hypothesis to issue tree nodes. In AI discovery, hypotheses often concern root cause ("40% of contacts are policy lookup"), feasibility ("authoritative corpus exists"), and value ("3 minutes savable per contact"). Kill weak branches early when evidence falsifies them.

Pitfalls.

  • Hypotheses that are vague ("stakeholders want AI") or unfalsifiable ("leadership supports innovation").
  • Confirmation bias—only interviewing sponsors who agree.
  • Keeping rejected hypotheses undocumented—teams repeat dead paths.
  • Treating pilot success as hypothesis proof without counterfactual or control.

Worked example. H1: "≥35% of Tier-1 contacts are retrievable policy lookups." Evidence: stratified sample of 200 contact logs coded by intent. H2: "Approved policy corpus is complete for top 20 enquiry types." Evidence: SME sign-off against enquiry taxonomy. H3: "Lookup time ≥4 minutes due to search friction, not training gap." Evidence: shadowing + system timestamps. If H1 fails, the business case for a lookup assistant collapses—reframe before build.

Five Whys and root-cause analysis

Definition. Five Whys iteratively asks why a symptom occurs to surface contributing causes—recognising that complex systems rarely have a single root cause. It complements issue trees for operational and incident-style problems.

Engagement use. Use when stakeholders anchor on a surface symptom ("AI will fix complaints"). Stop when you reach causes you can act on (policy ambiguity, missing RACI, tool fragmentation)—not abstract culture statements unless backed by observable behaviours. Document the chain; multiple branches are normal.

Pitfalls.

  • Stopping at blame ("human error") without systemic cause.
  • Single-threaded Whys when multiple parallel causes exist.
  • Using Whys instead of data when logs and samples are available.
  • Assuming root cause is always technical—organisational causes dominate many service failures.

Worked example. Symptom: complaints about "wrong information" on calls. Why? Agents give inconsistent answers. Why? Multiple policy versions in circulation. Why? No single published source after product change. Why? Product launch process lacks compliance publication gate. Why? RACI between Product and Compliance undefined for servicing docs. Intervention may be process and knowledge management—not GenAI—unless corpus governance is fixed regardless.

Interviews, workshops, and evidence collection

Definition. Structured interviews and workshops collect primary evidence for hypotheses: how work happens, where exceptions cluster, what "good" looks like, and what failed before. Good consulting treats interviews as evidence sessions, not sales conversations or unstructured venting.

Engagement use. Prepare interview guides mapped to issue tree leaves. Use consistent prompts across roles (ops, compliance, IT, frontline). Record: quote, observation, metric, system artefact. Triangulate—if three roles describe the same bottleneck differently, that disagreement is data. Workshops synthesise; they should not be the first time hypotheses are written down.

Pitfalls.

  • Only interviewing managers, not practitioners.
  • Leading questions ("Wouldn't AI help with…?").
  • No note template—insights stay in consultants' heads.
  • Workshop theatre: sticky notes without decision criteria or owners.

Worked example. For claims triage, interview adjuster (daily workflow), team lead (QA patterns), SIU (fraud referral triggers), compliance (fair treatment), IT (claims platform constraints). Guide sections align to tree: volume by claim type, time on document review, authority limits, known system workarounds. Synthesis day produces updated hypothesis register with confidence levels—not a slide of "themes."

Interview guide template (adapt per leaf node).

SectionPromptsCapture
ContextRole, tenure, volume handledBaseline credibility
Happy pathWalk through last typical caseSteps, systems, minutes
ExceptionsLast difficult caseHandoffs, approvals, rework
MetricsHow you are measuredKPIs, QA, incentives
WorkaroundsUnofficial tools/processesShadow IT, spreadsheets
Prior changeLast initiative that failedRoot cause of failure
AI probe (last)If answers were instant and cited?Fears, must-not-happen

Keep AI questions at the end so you do not anchor responses. Record verbatims where policy language matters.

Workshop facilitation patterns.

  • Branch-focused workshop: One issue tree branch only; pre-read includes hypothesis status.
  • Red team review: Risk/compliance challenges assumptions—not architecture—before steering.
  • Synthesis readout: 20-minute pyramid; appendix for evidence; Q&A captured in decision log.
  • Kill party: Explicit session to close falsified hypotheses and communicate scope reduction—builds trust.

Avoid "innovation theatre" workshops that produce idea walls without owners, metrics, or links to tree nodes.

Synthesis and storyline

Definition. Synthesis converts fragmented evidence into a coherent answer to the problem statement—typically pyramid structure: recommendation first, then supporting arguments grouped MECE, then evidence appendices. Storyline is the narrative executives can retell without you in the room.

Engagement use. After evidence pass, force a one-page synthesis: recommendation, three reasons, risks, next step. Check mutual exclusivity of reasons (not three versions of "AI is trendy"). Align with domain language from topic 03. Prepare explicit "what we are not recommending yet" to manage scope.

Pitfalls.

  • Listing findings without a decision.
  • Buried lead—ten slides before the answer.
  • Recommendations that are a bundle of every interview idea (non-MECE options).
  • Missing disconfirming evidence ("we didn't test X").

Worked example. Synthesis headline: "Fix authoritative policy publication and agent retrieval before any generative assistant pilot." Reason 1: 38% contacts are lookup—volume justifies investment. Reason 2: corpus incomplete for 6 of top 10 intents—GenAI would amplify errors. Reason 3: compliance rejects unsourced answers—citation architecture required regardless. Next: 6-week corpus programme + retrieval benchmark; revisit GenAI phrasing after SME sign-off.

Recommendation quality and decision facilitation

Definition. A quality recommendation states the decision, options considered, preferred path, trade-offs, dependencies, risks, success metrics, and what happens if assumptions fail—at a level appropriate for the decision body (steering committee vs technical design authority).

Engagement use. Use option sets where real alternatives exist: fix process, buy search, build RAG, change staffing. Score options against shared criteria (value, feasibility, risk, time to value)—not hidden weights favouring a preset vendor. Facilitate decisions by pre-socialising risk and compliance on assumptions, not only on final slides. Log decisions in a decision log with date, decider, rationale, revisit trigger.

Pitfalls.

  • Fake options (straw man "do nothing" when status quo is unacceptable).
  • Single-option "recommendation" that is a project plan in disguise.
  • No revisit triggers—teams persist when evidence has flipped.
  • Decision made in room without owner for follow-through actions.

Worked example. Steering committee choice: Option A corpus + guided search (lower risk, 12 weeks); Option B RAG assistant with citations (medium risk, 16 weeks); Option C autonomous agent (high risk, deferred). Recommend A then B gated on corpus completeness metric ≥95% on top intents. Decision log entry: "Proceed A; B conditional on KPI-CORPUS-01; C out of scope until complaint rate stable 2 quarters."

Frameworks and methods

MECE (Mutually Exclusive, Collectively Exhaustive)

Apply when: Decomposing problems, structuring slide decks, grouping findings, defining epics/user story maps.

Do not apply when: The problem space is inherently overlapping (some regulatory topics)—then use "approximately MECE" and document overlaps explicitly.

Test: Can two sibling nodes be true for the same case without contradiction? Is there an obvious missing sibling executives would ask about?

Issue trees vs hypothesis trees

ToolNode typeBest for
Issue treeQuestions/driversScoping discovery, workshop agendas
Hypothesis treeTestable beliefsEfficient evidence collection, killing branches

Use together: issue tree for structure, hypotheses on leaves for fieldwork.

Five Whys

Apply when: Operational symptoms, incident reviews, service failures, handoff errors.

Pair with: Process maps and sample data—not Whys alone.

SIPOC and process mapping

Apply when: Boundaries unclear between teams/systems; need to anchor metrics to process steps.

Output: Inputs/outputs per step feed hypothesis design ("which step owns the delay?").

MoSCoW and RICE (prioritisation)

MoSCoW: Must/Should/Could/Won't for scope negotiation when problem is structured but backlog is wide.

RICE: Reach, Impact, Confidence, Effort—for comparing solution options or use cases once problem is validated.

Caution: Do not RICE prioritise before problem validation—you will optimise the wrong backlog.

Value–risk–feasibility scoring

Three-axis scorecard for options after synthesis. Value (economic lever), Feasibility (data, integration, skills), Risk (regulatory, reputational, operational). Weight risk heavily in regulated industries.

Pyramid Principle (Barbara Minto)

Answer → supporting arguments → evidence. Mandatory for executive readouts and steering packs.

8D stage alignment

Map problem-solving artefacts to gates: problem statement and tree in Define/Discover; hypothesis register through Design; decision log through Deliver. See 8D Framework.

Consulting artefact map (what to produce when)

PhasePrimary artefactsQuality bar
Week 0–1Problem statement, draft issue tree, interview planSponsor agrees problem is worth solving
Week 2–3Hypothesis register, interview notes, process sketchAt least one hypothesis rejected or weakened
Week 3–4Synthesis one-pager, option set, decision log draftCompliance/risk pre-read complete
GateRAID, measurement plan, phased roadmapVALUE or 8D gate criteria met

Avoid producing architecture diagrams before the synthesis one-pager exists—teams confuse activity with progress.

First principles vs analogical reasoning

First principles ask what must be true for the outcome to improve (e.g., "agents must see the same policy version customers are sold"). Analogical reasoning asks what worked elsewhere ("Bank X deployed a copilot"). Use analogy to generate hypotheses, not to skip validation—different corpus governance or regulatory regime breaks the analogy. Document borrowed patterns with explicit transfer conditions.

Operating model for consulting on AI engagements

Problem-solving is not only analysis—it is how the engagement itself runs.

Single problem owner. Name one client-side owner for the problem statement—not only the project manager. Without this, synthesis debates restart every steering meeting.

Evidence cadence. Weekly hypothesis update: supported / rejected / open. Visible to client PMO. Prevents surprise pivots at month six.

Challenge roles. Assign a internal "red team" reviewer to knock down MECE and hypotheses before executive readout—partner, risk SME, or peer from another account.

Decision hygiene. No "verbal approval" on scope shifts; decision log entry or formal deferral. AI engagements churn scope when new model capabilities appear mid-stream—log what changed and why.

Workshop design. Opening: restate problem and tree branch focus. Middle: evidence review, not brainstorming. Close: decisions, owners, dates—not "more discovery needed" without bounded next tests.

Discovery week rhythm (example)
────────────────────────────────
Mon Update hypothesis register from prior week
Tue Frontline / SME interviews (leaf nodes)
Wed Data sample or shadowing
Thu Internal synthesis + red team challenge
Fri Client working session → decision log updates

Adapt cadence to client procurement rules—public sector may require longer notice for workshops.

Real-world scenarios

Scenario A: "AI will fix complaints" (retail banking servicing)

A UK retail bank steering group requests a GenAI customer assistant after complaint volume rises 18% year-on-year. Marketing frames the problem as "customers want digital answers faster."

You facilitate restructuring. Five Whys on complaint codes shows top drivers: conflicting policy guidance after a product change, not chat UX. Issue tree: (1) product/policy clarity, (2) agent knowledge access, (3) handoffs between teams, (4) vulnerable customer handling gaps. Interviews with team leads, compliance, and 12 agent shadow sessions: 41% of sampled complaints reference "was told different things"; corpus audit shows 23% of top-intent answers rely on deprecated PDFs.

Hypotheses: H1 policy/publication failure (supported); H2 agent search friction (supported); H3 understaffing (weak—AHT flat). Synthesis: Recommend corpus remediation + guided retrieval pilot for Tier-1 lookup intents; defer generative phrasing until compliance signs corpus completeness KPI. Measurable outcome: complaint rate on policy-conflict category down 15% in pilot branch; no increase in mis-selling flags.

Executive reaction: "You told us what we were afraid to admit—we were going to automate confusion." Credibility earned; scope reduced to controllable phase one.

Scenario B: "We need GenAI for underwriting" (commercial insurance)

A commercial insurer executive wants "GPT for underwriters" after a competitor press release. Underwriting cycle time up 22%; loss ratio pressure high.

Problem statement draft: "Reduce technical underwriting cycle time on mid-market property submissions by 15% without increasing referral errors or ESG policy breaches." Issue tree branches: submission quality, external data latency, rules engine maintenance, underwriter tooling, referral governance. Evidence: sample 150 submissions—40% time in re-keying and chasing missing surveys; only 8% in wording narrative. Hypothesis H-submission-data dominates.

Recommendation: Structured ingestion + validation for survey and loss data; pilot summarisation of loss runs with citations only after data quality threshold met—not open-ended GenAI on free-text alone. Option set: A) data fix only; B) data + summarisation; C) autonomous quote. Recommend B with gate on data completeness metric.

Measurable: cycle time −12% on pilot segment; referral rate unchanged; ESG exclusion rules remain rules-engine enforced, not model-interpreted. Partner uses this to align with Opportunity Discovery workshops rather than jumping to model selection.

Scenario C: Public sector service backlog (stretch)

A government agency faces backlog in permit applications; minister demands "AI clearance." Union sceptical; algorithmic transparency policy applies.

Problem structuring separates throughput from decision quality and appeal rights. Issue tree includes legal mandate constraints—not optional. Hypotheses test whether backlog is completeness checks vs assessment vs post-decision correspondence. Interviews show 55% resubmissions due to incomplete forms, not assessor capacity.

Recommendation: assisted completeness checking and citizen-facing guidance (rules + retrieval over public guidance)—not automated permit decision. MECE options presented with transparency artefacts required for each. Decision log triggers revisit if resubmission rate does not fall within two reporting periods.

Shows consulting skill in non-commercial incentives: political urgency vs lawful process design.

Scenario D: SaaS customer support "deflect with AI" (technology sector)

A B2B SaaS vendor ($120M ARR) wants to "deflect 30% of tickets with AI" after support headcount grows faster than revenue. NRR already soft; CS leadership fears slow responses hurt expansion.

Problem statement: "Reduce median first-response time on Tier-1 how-to tickets by 25% without lowering CSAT or increasing escalation rate to engineering, for paid accounts in NA and EU." Issue tree: (1) ticket taxonomy accuracy, (2) documentation quality and findability, (3) agent tooling, (4) product UX confusion driving tickets, (5) staffing/skill mix. Evidence: ticket sample shows 52% are "how-to" but only 31% match published KB articles; product release notes lag docs by average 11 days. Hypotheses: H-docs (supported), H-product-doc drift (supported), H-headcount (weak—utilisation 78%).

Options: A) doc refresh + search; B) RAG in agent console with citations; C) customer-facing bot. Recommend B then C: agents validate retrieval quality before customer exposure; CSAT guardrail mandatory. Not recommended: customer bot first—would automate wrong answers from stale docs.

Measurable: first-response time −28% on pilot queue; CSAT ±0.2 points; escalation rate flat. Demonstrates problem-solving in a sector with faster procurement but still doc-governance root cause—parallel to regulated clients.

Worked example: Full mini issue tree (insurance FNOL complaints)

Each leaf gets one hypothesis, one evidence method, one owner. Synthesis might prioritise status notification gaps over GenAI—fix push/SMS templates and portal status before chatbot.

Practice exercises

Primary exercise: MECE issue tree and hypothesis register (2–3 hours)

Take the prompt: "We need GenAI for our contact centre." (Use retail banking, insurance, or telco—you choose.)

Deliver:

  1. Problem statement (one paragraph) — no AI in the first sentence; include metric target, population, constraints.

  2. Two-level MECE issue tree — minimum eight leaf nodes; label any MECE compromises.

  3. Hypothesis register — minimum six falsifiable hypotheses mapped to leaves; specify evidence method (sample size, interview role, system report).

  4. Draft recommendation (half page) — option set of three paths; preferred path with revisit trigger.

Acceptance criteria: A reviewer can identify which hypothesis, if false, would kill your preferred path; tree contains no "build chatbot" nodes at level 1 or 2.

Stretch exercise: Five Whys to synthesis memo (2 hours)

Pick a real or fictional operational incident (wrong bill, claim delay, clinical prior auth denial, network outage ticket flood). Run a Five Whys chain (allow branching). Produce a one-page synthesis with recommendation that is not primarily AI. Add a short paragraph: "Where AI could help later—and only after which non-AI fixes."

Acceptance criteria: Root causes include at least one process/governance factor; AI is sequenced, not default.

Reflection exercise: Challenge your own recommendation (30 minutes)

Take any AI proposal you are currently shaping (client or internal). Write:

  1. The problem statement—then strike through every technology word and rewrite.
  2. Two hypotheses that would disprove your preferred approach if true.
  3. One interview you have not done that could surface disconfirming evidence.
  4. The honest "do nothing / fix process only" option—and one metric it might improve.

File this with your pattern library entry. Managers use this in 1:1s to test consulting rigour before staffing senior client meetings. Repeat when scope or sponsor narrative shifts materially mid-engagement—before the next gate review.

Questions you should be able to answer

  1. What is the problem in one sentence without naming AI, vendors, or models?
  2. What decision is this engagement trying to produce, and who is the decision body?
  3. What are the level-1 MECE drivers of the problem—and where might they overlap?
  4. Which three hypotheses, if false, would kill the current proposal?
  5. What evidence have you collected vs what remains assumption?
  6. What did frontline staff say that contradicted the sponsor narrative?
  7. What options did you consider besides the recommended path—and why were they rejected?
  8. What is explicitly out of scope for this phase, and how was that negotiated?
  9. What baseline metrics anchor the value case—and who owns the measurement?
  10. What would change your recommendation in the next two weeks of discovery?
  11. How does the recommendation map to domain constraints from industry knowledge work?
  12. What risks are you asking the client to accept—and who signed off?
  13. What failed initiatives touched this problem before, and why did they fail?
  14. Can you retell the synthesis pyramid in 60 seconds: answer, three reasons, next step?
  15. What is logged in the decision log, and when should the decision be revisited?

Negative cases

Weak problem-solving produces expensive clarity failures. Recognise these patterns early.

Solution-first discovery. Workshops enumerate AI features; no shared problem statement. Fix: re-run Define with executive sponsor; refuse architecture until statement signed.

Issue tree theatre. Beautiful tree never updated after contradictory evidence. Fix: hypothesis register with kill criteria; weekly synthesis updates.

MECE violations double-count value. Benefits summed from overlapping branches ("efficiency" + "automation" counting same hours). Fix: assign metrics to single owner nodes.

Interview echo chamber. Only digital transformation allies interviewed; frontline and compliance absent. Fix: role-balanced guide; explicit disconfirming interviews.

Analysis paralysis. Fourth workshop without recommendation or experiment. Fix: time-box discovery; define minimum evidence for phase gate per 8D.

Recommendation without owner. Steering committee agrees; no one owns corpus remediation dependency. Fix: RAID log with named owners and dates at decision time.

GenAI as substitute for governance. Model proposed where policy ambiguity is root cause—automates inconsistency at scale. Fix: Five Whys + compliance on corpus authority.

Unfalsifiable success criteria. "Improve employee experience" without metric. Fix: tie to operational KPIs already reported to leadership.

Hidden vendor bias. Options weighted toward pre-selected platform. Fix: score options before vendor shortlist; document weights.

Decision log absent. Teams revisit the same debate monthly. Fix: decision log with revisit triggers shared with client PMO.

Premature convergence. Sponsor pressure closes discovery before compliance interview. Fix: minimum evidence checklist per 8D gate; named dissent in decision log.

Confusing correlation with causation. Contact volume dropped same quarter as pilot—claimed as AI win without control branch. Fix: matched cohort or stepped-wedge design; pre-register metrics.

Integration with Discovery and 8D

Problem-solving artefacts are the inputs to Opportunity Discovery delivery and the evidence for 8D Framework stage gates—not parallel paperwork.

In Discovery, the problem statement and issue tree become workshop anchors: stakeholder maps attach to tree nodes; workflow maps validate handoffs; constraint lists attach to regulatory branches. When Discovery produces a "use case shortlist," each item must trace to a supported hypothesis—not a brainstorm sticky note.

In 8D, typical alignment is:

  • Define (D1–D2): Signed problem statement; initial issue tree; sponsor and problem owner named.
  • Discover (D3): Hypothesis register populated; interview evidence; updated tree.
  • Design (D4–D5): Option set with value–risk–feasibility scores; preferred path with dependencies.
  • Develop/Deliver (D6–D8): Decision log maintained; revisit triggers acted on when assumptions fail.

VALUE reviews should challenge weak problem structure before challenging weak architecture—many "AI quality" failures are mis-scoped problems wearing a model metric.

Practice checklist

  • I can state the problem in one sentence without naming AI
  • I built a MECE issue tree and flagged any deliberate overlaps
  • I maintain a hypothesis register with at least one rejected hypothesis documented
  • I collected evidence from frontline roles—not only sponsors
  • I produced a synthesis with explicit "not recommending yet" scope
  • I presented at least three real options with trade-offs
  • I logged a decision or explicit deferral with revisit trigger
  • I completed the primary exercise and filed artefacts in my pattern library

Supplemental scenario — logistics carrier dispatch delays

Context: European parcel carrier; sponsor claims "AI routing" will fix +14% late deliveries. Discovery week 0: issue tree before any model discussion.

Problem statement (outcome language): "Reduce late-first-attempt deliveries from 11.2% to ≤9.0% within 12 months without increasing driver overtime above +3%."

Issue tree (MECE excerpt):

BranchDriver hypothesisEvidence collectedStatus
Demand forecast errorWrong parcel volume by depot6-month MAPE 18% at 3 depotsSupported
Route static rulesLegacy cut-off timesProcess map shows 2 manual overrides/daySupported
Driver skill gapTraining insufficientSurvey n=120; not primaryRejected
GenAI copilotFaster exception handlingShadowing 8 dispatchersPartial

Synthesis (pyramid top): Recommend rules + forecast fix first (£420k); defer GenAI copilot until baseline late rate stable 2 quarters. Not recommending end-to-end "AI routing platform"—hypothesis register shows 62% of delay minutes from forecast + cut-off rules, not search time.

Decision log: Sponsor Agrees to Phase 0 rules; Defers copilot with revisit trigger "late rate ≤9.5% for 8 weeks."

Lesson: Fast kill on AI branch saved £1.1m mis-scoped build (estimated vendor quote)—problem structure before technology.

Discussion

Comments

Share feedback or questions about this page. No account required.

Loading comments…