Product Management
Executive view
Technical view
Why this matters
Most enterprise AI failures are product failures wearing model metrics. A RAG assistant can score 87% on a golden set while 12% of target users open it monthly—because the job-to-be-done was "draft regulator-ready response with audit trail," not "answer questions faster," and nobody wrote that in a PRD. Product management is how AI Solution Engineers prevent building the right model for the wrong product.
Discovery without product translation wastes Stage 1–2 work. Issue trees and opportunity scores sit in decks while engineering improvises scope from demo feedback. A PRD linked to discovery hypotheses forces traceability: every epic should map to a supported problem branch or explicit new learning goal. When scope debates arise, you return to the PRD and MVP boundary—not re-run politics.
MVP vs platform decisions determine whether you earn expansion revenue or drown in integration. Platforms promise reuse; MVPs promise fast validated learning on one job, one persona, one metric. AI amplifies the trade-off: platform teams build ingestion pipelines nobody consumes; MVP teams ship a narrow assistant that proves chargeback and override patterns reusable later. Product management names what is in, out, and later—with acceptance criteria stakeholders sign.
Adoption metrics separate launch from value. Executives conflate deployment with success. Product discipline defines leading indicators (WAU, tasks completed, time-on-task delta, escalation rate, override rate, thumbs feedback) and lagging indicators (handle time, error rate, NPS segment, cashable hours). These feed Commercial and Financial Modelling benefits realisation and Change Management and Adoption interventions.
AI-specific product requirements—confidence UX, citation, human override, hallucination handling, escalation, continuous evaluation—are not UX niceties. They are what regulators and unions ask for first. Product managers who omit them force expensive rework when risk reviews the near-finished build. Pair this topic with Adoption guide, AI Product Management roadmap, and topic 21 (evaluation).
Learn
From discovery outcomes to product backlog
Definition. A product backlog is an ordered list of items (epics, stories, spikes) that translate validated discovery into buildable work—each item linked to a user outcome, acceptance criteria, and evidence it addresses a discovery hypothesis or explicit learning goal.
Engagement use. After discovery, run a translation workshop: import opportunity shortlist rows; map each to epic candidates; kill or defer items with no hypothesis link. Split epics into user stories in INVEST shape (Independent, Negotiable, Valuable, Estimable, Small, Testable). Tag stories: must-mvp, should-v1, platform, spike. Maintain discovery trace IDs on backlog items so steering can audit scope.
Pitfalls.
- Backlog as vendor feature checklist ("enable agents," "add GPT-4").
- No linkage to discovery—scope re-litigated every sprint.
- Everything priority-one; no ordering model.
- Missing spikes for unknowns (corpus completeness, tool reliability).
- Backlog hidden in Jira without executive-readable roadmap slice.
Worked example. Discovery supported H2: "≥38% Tier-1 contacts are policy lookup with authoritative corpus gaps on 6 intents." Backlog epics: (E1) corpus publication workflow integration; (E2) retrieval + citation assistant for lookup intents only; (E3) agent console UX; (E4) eval dashboard. MVP includes E1 partial (top 10 intents) + E2; E3/E4 deferred. Story example: "As an agent, I see cited policy excerpts with version date so I can read answer aloud compliantly."
PRDs for AI products
Definition. A Product Requirements Document (PRD) describes problem, users, scope, functional/non-functional requirements, success metrics, dependencies, risks, and release phases—at sufficient detail for engineering, design, risk, and ops to estimate and build without guessing intent.
Engagement use. PRD sections for AI: job statement, personas, in/out scope, user journeys, functional reqs (retrieve, cite, override, escalate, log), non-functional (latency p95, languages, accessibility), data/corpus reqs, eval acceptance, human-in-the-loop rules, telemetry, rollout/phasing, open questions. Version PRD; link to business case benefits line items. Socialise with compliance before sprint commitments.
Pitfalls.
- PRD describes model behavior vaguely ("accurate," "helpful").
- No negative requirements ("must not execute transactions").
- Missing eval acceptance thresholds tied to release gates.
- Single persona when multiple roles have conflicting jobs.
- PRD written after build started—change control chaos.
Worked example. Claims summarisation PRD: Primary user = property adjuster; job = "produce first-pass summary for complex water damage claims in ≤5 minutes with citations to uploaded photos and policy clauses." Out of scope: coverage determination, payment authorization. NFR: p95 3.5s interactive; EN+FR; WCAG 2.1 AA. Eval gate: ≥88% citation presence on 200-claim golden set; ≤2% critical hallucination class on red-team set.
Personas, jobs-to-be-done, and journeys
Definition. Personas are research-based archetypes (goals, constraints, systems, fears—not demographics alone). Jobs-to-be-done (JTBD) describe progress users seek ("when I receive a 40-page submission, I need a defensible summary fast"). Journey maps sequence steps, pain points, and moments where AI helps or harms trust.
Engagement use. Build personas from discovery interviews—not imagination. One primary persona for MVP; secondary personas explicit with phased support. Map happy path and exception path (low confidence, missing corpus, angry customer, vulnerable user). Mark human authority moments on journey (sign-off, escalation).
Pitfalls.
- Generic "business user" persona.
- Journey ends at AI answer—omits verification and downstream systems.
- Ignoring manager persona who owns QA metrics and incentives.
- JTBD written as solution ("use chatbot") not outcome.
Worked example. Underwriter persona "Sam": 8–15 submissions/day; pain = re-keying loss runs; fear = ESG breach on excluded sectors. JTBD: "Quickly see prior losses and survey gaps without missing referral triggers." Journey: email arrival → portal upload → AI summary draft → Sam edits → rules engine referral → sign-off. AI touchpoints only at summary draft with citations; referral stays rules engine.
Prioritisation: value, risk, learning, and capacity
Definition. Prioritisation orders backlog items using explicit criteria—typically value to user and business, risk reduction, learning uncertainty, dependencies, and team capacity—so trade-offs are visible to sponsors.
Engagement use. Use RICE (Reach, Impact, Confidence, Effort) or WSJF on epics after MVP scope frozen for must-haves. For AI, add risk reduction weight (compliance, safety). Publish prioritisation rationale in steering packs—one paragraph per deferred epic. Re-prioritise on eval or adoption signals, not demo applause.
Pitfalls.
- HiPPO (highest paid person's opinion) overrides scoring.
- Effort underestimated for integration and eval harnesses.
- No confidence score—pet projects treated as certain value.
- Prioritising visible UI over corpus/eval foundations.
Worked example. Epic scores: E2 retrieval assistant RICE 420; E5 "agent auto-closes tickets" RICE 95 (low confidence, high risk)—deferred. Steering agrees deferral documented in decision log with revisit trigger: "Revisit E5 if override rate <8% and eval green for 2 quarters."
MVP vs platform: thin vertical vs horizontal reuse
Definition. MVP (Minimum Viable Product) is the smallest release that validates the core job and metric for one segment. Platform investments (shared ingestion, model routing, eval framework, identity) enable multiple products—but delay first value if over-built upfront.
Engagement use. Decide MVP boundary in writing: one BU, one locale, one intent cluster, one integration path. Platform elements allowed only if blocking MVP (e.g., auth, logging) or contractually required. Use platform roadmap lane for reuse after MVP adoption ≥ target. Document technical debt accepted for speed vs platform tax avoided.
Pitfalls.
- "MVP" includes 12 integrations and 4 languages.
- Building data lake before one working assistant.
- No MVP success criteria—platform justified by architecture elegance.
- Duplicate MVPs across BUs without shared learnings.
Worked example. Bank chooses thin vertical MVP: UK retail mortgage servicing—policy lookup only, English, Salesforce embed, 800 agents. Platform team builds only shared auth + telemetry. Phase 2 platform: shared retrieval service if MVP hits ≥50% WAU and −18% AHT on lookup cohort. Wealth BU waits—explicit Won't in MVP PRD.
User stories and acceptance criteria for AI
Definition. User stories express user value in one sentence; acceptance criteria are testable conditions (Given/When/Then or checklist) including AI-specific behaviors (cite, refuse, escalate, log override).
Engagement use. Every story includes: happy path, low-confidence path, missing-data path, telemetry events, eval case IDs where applicable. Separate model stories from workflow stories. Definition of Done includes eval regression pass on linked goldens.
Pitfalls.
- Acceptance criteria only cover sunny-day phrasing.
- No criteria for citation format or source version display.
- Stories too large ("implement RAG")—not estimable.
- Missing audit log requirements for overrides.
Worked example. Story: "As a nurse, I receive cited clinical guideline excerpts for drug interaction checks so I can verify before administering." AC: (1) Given question on known interaction, When query submitted, Then response includes ≥1 citation from approved corpus with date; (2) Given confidence below threshold, Then system shows "verify manually" without dosing advice; (3) When nurse overrides, Then reason code captured and sent to review queue.
AI-specific product: trust, override, and evaluation loops
Definition. AI products require uncertainty communication (confidence, ranges, "I don't know"), human override paths with feedback capture, hallucination handling (refusal, escalation), and continuous evaluation feeding releases—not one-time UAT.
Engagement use. PRD must specify: when to show citations; when to block auto-send; override UI; feedback taxonomy; eval cadence (weekly smoke, release regression); who reviews failures (ops, compliance). Link eval metrics to release gates in topic 21.
Pitfalls.
- Binary answers on ambiguous policy questions.
- Override data not routed to improvement backlog.
- Eval only on launch; corpus drift ignored.
- Trust UI cosmetic—no effect on workflow.
Worked example. Servicing assistant: auto-send disabled for payment changes; amber band 0.65–0.82 confidence shows draft + sources; red <0.65 forces search fallback. Override reasons: wrong doc, outdated, misinterpreted—fed to weekly corpus + prompt review. Target: override rate declining from 14% → 8% over two quarters as corpus improves.
Roadmaps, releases, and stakeholder communication
Definition. A roadmap communicates sequence of outcomes across time horizons (now/next/later)—not a commitment to fixed dates for uncertain AI capabilities. Release plans tie versions to eval gates and change readiness.
Engagement use. Roadmap rows: outcome, persona, metric target, dependency, risk. Separate committed vs aspirational lanes. Align with Delivery and Programme Management milestones. Update when eval or adoption falsifies assumptions.
Pitfalls.
- Date-driven roadmap for research-heavy features.
- Roadmap as Gantt of technical tasks only.
- No communication to affected ops teams until launch week.
- Promising model upgrades vendor has not shipped.
Worked example. Now: lookup MVP + adoption programme; Next: summarisation for complex claims with human QA queue; Later: multilingual if EN adoption ≥55% WAU. Each column lists metric gate before progression.
Adoption metrics and product analytics
Definition. Adoption metrics measure whether target users use the product as intended and whether usage correlates with outcome KPIs. Product analytics instrument events (session start, task complete, override, escalation, time saved proxy).
Engagement use. Define North Star (e.g., successful assisted tasks/week) and guardrails (error rate, escalation, complaints). Segment by team, tenure, channel. Baseline before launch; cohort compare pilot vs control where possible. Feed monthly benefits realisation reviews (topic 07).
Pitfalls.
- Tracking logins only—not completed jobs.
- No baseline pre-AI metric.
- Vanity metrics (messages sent) vs outcomes (handle time).
- Analytics without privacy review on prompt logging.
Worked example. Metrics dashboard: WAU/MAU ratio target 0.62; tasks completed per active user ≥8/week; override rate ≤10%; median time-to-first-citation <2.5s; lagging: AHT on assisted queue −15% vs matched control after 90 days.
Feedback loops: qualitative and quantitative
Definition. Feedback loops combine telemetry, user interviews, support tickets, QA sampling, and eval failures into prioritized product improvements—closing the loop with users who reported issues.
Engagement use. Weekly triage: top override reasons, failed goldens, support themes. Monthly voice of user session with frontline. Publish you said / we did notes to build trust. Link to backlog reprioritisation.
Pitfalls.
- Collecting thumbs without taxonomy or follow-up.
- Only engineering triages—product absent.
- No closure communication—users stop reporting.
- Overfitting to vocal minority power users.
Worked example. Override taxonomy shows 22% "citation missing appendix"—epic added for appendix chunking; fix shipped in 3 weeks; override reason trend drops 22% → 11% in four weeks post-release.
Release planning, feature flags, and phased rollout
Definition. Release planning sequences deployable increments with eval and change readiness gates. Feature flags control exposure by segment, region, or role—enabling canary releases and instant rollback without redeploy.
Engagement use. Map releases to PRD phases: internal dogfood → pilot BU → regional expansion. Flags for: new model route, new intent pack, auto-send enablement. Each flag has owner, eval requirement, and rollback runbook link (topic 32).
Pitfalls.
- Big-bang go-live without canary.
- Flags without telemetry—cannot compare cohorts.
- Forgetting to retire flags—technical debt and security surface.
- Pilot cohort so small metrics are noise.
Worked example. Release R1: 80 agents in one site, flag lookup_assistant_v1; R2: 620 agents after 21 days with WAU ≥45% and override ≤12%; R3: national after corpus KPI met—each gate in decision log.
Stakeholder alignment and scope negotiation
Definition. Stakeholder alignment makes trade-offs visible before build: who wins, who loses capacity, who owns risk, who signs Won't list.
Engagement use. Run scope negotiation workshop with MoSCoW output published to steering. Pre-wire compliance and union/works council where applicable. Document dissent in decision log—not hidden offline vetoes.
Pitfalls.
- Silent veto after PRD sign-off.
- Scope negotiation only with IT—missing ops and frontline.
- Promising dates without dependency acknowledgment.
- Conflating sponsor enthusiasm with org readiness.
Worked example. Union requires opt-in pilot pools; PRD adds metric "voluntary adoption rate" separate from mandated rollout—prevents gaming and political backlash that kills phase 2 funding.
Competitive and alternative analysis (product lens)
Definition. Product managers compare alternatives users employ today: status quo manual process, search-only, outsourced BPO, non-AI SaaS—not only other AI vendors.
Engagement use. Jobs-to-be-done interview question: "What did you do last time this failed?" Alternatives inform minimum viable improvement bar. If search upgrade solves 70% at 1/5 cost, AI must justify incremental value honestly.
Pitfalls.
- Straw-man "do nothing" alternative.
- Competitor AI demo drives scope without JTBD match.
- Ignoring process fix because AI is funded.
Worked example. Claims team already has OCR pipeline; alternative analysis shows structured extraction reuses OCR output—MVP skips re-scanning epic, saves 9 weeks and £140K build; product manager reframes AI layer as summary on existing JSON, not greenfield doc AI.
Frameworks and methods
MoSCoW for MVP scope
Must: Required for MVP launch and eval gate. Should: Important but deferrable one sprint. Could: Nice-to-have. Won't: Explicitly out—prevents scope creep.
Apply when: Steering pressure expands scope mid-build. Do not apply when: Items are legally mandatory—mark Must with compliance citation.
RICE prioritisation
Reach × Impact × Confidence / Effort. Use Reach as number of users/transactions per quarter; Impact as expected KPI delta; Confidence 0.5–1.0 from discovery evidence; Effort person-weeks including eval and change.
Kano for AI features
Basic: Citations, audit log, override—expected; absence causes dissatisfaction. Performance: Latency, accuracy—more is better. Delight: Proactive suggestions—use sparingly in regulated contexts; can increase risk.
Opportunity Solution Tree (Teresa Torres)
Outcome → opportunities (from discovery) → solutions → experiments. Keeps backlog tied to outcomes, not ideas.
AI PRD template (minimum sections)
| Section | Purpose |
|---|---|
| Problem & job | Link to discovery problem statement |
| Personas & journeys | Primary user, exceptions |
| Scope in/out | MVP vs platform |
| Functional requirements | Retrieve, cite, override, escalate |
| NFRs | Latency, languages, a11y, residency |
| Data & corpus | Sources, freshness, ownership |
| Eval acceptance | Golden metrics, red-team thresholds |
| Metrics | North Star, guardrails, baselines |
| Rollout | Segments, phasing, training |
| Risks & dependencies | RAID hooks |
| Open questions | Spikes with owners |
Definition of Ready / Done (AI-adjusted)
Ready: Persona clear; AC include failure paths; eval cases identified; compliance pre-read scheduled. Done: Eval regression pass; telemetry live; runbook stub; change comms sent; docs updated.
Real-world scenarios
Scenario A: Retail banking servicing assistant (UK)
Context. Tier-1 contact centre 4,200 agents; discovery validated 41% policy lookup contacts; pilot branch 620 agents.
Product choices. PRD v1.2: primary persona = frontline agent; job = resolve policy lookup in one call with auditable citation; MVP = 12 top intents, English, desktop embed in CRM; Won't = payment execution, vulnerable-customer automated advice.
Backlog prioritisation. Must: corpus sync for top intents, retrieval+citation, confidence bands, override logging. Should: suggested reply draft. Could: proactive intent detection. RICE drove Should deferral until override rate stable.
MVP vs platform. Shared telemetry only—no enterprise "AI hub" until WAU ≥48% on pilot.
Adoption metrics (90 days). WAU 51%; assisted lookups 6.8/agent/day; override 9.2% (target ≤10%); citation presence 96.1% on QA sample; lagging AHT on lookup contacts −17% vs control branch.
Outcome. Phase 2 PRD adds Welsh + 8 intents; platform retrieval service approved with FinOps chargeback model (see topic 07).
Scenario B: B2B SaaS support copilot (North America)
Context. $95M ARR SaaS; 38 support engineers; ~14K tickets/month; 44% Tier-1 how-to; documentation stale 9 days avg after release.
Product choices. Persona = support engineer (not customer-facing bot in MVP). JTBD = find correct fix article and draft customer reply with citations. MVP integrates Zendesk + Confluence sync; customer bot Won't until engineer override rate <12% for 60 days.
Discovery → backlog. Epic 1 doc freshness pipeline; Epic 2 engineer copilot; Epic 3 eval on 300 historical tickets; Epic 4 customer deflection—Later.
Metrics. Engineer WAU target 70%; median research time −35% (self-report + ticket timestamps); customer CSAT on assisted tickets ≥ baseline −0.1; escalation to L2 flat.
Numbers at 16 weeks. WAU 74%; research time −31%; override 10.5%; top override = wrong product version doc—feeds doc pipeline epic. CFO sees path to $420K/year capacity equivalent without headcount (links to business case).
Scenario C: Public hospital clinical reference (stretch — regulated)
Context. NHS trust; 1,100 ward nurses; policy = AI decision support only, not diagnosis or prescribing without pharmacist review.
Product constraints. PRD mandates refusal on dosing; citations from approved formulary only; override mandatory log; no patient identifiable data in prompts (redaction layer product req).
MVP. One ward type (med-surg); 4 high-volume interaction checks; tablet workflow.
Adoption. Champion nurses per shift; training tied to MDT governance. Metric: safe usage sessions/week not raw query count. 62% WAU at 8 weeks; zero critical safety escalations on audit sample n=180; benefits tracked as time saved on non-clinical lookup, not "better diagnoses."
Shows product management under hard negative constraints—scope smaller, metrics conservative, platform deferred.
Scenario E: Pharmaceutical medical information (MLR-governed)
Context. Pharma field medical team; 850 reps; MLR-preapproved content only; no generative paraphrase outside approved text library in MVP.
Product implications. PRD defines retrieval + exact quote assembly—not abstractive summarization. Backlog prioritizes version-controlled content sync and audit trail over conversational UX. Adoption metric = compliant lookups/week, not chat engagement.
Numbers. 71% WAU at 10 weeks; zero MLR violations on 240 audit sample; average lookup −3.8 minutes vs manual PDF search; phase 2 generative draft deferred until legal defines bounded paraphrase rules with eval set ≥500 items.
Lesson. Product scope follows regulatory product class—attempting standard "copilot" PRD would fail MLR; product manager reframes as compliant find-and-cite product with different metrics and roadmap.
Product documentation and sales enablement (avoid promise drift)
Product managers maintain external-facing capability bounds so presales (topic 29) does not oversell MVP. Deliver: capability one-pager (in/out, metrics, known limits), FAQ for account teams, and demo script tied to PRD acceptance criteria—not ad-hoc prompts.
Worked example. Presales claims "works in 12 languages"; PRD MVP is English only—product publishes roadmap language gates (FR/DE after EN WAU ≥50%) and trains sales on approved talk track; RFP win rate improves because proposals match deliverable reality.
Product enablement is part of adoption design: if account teams promise capabilities the MVP excludes, frontline users receive broken expectations before day one—override rates and NPS suffer for reasons no model change can fix.
Practice exercises
Primary exercise: Discovery-to-backlog translation + PRD skeleton (3–4 hours)
Use this prompt: Discovery validated a claims FNOL assistant for adjusters—38% of rework is missing document summary; corpus partial on water damage.
Deliver:
-
Two personas with JTBD (primary + secondary).
-
MVP vs platform table — minimum 8 in/out rows with rationale.
-
Prioritised backlog — 2 epics broken into 6 user stories with AI-specific acceptance criteria (cite, override, escalate).
-
PRD skeleton — all template sections filled (one paragraph each minimum).
-
Adoption metric framework — North Star, 3 guardrails, baselines to collect, 90-day targets.
Acceptance criteria: Every backlog item traces to stated discovery finding; at least one story covers low-confidence behavior; Won't list includes at least one sponsor-requested feature explicitly deferred.
Stretch exercise: MVP scope negotiation role-play (2 hours)
Write a one-page steering memo defending MVP scope against a sponsor who demands: (a) customer-facing chatbot, (b) auto-approval of claims under £500, (c) 5-language launch. Use RICE or MoSCoW; cite risk and eval readiness; propose phased roadmap with metric gates.
Acceptance criteria: Memo includes numeric adoption targets and explicit revisit triggers; auto-approval rejected with compliance reasoning; languages phased.
Reflection exercise: Override taxonomy design (45 minutes)
For an assistant you know, draft 8 override reason codes and map each to a backlog item type (corpus, prompt, UX, training, out-of-scope). Define how weekly triage would reprioritise.
Questions you should be able to answer
- What job is the primary user hiring the product for—in their words, not yours?
- Which discovery hypothesis does each MVP epic trace to?
- What is explicitly out of MVP, and who signed that boundary?
- How do users know when the system is uncertain, and what can they do about it?
- What happens when a user overrides an answer—where does that data go?
- What eval thresholds gate release, and who owns red-team sign-off?
- What is your North Star metric vs guardrail metrics?
- What baseline will you measure against pre-launch?
- Why is this MVP thin vertical rather than platform-first—or vice versa?
- How will you know adoption succeeded at 30, 60, and 90 days?
- Which persona is not served in MVP, and when are they phased in?
- What is the top predicted override reason, and which backlog item addresses it?
- How does roadmap Next differ from Later, and what metric unlocks progression?
- What product telemetry requires privacy or union review before switch-on?
- If adoption misses target by 20%, what is the first product lever—not model swap?
Negative cases
Requirements soup. Fifty stakeholder requests in backlog; no MVP. Fix: MoSCoW + signed PRD; defer with decision log.
Demo-driven PRD. Scope mirrors sales demo on clean data. Fix: eval acceptance on client goldens; Won't for unvalidated intents.
Platform tax bankruptcy. Two years building ingestion; no user-facing MVP. Fix: thin vertical with shared telemetry only.
Metric theatre. Dashboard shows queries; handle time unchanged. Fix: outcome-linked North Star; control cohort.
Override black hole. Users correct system; nothing improves. Fix: taxonomy + weekly triage to backlog.
Persona collapse. Product built for manager dashboard; frontline never adopts. Fix: primary persona = daily user; shadow testing.
Trust UX bolt-on. Compliance asks for citations at UAT. Fix: citations in MVP AC from day one.
Roadmap as promise. Dates committed before eval. Fix: outcome gates; aspirational lane labeled.
No Won't list. Scope expands silently. Fix: publish Won't in steering; revisit triggers only.
Feedback fatigue. Thumbs with no follow-up. Fix: you said / we did loop.
AI replaces training. Product launches without change plan (see topic 26). Fix: adoption workstream in PRD.
Multi-BU premature merge. One MVP forced for all regions with conflicting corpus. Fix: sequence BUs after first metric proof.
Operating model: discovery handoff to delivery
Product management sits at the handoff between consulting-shaped discovery and engineering-shaped delivery. Without an operating model, PRDs arrive late, backlogs swell, and steering loses traceability.
Handoff ceremony (recommended). Within one week of discovery exit, run a translation session with product owner, tech lead, design, compliance liaison, and client problem owner. Inputs: opportunity scorecard, hypothesis register (supported/rejected), process maps, constraint list. Outputs: signed MVP boundary, backlog v1, open spikes with owners, metric baselines to collect in next two weeks.
Roles.
| Role | Product accountability |
|---|---|
| AI Solution Engineer / PM | PRD, backlog, prioritisation, metric framework |
| Client product owner | Scope sign-off, adoption sponsorship, Won't negotiations |
| Tech lead | Estimation, spike outcomes, eval harness feasibility |
| Design / UX | Journey, trust patterns (pairs with topic 23) |
| Compliance / risk | Negative requirements, eval gates |
| Change lead (topic 26) | Training plan inputs, champion map |
Cadence. Weekly backlog refinement with eval failure review; bi-weekly steering slice showing metric trends not only burn-down; monthly roadmap refresh with explicit deferrals.
Artefact versioning. PRD v1.0 = MVP commit; v1.x = additive; v2.0 = new phase only after adoption gate. Link every version to decision log entry—prevents "scope whisper" via Slack.
Discovery artefact → backlog mapping (worked table)
| Discovery output | Backlog artefact | Example |
|---|---|---|
| Supported hypothesis H2 | Epic | "Lookup assistant for top 12 intents" |
| Rejected hypothesis H5 | Won't + revisit trigger | "Auto payment changes — revisit if override <6%" |
| Process map pain point | User story | "Reduce re-keying on loss run screen" |
| Compliance constraint | Non-functional AC | "No unsourced payment advice" |
| Opportunity score RICE | Priority rank | Epic order in sprint planning |
| Open data question | Spike | "Corpus completeness audit — 5 days" |
Experiment design for AI products
Not every uncertainty needs a six-month build. Product managers specify experiments with time boxes: concierge MVP (human behind curtain), shadow mode (AI suggests, human sends), A/B on cohorts, wizard-of-oz for intent validation.
Experiment card template: hypothesis, success metric, duration, sample size, ethical/compliance guard, build vs manual effort, decision rule (ship/kill/pivot).
Worked example. Before customer-facing bot, run 4-week shadow mode in engineer console: measure whether draft replies would reduce handle time without sending—target ≥25% time delta on sample n=400 tickets; if <10%, kill customer bot epic; invest in doc pipeline (Scenario B pattern).
Platform product patterns (when MVP succeeded)
After MVP adoption gates pass, platform investments must still be product-managed, not "build it and they will come."
Platform PRD differences: internal customers as personas; SLAs for ingestion latency; self-service onboarding checklist; FinOps chargeback hooks (topic 22); golden path templates; versioned API contracts.
Anti-pattern: platform team measures components shipped; zero BUs onboarded. Fix: platform OKR = adopted products on shared services and £/success improvement vs bespoke stacks.
Instrumentation and analytics implementation guide
Product metrics fail when instrumentation is an afterthought. Specify events in PRD before sprint one.
Minimum event schema: session_id, user_id (hashed role), intent, corpus_version, model_route, confidence_band, citation_count, override_flag, override_reason, task_success, latency_ms_total, tokens_in/out.
Privacy: legal review on prompt logging; default metadata-only logging in production with sampled content capture for eval.
Dashboards: product (adoption), FinOps (cost per success), quality (eval smoke), ops (error rate)—same trace ID across all four.
Scenario D: Energy utility field-service scheduling assistant (stretch)
Context. EU utility; 1,400 field technicians; scheduling calls drive 23% of dispatcher time; discovery shows 31% calls are "where is my slot / what's the scope"—not complex optimization.
Product decision. MVP does not optimize routes with OR-Tools in phase one—sponsor wanted "AI optimizer." PRD Won't: autonomous schedule changes; MVP does retrieve job pack + customer history + SLA text with citations in mobile app.
Backlog. Epic 1 mobile retrieval UX; Epic 2 integration work order API; Epic 3 dispatcher console mirror; Epic 4 optimization — Later gated on ≥55% technician WAU and dispatcher time −20%.
Metrics at 14 weeks. Technician WAU 58%; dispatcher assisted calls −24%; override 7.8% (mostly stale job status—integration fix); zero unauthorized schedule mutations.
Lesson. Product manager protected scope against technically exciting but unvalidated optimization—discovery hypothesis was information access, not NP-hard routing.
Integration with Stage 5 topics
Product management is the hub for Stage 5 delivery:
- Topic 07: PRD metrics become benefits register baselines; MVP scope drives TCO build/run split.
- Topic 22: Success definition in PRD defines cost per successful task; performance budgets are non-functional requirements.
- Topic 23: Trust UX flows are user stories with acceptance criteria—not design-only polish.
- Topic 25: Roadmap now/next/later aligns to programme milestones and RAID dependencies.
- Topic 26: Adoption plan consumes persona list and champion metrics from PRD.
- Topic 21: Eval acceptance criteria are release gates in Definition of Done.
Weak integration symptom: business case cites 50% time savings while PRD measures logins—align in handoff ceremony.
Related playbook content
- AI Product Management roadmap — end-to-end product leadership path
- The Complete Product Manager Roadmap — deep article companion
- AI Product Builder roadmap — builder-oriented product delivery
- The Complete AI Product Builder Roadmap — hands-on product build guide
- Adoption guide — change, training, and usage measurement
- Opportunity Discovery — upstream discovery workshops and artefacts
- Business case and prioritisation — fund the roadmap with credible economics
- AI Opportunity Discovery — scored opportunity shortlist input
- AI Evaluation and Quality Assurance — eval gates in PRD
- User Experience and Human Factors — trust and interaction design
- Commercial and Financial Modelling — benefits linked to adoption metrics
- 8D Framework — stage gates for product evidence
- How to use this Learning Map — study loop and artefact standards
Practice checklist
- I can explain the primary user's job without naming AI or vendors
- Every MVP epic links to a discovery outcome or explicit learning goal
- PRD includes negative requirements and eval release gates
- MVP vs platform boundaries are signed with at least five explicit Won'ts
- User stories have acceptance criteria for cite, override, and escalate paths
- North Star and guardrail metrics have baselines and 90-day targets
- I completed the primary exercise and filed artefacts in my pattern library
- I can describe what happens to override data within one week of capture
Discussion
Comments
Share feedback or questions about this page. No account required.
Loading comments…