Leadership and People Management
Executive view
Assess leaders on decision velocity, team health and succession—not personal heroics. Ask: Are decision memos used? Is coaching developing client-side capability? Are ethics and quality issues escalated early?
Decision required: Does this leader leave the organisation stronger after the engagement ends?
Technical view
Practice lead-without-authority: decision criteria, ADRs, time-boxed spikes, feedback with behaviour examples. Build skills matrices before asking HR for headcount.
Pair with [Career and capability roadmap](/ai-solution-engineering/career-roadmap) for role expectations at each level.
Why this matters
AI programmes combine research uncertainty with enterprise delivery pressure. Teams fail when:
- Seniors debate architecture indefinitely while juniors wait
- Leaders solve instead of develop others—bus factor stays at one
- Performance conversations avoid eval metrics and client outcomes
- Psychological safety is low—teams hide bad benchmarks until launch
- Staffing mixes three ML PhDs and zero integration engineers for an ERP copilot
- Leaders avoid escalating risk to protect relationships—until Sev 1
Leadership in this context is not generic "people skills." It is structuring ambiguity so multidisciplinary teams can decide, learn and deliver—while connecting to People and Organisational Capability and Exceptional Culture at firm scale.
Strong people leadership produces decision logs, growing juniors, staffing fit for use case, and early escalation of blockers. Weak leadership produces deadlock, burnout, cover-ups, and client dependency on one consultant.
Learn
Leadership vs management in AI delivery
| Management | Leadership |
|---|---|
| Plan, allocate, track | Set direction and meaning |
| Process compliance | Judgment under uncertainty |
| Resource assignment | Develop people and coalitions |
| Status reporting | Escalate with options |
You may manage without leading (checkbox PM) or lead without managing (principal IC). AI Solution Engineers at senior levels need both lenses—clarity on outcomes and care for team capacity.
Delegation — outcomes, not tasks
Bad delegation: "Implement RAG" with no acceptance criteria.
Good delegation: "Deliver faithfulness ≥88% on Eval-v2 by Friday; constraints: no PII in logs; escalate if retrieval empty rate >5%."
Levels of delegation:
- Do exactly this — Junior on well-defined script
- Research and recommend — Mid on options paper
- Decide and inform — Senior on bounded domain
- Own outcome — Lead on workstream with client interface
Match level to skill and risk tier—do not delegate tier-high launch approval to unbriefed junior.
Coaching — grow problem-solvers
Coaching differs from teaching (transfer knowledge) and mentoring (career counsel).
GROW model adapted for AI:
- Goal — What decision or skill this week?
- Reality — Current eval, blocker, stakeholder map
- Options — Architecture paths, spike scopes
- Will — Commitment, date, escalation trigger
Coaching questions:
- "What evidence would change your recommendation?"
- "If you were Decide role, what would you choose?"
- "What's the smallest spike to de-risk this?"
Avoid answer dumping unless incident Sev 1—then coach after contain.
Feedback — behaviour, impact, future
SBI framework: Situation, Behaviour, Impact—plus requested change.
Example: "In yesterday's client demo (S), you improvised accuracy claims we haven't validated (B), which put legal on alert and delayed sign-off (I). Next demo, stick to approved metrics or flag gaps 24h ahead."
Positive feedback must be specific—"great job" reinforces nothing.
Cadence: Weekly 1:1 25–30 min; no status-only meetings—status belongs in tools.
Leading without authority
Common when consulting or matrixed:
| Lever | Action |
|---|---|
| Credibility | Accurate summaries; admit unknowns |
| Pre-work | Decision memos before steering |
| Coalitions | Align process owner before exec session |
| Reciprocity | Help risk team with control language |
| Transparency | Decision log; no back-channel surprises |
See topic 27 for stakeholder maps; leadership adds team dimension inside coalition.
Managing seniors and experts
Seniors resist when they feel overruled without criteria or used as staff aug.
Tactics:
- Frame as peer problem-solving—"We need a decision by Thursday; here are criteria."
- Public credit for their architecture; private challenge on gaps
- Time-box disagreement—spike, not debate
- ADR with named author—even if you decide
Never humiliate expert in client forum—lose team forever.
Decision-making under uncertainty
Pattern:
- Define decision and Decide role (RAPID)
- Set criteria (eval, cost, risk, time)
- Time-box spike if evidence missing
- Decide; document ADR; communicate dissent
- Review date—revisit if assumptions wrong
Example deadlock: agent vs deterministic workflow. Criteria: change frequency, audit requirement, integration count. Spike 3 days both on one journey; score; decide.
Staffing AI teams — skills matrix
| Skill area | Roles | Typical gap signal |
|---|---|---|
| ML / eval | DS, ML engineer | Good offline metrics; prod blind |
| App / integration | Backend, full-stack | API works; no auth on tools |
| Data engineering | DE | Pipeline late; corpus dirty |
| Platform / SRE | Platform, SRE | No rollback; alert fatigue |
| Security / privacy | Security architect | Late veto |
| Change / UX | Change lead, UX | Low adoption |
| Domain | BA, SME | Wrong workflow automated |
Staffing ratio varies by use case—RAG-heavy needs DE + eval; agent-heavy needs integration + security.
Honest matrix: RAG green, agent auth red → hire or upskill before promising agent scope.
Performance management — connect to outcomes
Performance = behaviours + outcomes aligned to client and team goals.
AI-specific outcomes:
- Eval thresholds met sustainably
- Incidents contained per runbook
- Knowledge transferred to client
- Reusable patterns contributed
- Ethics issues escalated appropriately
Underperformance: Early clarity on gap; support plan 30 days; document; escalate to HR per firm policy—not surprise at review cycle.
Psychological safety — Amy Edmondson applied
Team members speak up when:
- Bad eval news is rewarded with problem-solving, not blame
- Questions in client call are welcomed
- Incidents trigger PIR, not scapegoating
- Dissent recorded in decision log without career penalty
Leader behaviours: Model "I don't know"; thank reporter of bad news; never punish messenger.
Measure: Retrospective prompts; optional anonymous pulse; track eval disclosure latency (how fast bad results rise).
Workforce planning and hiring
Before headcount request:
- Skills matrix gap analysis
- Duration of need (engagement vs permanent)
- Location / clearance constraints
- Diversity of thought—avoid clone hiring
Interview loop for AI roles: practical eval design exercise, integration scenario, ethics question—not only LeetCode.
Escalation — early, with options
Escalate when: blocker >5 days; risk tier mismatch; ethical line approached; sponsor nominal; team capacity broken.
Escalation memo: BLUF, impact, options, recommendation, Decide role—≤1 page.
Anti-pattern: Vent to peer without memo—wastes political capital.
Calm under pressure — incident and launch
During Sev 1 or launch week:
- Single incident commander
- Communicate on schedule even if "still investigating"
- Protect team from random executive scope injection
- Debrief after—sleep deprivation decisions get review date
Leader composure is contagious—panic spreads faster than incidents.
Frameworks and methods
Decision memo template
Title: [Decision needed by DATE]
BLUF: Recommend Option B because [criteria].
Context: 3–5 sentences.
Options: A / B / C with pros/cons table.
Criteria: Eval, cost, risk, time, alignment.
Recommendation: B.
Dissent: [Name] prefers A because [X]—accepted/rejected because [Y].
Review: Revisit if [assumption] false by [date].
ADR (Architecture Decision Record)
Status, context, decision, consequences. Link from decision memo for technical choices.
Skills matrix — RAGY scoring
| Skill | R required | A available | G growing | Y gap |
|---|---|---|---|---|
| Eval design | 1 FTE | ✓ | ||
| Agent tool auth | 0.5 FTE | ✓ |
Action column: hire, train, contractor, defer scope.
1:1 agenda template
- Their agenda (10 min)
- Blockers and decisions needed (10 min)
- Coaching on one growth goal (10 min)
- Feedback exchange (5 min)
Performance conversation structure
- Review outcomes vs goals
- Strengths with examples
- One development priority
- Support from leader
- Document agreement
Team charter for AI squad
- Mission and non-goals
- Decision rights (RACI summary)
- Working agreements (eval honesty, no Friday tier-high deploy)
- Ceremonies (stand-up, retro, eval review)
- Definition of done including eval and docs
Conflict resolution — security vs delivery speed
- Name shared goal (launch safely with acceptable risk)
- Separate must-haves vs nice-to-haves
- Third option creative (phased launch, reduced autonomy)
- Decide role speaks—not longest argument wins
- Record; schedule review
Pairs with topic 27 conflict protocol.
Real-world scenarios
Scenario A — Deadlock: agent vs workflow automation
Context: Global insurer; claims FNOL intake; Senior Architect A pushes autonomous agent; Senior Architect B pushes deterministic workflow with LLM assist only. 12-person team idle 9 days; client demo in 15 days.
Leadership intervention:
- Decision memo criteria: audit trail, change frequency, MLR approval, time
- 3-day spike: both on one FNOL path; score on eval + audit + build time
- RAPID: client AI lead Decide; you Recommend after spike
- Result: Workflow + LLM assist wins; agent deferred phase 2
- ADR published; A credited for agent research roadmap
Outcome: Demo proceeds; team velocity restored; dissent documented.
Numbers: Idle cost ~£47k (9 days × 12 FTE blended); spike cost £11k—cheaper than continued deadlock.
Scenario B — Staffing: ERP copilot understaffed on integration
Context: Manufacturer; SAP PM copilot; team 4 ML, 1 app dev; integration backlog 6 weeks; client COO escalating.
Skills matrix: Integration red; eval green; change amber.
Intervention:
- Honest client conversation: scope cut to read-only SAP queries phase 1
- Request 2 integration engineers for 10 weeks via partner
- Upskill one ML engineer on SAP OData pairing with client dev
Outcome: Launch 4 weeks late vs original impossible date; faithfulness 89%; client COO accepts phased roadmap.
Numbers: Partner cost €180k vs €400k+ rework if launched with broken integration; adoption 62% at 90 days vs projected 35% under full-scope failure.
Scenario C — Psychological safety: hidden eval failure
Context: Bank copilot; junior DS discovers faithfulness 74% on new eval set Monday; previous lead shouted down "bad news" in retro month prior; DS stays silent until Thursday client steerco.
Impact: Steerco surprise; sponsor trust damaged; launch delayed 3 weeks.
New lead intervention:
- Public retro rule: eval metrics first; thanks for bad news
- Private coaching former lead on behaviour (SBI feedback)
- Weekly eval review ceremony with no blame
- DS presents fix plan in steerco—credibility restored
Outcome: Culture shift slow but measurable—next eval drop reported same day.
Numbers: 3-week delay cost ~$220k; compare to potential Sev 2 if launched at 74%.
Scenario D — Escalation: nominal sponsor
Context: Public sector digital programme; £3.2m AI budget; sponsor (CDO) missed 4 steercos; CISO conditional approval expiring 45 days; usage 28% vs 60% target.
Escalation memo to buyer (Permanent Secretary delegate):
- BLUF: Programme at amber-red; recommend extraordinary steerco
- Options: (A) continue with new sponsor; (B) pause and resize; (C) terminate vendor
- Recommend A with named active sponsor (COO)
- Risk of inaction: wasted £890k spend to date
Outcome: COO becomes sponsor; CISO extension granted; adoption plan reset with union input.
Numbers: Usage climbs to 54% in 90 days after sponsor active; avoided £1.1m write-off scenario in option C analysis.
RACI for AI squad — worked example (production RAG)
| Activity | Eng lead | ML | DE | App | Security | Client PO |
|---|---|---|---|---|---|---|
| Weekly eval review | A | R | C | I | I | C |
| Prompt change (safety) | A | R | I | C | C | I |
| Corpus ingest | I | C | R | I | C | A |
| Production launch | C | C | C | R | C | A |
| Incident harmful output | R | C | I | C | C | I |
A = exactly one per row.
Influence without authority — case playbook
Situation: Client IT blocks API access; business escalates to you as "lead."
Steps:
- Map Decide role—usually client CIO delegate, not you
- Memo: cost of delay £X/week; options with integration effort
- Coalition: business sponsor + user champion pre-call
- Offer spike to reduce IT uncertainty—time-boxed
- Never bypass IT in production path—credibility preserved
Performance improvement plan — AI role example
Gap: Engineer repeatedly merges prompts without eval.
Plan 30 days:
- Week 1: Pair on eval harness; shadow merge blocked
- Week 2: Solo merge with pre-merge eval checklist signed
- Week 3: Independent merges; spot audit 100%
- Week 4: Review; close or extend with HR
Measurable behaviours—not personality.
Hiring interview loop — AI Solution Engineer
| Stage | Assessor | Exercise |
|---|---|---|
| Screen | Recruiter | Role fit |
| Technical 1 | Architect | Architecture whiteboard |
| Technical 2 | ML lead | Eval design for use case |
| Behavioural | Engagement lead | Stakeholder scenario |
| Values | Partner | Ethics red line question |
Scorecard weighted equally technical/behavioural—no "brilliant jerk" hire.
Practice exercises
Primary exercise — Decision memo for security vs speed (45 minutes)
Brief: Retail bank; marketing wants public LLM for campaign copy generation; CISO blocks external API with customer segments data.
Tasks:
- Write decision memo: BLUF, three options, criteria, recommendation, Decide role.
- Define 3-day spike scope if evidence needed.
- One paragraph ADR summary for technical path.
- Note who to coach vs inform.
Acceptance criteria:
- ≤500 words memo body
- No villain narrative
- Dissent path documented
- Criteria include risk tier and eval where relevant
Stretch exercise — Team charter and skills plan (half day)
Brief: Healthcare analytics squad 14 FTE; predictive readmission model + clinician copilot; 6-month timeline; mixed vendor and client staff.
Tasks:
- Skills matrix RAGY with gaps.
- Staffing plan: hire, upskill, partner, defer scope.
- Team charter: mission, non-goals, working agreements, ceremonies.
- Coaching plan for one underperforming integration engineer (fictional but realistic).
- 1:1 template customised for AI squad.
- Link one growth goal to Career and capability roadmap competency.
Acceptance criteria:
- Non-goals explicit (e.g. no autonomous diagnosis)
- Psychological safety working agreement included
- Performance metric tied to client outcome
Scenario G — Merger integration two AI teams
Context: Client merges Bank A + Bank B; two copilot teams; hostile culture; duplicate platforms £1.8m/year.
Leadership:
- Neutral integration charter—best eval wins architecture
- No layoff promises without HR—but honest role mapping
- Joint retro on fear; executive steerco on single sponsor
Outcome: One platform 11 months; £1.1m annual save; 23 voluntary redundancies with packages—handled with union.
Numbers: Combined faithfulness 89% vs 84% / 87% pre-merge best-of-breed hybrid.
Scenario H — Leading up when client PM is toxic
Context: Client PM belittles your juniors in calls; scope creep daily.
Actions:
- Document 3 incidents with dates
- Escalate to your partner with facts—not emotion
- Request steerco restructure with senior client counterpart
- Protect team: you join all client calls until behaviour shifts
Retention of 2 juniors saved £180k rehire/training vs doing nothing.
Leadership reading companion — map cross-links
| Topic | Leadership connection |
|---|---|
| 27 Stakeholder | Influence without authority |
| 28 Communication | BLUF, steerco |
| 26 Change | Workforce honesty |
| 34 Ethics | Escalate harm |
| 35 Personal | Sustainable pace |
Topic 33 integrates prior map—leadership is how you deploy technical knowledge through people.
FAQ — leadership and people management
Q: How decide without authority?
A: Decision memo + RAPID Decide role + time-boxed spike—never impose unless client charter gives you sign-off.
Q: How often 1:1s?
A: Weekly for direct reports; fortnightly for extended team during crunch with explicit end date.
Q: What if senior refuses decision criteria?
A: Escalate to engagement lead/client sponsor with facts—document repeated deadlock as programme risk.
Q: How measure psychological safety?
A: Leading indicators: eval disclosure speed, retro participation, incident reporting—optional anonymous pulse.
Q: Coaching vs doing—where's the line?
A: Sev1 you lead; otherwise 24-hour rule—if they haven't progressed, pair; if repeated, direct with clearer spec.
New lead first 30 days checklist
- Meet each team member 1:1
- Read last 10 decision logs / ADRs
- Skills matrix draft
- Sponsor alignment call
- Retro facilitation with safety norms set
- Identify one quick win delegation
- Link personal development to Career and capability roadmap
Team health indicators — review monthly
| Indicator | Green | Red flag |
|---|---|---|
| Attrition risk (1:1 sense) | Stable | 2+ key people interviewing |
| Eval honesty | Bad news same day | Hidden >48h |
| Decision velocity | ADR/week healthy | Same debate 3+ weeks |
| Overtime | Sprint spikes only | 4+ weeks sustained |
| Client feedback | Constructive | Personal attacks on team |
Red flags trigger escalation to engagement partner—not ignore until quit.
Decision log — shared team discipline
Every major decision logged:
ID: DEC-2026-014
Date: 2026-07-28
Decision: Workflow not agent for FNOL phase 1
Decide: Client AI lead
Recommend: Engagement architect
Dissent: Senior A (agent) — deferred to phase 2 eval gate
Review: 2026-11-01 post-pilot
Searchable wiki—new joiners read history not re-debate.
Coaching conversation worked example (GROW)
Goal: Mid ML engineer leads eval review confidently by month end.
Reality: Presents metrics but avoids interpreting failures; waits for lead.
Options: (A) Shadow lead 2 reviews then co-lead; (B) Own one metric deep-dive; (C) Training course only.
Will: Option A+B—co-lead review 2026-08-12; prepare interpretation doc 2026-08-05.
Document in coaching plan—review in 1:1 2026-08-14.
Staffing AI programme — FTE estimation heuristic
| Phase | Typical FTE multiplier on core build team |
|---|---|
| Discovery | 0.5× |
| Build | 1.0× |
| Pilot | 0.8× |
| Scale | 0.6× + ops 0.3× |
Adjust for tier-high compliance +0.2× security/governance across phases.
Use with skills matrix—FTE without skills is meaningless.
Psychological safety — team working agreements (sample bullets)
- We report eval drops same day—no hiding to "fix later"
- We disagree in meetings with criteria, not personal attack
- We document decisions—no secret architecture in DMs
- We escalate ethics concerns before launch, not after incident
- We protect sleep—no heroics as default culture
Post in team channel; revisit retro quarterly.
Leading multidisciplinary retro — AI squad agenda (45 min)
- Eval metrics truth (10 min) — no blame
- What shipped / learned (10 min)
- Stakeholder or ops surprise (10 min)
- One process tweak (10 min) — owner named
- Appreciation (5 min)
Facilitator rotates—builds leadership bench.
Engagement lead weekly rhythm (reference)
| Day | Leadership focus |
|---|---|
| Mon | Priority card; sponsor pulse |
| Tue | 1:1s batch |
| Wed | Decision log review; unblock |
| Thu | Client steerco prep; coalition |
| Fri | Retro or team health; weekly review |
Adjust for travel—minimum preserve 1:1s and decision log review.
Mentoring across the Learning Map — 12-week plan
| Week | Mentee topic | Artefact |
|---|---|---|
| 1–2 | 05 Discovery | Opportunity brief |
| 3–4 | 13 RAG | Architecture sketch |
| 5–6 | 21 Eval | Eval plan |
| 7–8 | 27 Stakeholder | RACI |
| 9–10 | 30 Contract | Acceptance annex |
| 11–12 | 33 Leadership | Decision memo |
Leadership includes developing others through the map—not solo completion.
Performance calibration — AI squad contribution dimensions
When calibrating team performance, weight:
- Eval integrity — honest metrics, no gaming
- Client outcome — adoption, SLA, incident handling
- Collaboration — cross-discipline, psychological safety
- Reuse — patterns, docs, runbooks left behind
- Growth — skills matrix movement quarter over quarter
Avoid single metric (lines of code, demos)—AI delivery is multidimensional.
Sponsor relationship — engagement lead responsibilities
- Brief sponsor before steerco—never surprise with bad eval news
- Provide BLUF email after major incidents within 24h
- Ask sponsor monthly: "What blocker can you remove?"—document answer
- Escalate nominal sponsorship per topic 27 protocol—not after month six
Leadership is upward management of sponsorship quality—not only downward team care.
Readiness for people leadership — self-assessment
Rate 1–5: delegation, difficult feedback, decision memos, coaching others, escalating ethics, staffing honesty, calm in incidents. Any ≤2 is learning backlog priority before taking engagement lead on tier-high programme—protect team and client from hero-IC pattern.
Questions you should be able to answer
- What decision am I uniquely positioned to make this week?
- Who on the team needs coaching vs teaching vs delegation?
- What is blocked, and who is the Decide role for escalation?
- Does our skills matrix match the use case—not generic "AI team"?
- When did I last give specific negative feedback with SBI?
- Is psychological safety high enough that bad evals surface same day?
- What deadlock needs a time-boxed spike and decision memo?
- Who are my seniors, and how do I disagree without humiliation?
- What risk am I avoiding escalating, and what is the cost of delay?
- What outcomes define performance for each role on the squad?
- What is in the team charter for eval honesty and deploy discipline?
- How do I lead without authority on this engagement?
- What succession or client capability transfer am I driving?
- How do I stay calm and structure incident leadership?
- What development goal links to career roadmap for my mentee?
Building psychological safety — 90-day plan for new leads
Days 1–30: Model vulnerability—admit mistakes in retros; thank first person who surfaces bad eval.
Days 31–60: Introduce eval-first retro agenda item; no action items without owner.
Days 61–90: Measure time-to-report eval drops; celebrate reduction; 360 or pulse optional.
Track leading indicators: questions in client calls, dissent in decision logs, incident reports without blame.
Staffing AI programmes — role definitions sample
| Role | FTE | Accountability |
|---|---|---|
| Engagement lead | 0.5 | Client outcomes, steering |
| AI architect | 1 | ADRs, eval strategy |
| ML engineer | 2 | Eval harness, model integration |
| Data engineer | 2 | Corpus, pipeline, index |
| App engineer | 2 | API, UI, tool auth |
| Security architect | 0.3 | Threat model, gates |
| Change lead | 0.5 | Adoption, training |
| Product owner (client) | 0.5 | Backlog, acceptance |
Adjust ratios: RAG-heavy → more DE; agent-heavy → more app + security.
Delegation ladder exercise for leads
Each direct report: classify current tasks into levels 1–4; one task upgraded one level this month with support plan.
Document in coaching plan—review in 1:1.
Difficult conversation preparation — feedback script
Before SBI delivery:
- Intent: development not punishment
- Facts: specific instances with dates
- Ask: their perspective 2 min uninterrupted
- Agree: one behaviour change + support
- Follow-up: date in calendar
For performance plan territory—involve HR early; never solo improvisation.
Multidisciplinary ritual design
| Ceremony | Purpose | Cadence |
|---|---|---|
| Stand-up | Blockers | Daily 15m |
| Eval review | Quality truth | Weekly 30m |
| Architecture forum | ADRs | Fortnightly |
| Retro | Safety + process | Fortnightly |
| Steerco prep | Alignment | Weekly internal |
Eval review is non-optional for AI squads—prevents hiding bad news.
Leading client-side capability transfer
Goal: client can operate without you at engagement end.
| Activity | When |
|---|---|
| Pair client engineer on eval harness | Month 2+ |
| Client presents steerco section | Month 4+ |
| Runbook co-authorship | Pre go-live |
| Office hours → ticket handoff | Hypercare week 3 |
Success metric: client-owned eval run two consecutive weeks before exit.
Career development conversations — link to roadmap
Use Career and capability roadmap competencies in 1:1:
- "For Principal path, you need two client steercos and one published pattern—let's plan."
- Document in coaching plan; revisit quarterly.
Avoid vague "be more strategic"—name artefacts and dates.
Scenario E — Cross-cultural team friction
Context: Global programme; UK delivery lead; India engineering centre; US client; video calls tense; UK team feels "work thrown over wall."
Intervention:
- Team charter timezone overlap rules—core hours 4h shared
- Definition of done includes docs—not only code
- Rotate demo lead by region weekly
- Decision memos written before calls—reduce misunderstanding
Outcome: Velocity +18% over 8 weeks; attrition risk down in pulse.
Numbers: Rework hours 120/month → 45/month; client NPS +12 points.
Scenario F — High performer toxic to juniors
Context: Star ML engineer; eval scores best on team; interrupts juniors publicly; two juniors request transfer.
Intervention:
- SBI feedback with specific incidents
- Code review norms enforced—no public humiliation
- Pair high performer as mentor with success metric = mentee demo
- HR engaged when behaviour repeats
Outcome: One junior retained; high performer moves to research spike with less client contact—role fit discussion.
Lesson: Performance ≠ values; leaders protect team health over solo hero metrics.
Reading list integration — leadership articles
Deepen with firm content after exercises:
- Leadership: Remove Delivery Blockers — when to escalate for exec air cover
- Leadership: Make Risk, Governance and Assurance Decisions — tier-high judgement
- The Complete Engineering Manager Roadmap — EM transition craft
Monthly leadership rhythm for AI engagement leads
| Week | Focus |
|---|---|
| 1 | Skills matrix update; hiring pipeline |
| 2 | Decision log review; unresolved deadlocks |
| 3 | Client sponsor relationship; escalation scan |
| 4 | Team retro; coaching plan progress; ethics pulse |
Consistency beats episodic team dinners.
Extended Learn — situational leadership for AI teams
| Team maturity | Leadership style | Example |
|---|---|---|
| D1 Low skill, high will | Directing | Junior first eval task with checklist |
| D2 Some skill, low confidence | Coaching | Mid after first failed client demo |
| D3 High skill, variable commitment | Supporting | Senior bored on maintenance |
| D4 High skill, high will | Delegating | Tech lead owns architecture stream |
AI teams mix D1–D4 on same squad—avoid one-size management.
Extended Learn — remote and hybrid leadership
- Documentation-first decisions—verbal-only creates timezone injustice
- Record architecture forums for absent members
- Overlap hours sacred for pair programming on eval harness
- Visit client site at least once per major phase—trust accelerates
Topic 33 applies globally distributed AI delivery common in 2026.
Extended Learn — succession when lead rotates off
30-day rotation plan:
- Week −4: Shadow successor leads internal stand-up
- Week −2: Successor runs steerco prep with review
- Week −1: Decision log handover; sponsor intro
- Week 0: Lead on-call backup only
- Week +2: Lead available office hours
Client retention depends on smooth rotation—not single hero.
Extended Learn — diversity in AI team staffing
Diverse teams catch bias, UX, and ethics gaps earlier:
- Include non-ML voices in eval design review
- Include junior "beginner mind" in user acceptance
- Avoid all-male team on HR or health copilot without review
Staffing plan documents D&I goals per firm policy—impacts quality not only HR metric.
Extended Learn — leadership under client political stress
When client internal politics turn toxic:
- Do not take sides between executives
- Do document decisions and criteria in writing
- Escalate to your engagement partner if pressured to bypass controls
- Protect team from client blame language in calls—redirect to facts
Firm Leadership: Protect Quality, Independence and Trust relevant.
Extended Learn — coaching for client-side leaders
Sometimes you coach client product owner—not only your team:
- Teach eval interpretation
- Teach steering pack BLUF (topic 28)
- Teach when to escalate to CISO
Client capability transfer is leadership outcome—billable and valued.
Negative cases — when people leadership fails
Hero lead
Symptom: Lead codes every fix; team stops thinking.
Impact: Attrition; client dependency; leader exit collapses programme.
Fix: Delegation with criteria; coach through options; measure team decisions.
Avoidance
Symptom: Conflict "offline" forever; no Decide.
Impact: Launch gate explosion; blame culture.
Fix: Decision memo; RAPID; time-box.
Clone staffing
Symptom: Team all ML backgrounds; no integration or change.
Impact: Beautiful model; no adoption; security late.
Fix: Skills matrix before hiring; diverse roles.
Shoot the messenger
Symptom: Bad eval reporter punished; silence next time.
Impact: Launch disaster; cover-up culture.
Fix: Thank messenger; PIR blameless; leader models vulnerability.
Status-only 1:1s
Symptom: No coaching, feedback or growth.
Impact: Stagnation; surprise performance issues.
Fix: Agenda template; development goals documented.
Nominal leadership
Symptom: Title without escalation or sponsor challenge.
Impact: Programme drifts; wasted spend.
Fix: Escalation memo; programme risk register.
Practice checklist
- I wrote a decision memo with criteria and dissent
- I completed a skills matrix honest about gaps
- I practiced SBI feedback on a real or role-play scenario
- I can explain psychological safety with team behaviours
- I identified lead-without-authority levers for my context
- I linked one mentee goal to career roadmap
- I documented one people leadership negative case
Related playbook content
- Career and capability roadmap — Role levels and competency expectations
- People and Organisational Capability roadmap — Firm talent systems
- Leadership: Develop People and Organisational Capability — Deep people leadership article
- Exceptional Culture roadmap — Culture as operating system
- Leadership: Create an Exceptional Culture — Psychological safety at scale
- Stakeholder Management — Influence and decision rights
- Communication and Executive Articulation — Executive memos and BLUF
- Delivery and Programme Management — Programme governance
- Leadership — Stage hub and exit criteria
- Engineering Manager roadmap — EM craft crossover
- How to use this Learning Map — Reference-depth study method
Discussion
Comments
Share feedback or questions about this page. No account required.
Loading comments…