RFP, Procurement and Contracting
Executive view
Focus on whether proposals commit to measurable outcomes (faithfulness, latency, availability, human escalation)—not marketing superlatives. Ask: who signs acceptance, what happens when the model provider changes terms, and whether liability matches the risk tier.
Decision required: Are evaluation criteria weighted toward client outcomes and exit portability—not vendor demo polish?
Technical view
Translate architecture into contract language: data flows, subprocessors, IP in prompts and fine-tunes, rollback obligations, observability access, and SLAs tied to SLOs you can actually monitor.
Never sign technical schedules you cannot test. Draft acceptance tests alongside the SOW, not after build.
Why this matters
AI procurement fails in predictable ways. Vendors promise "enterprise-grade accuracy." Business sponsors paste those phrases into RFPs. Delivery teams inherit commitments that no eval suite can prove. Legal teams insert liability caps that leave the client exposed when a model hallucinates in a regulated workflow. Procurement optimises for licence cost while ignoring retraining, egress, and exit migration.
An AI Solution Engineer sits between these functions. You are not a lawyer, but you must translate probabilistic systems into contract-grade language—thresholds, test methods, escalation paths, and explicit exclusions. Weak contracting produces:
- Acceptance deadlock — UAT fails but contract says "substantially complete"
- Scope creep without CR — "just add another corpus" with no change control
- IP disputes — who owns prompts, eval sets, fine-tunes, and synthetic training data
- SLA theatre — 99.9% uptime on the API while retrieval quality collapses
- Vendor lock-in — no exit clause when the model provider deprecates a deployment
Strong contracting produces testable acceptance, aligned incentives, clear IP and data rights, and commercial terms linked to FinOps so run-rate cost does not surprise finance after go-live.
Learn
RFI, RFP and RFQ — when each applies
| Instrument | Purpose | Typical duration | AI relevance |
|---|---|---|---|
| RFI (Request for Information) | Market scan, capability discovery | 2–4 weeks | Model hosting options, residency, agent platforms |
| RFP (Request for Proposal) | Competitive selection with solution design | 6–12 weeks | End-to-end SI + platform + change |
| RFQ (Request for Quotation) | Price for defined spec | 2–6 weeks | Known SKU: API seats, vector DB capacity |
Rule: Do not skip RFI when the client has never procured GenAI at scale—otherwise RFP questions embed vendor marketing as requirements.
SOW, MSA, CR and order forms
MSA (Master Services Agreement) — umbrella legal frame: liability, indemnity, IP, confidentiality, termination, governing law.
SOW (Statement of Work) — specific delivery: scope, milestones, acceptance, fees, assumptions, dependencies.
CR (Change Request) — formal amendment when scope, data, model tier, or autonomy changes.
Order form / schedule — commercial SKU, term, renewal, price holds.
For AI, the SOW must reference technical schedules: data processing, security, model use restrictions, subprocessors, and acceptance annex.
Evaluation criteria — functional, technical, commercial
Procurement scorecards typically weight:
| Category | Example sub-criteria | AI-specific nuance |
|---|---|---|
| Functional fit | Use-case coverage, workflow integration | Human-in-loop design, not feature count |
| Technical fit | Architecture, security, scalability | RAG eval methodology, not demo latency alone |
| Security & compliance | SOC2, ISO 27001, DPA | Model training opt-out, residency, audit logs |
| Delivery | Methodology, team, references | Prompt/version control, MLOps maturity |
| Support | SLAs, escalation, training | Incident playbooks for harmful output |
| Commercial | TCO 3–5 years, exit cost | Token/run-rate modelled, not licence line only |
| Innovation / ESG | Roadmap, sustainability | Model routing for energy; inclusion |
Anti-pattern: 40% weight on "presentation quality" — rewards storytelling over operability.
Writing an AI RFP response — structure
- Executive summary — BLUF: fit, differentiators, risks acknowledged
- Understanding of requirements — mirror client language; flag ambiguities
- Solution overview — architecture diagram, data flows, human oversight
- Delivery approach — phases, gates, eval cadence, change control
- Team and governance — named roles, RACI, steering cadence
- Security and compliance — control mapping to client schedule
- Commercial model — capex/opex, consumption assumptions, FinOps linkage
- Acceptance and success criteria — measurable, testable, tiered
- Assumptions and dependencies — data readiness, API access, sponsor time
- References and case studies — numbers, not adjectives
Responses that negotiate unsafe requirements in writing ("we propose faithfulness ≥92% on held-out set X with citation requirement, not 100% accuracy") score higher with risk-aware evaluators.
Acceptance criteria for probabilistic AI
Absolute accuracy guarantees are scientifically and legally unsafe. Replace with:
| Measure | Definition | Example threshold |
|---|---|---|
| Faithfulness / groundedness | Answer supported by retrieved or cited source | ≥90% on agreed eval set |
| Citation coverage | Material claims linked to source span | 100% for regulated facts |
| Refusal appropriateness | Correct "I don't know" when evidence missing | ≥95% on abstention set |
| Harmful output rate | Policy violations per 10k queries | Below agreed ceiling |
| Latency p95 | End-to-end under load | ≤4s at 500 concurrent |
| Availability | API + orchestration layer | 99.5% monthly |
Acceptance should specify: eval dataset ownership, who runs UAT, retest trigger (model version change, corpus >20% delta), and human review queue for failures.
SLAs — what to measure and what to exclude
Standard IT SLAs (uptime, incident response) are necessary but insufficient for AI.
Include:
- API/orchestration availability and p95 latency
- Incident severity matrix (harmful output = Sev 1)
- Mean time to contain (disable feature, rollback model)
- Support hours and escalation path to engineering
Exclude or carve out:
- Third-party foundation model outages beyond failover design
- Quality degradation caused by client corpus pollution without change control
- Usage spikes beyond contracted rate limits without CR
Service credits should align with business impact—not token refunds alone when a copilot gave wrong policy advice.
Liability, indemnity and limitation of liability
Client concerns: Who pays when AI output causes financial loss, regulatory breach, or reputational harm?
Vendor concerns: Cap liability at fees paid; exclude consequential damages; indemnity only for IP infringement of vendor code.
AI Solution Engineer input:
- Map use case risk tier to liability discussion (informational vs decision support vs autonomous action)
- Document human-in-loop and disclaimers — they affect but do not eliminate duty of care
- Flag indemnity for training data — client data must not indemnify vendor for misuse in foundation training if prohibited
- Ensure subprocessor flow-down — model provider terms do not void client protections
Never advise on legal conclusions; surface trade-offs for legal and procurement.
IP — prompts, outputs, models and fine-tunes
| Asset | Typical client position | Typical vendor position | Negotiation note |
|---|---|---|---|
| Client documents / corpus | Client owns | Licence to deliver | Scope: processing only, no training |
| Prompts and system instructions | Client owns | Vendor retains templates | Project-specific prompts → client |
| Eval sets and rubrics | Client owns | Joint if vendor-built | Critical for regression |
| Fine-tuned weights | Client owns if paid | Vendor licence | Export format in contract |
| Generated outputs | Client owns | Varies | Clarify for regulated records |
| Pre-existing vendor IP | Licence | Vendor owns | Background IP schedule |
Fine-tune clause must address: export on termination, portability format, and prohibition on vendor reuse of client-specific weights in other accounts.
DPA, security schedules and subprocessors
AI almost always involves subprocessors: cloud host, model API, observability SaaS, annotation vendor.
DPA must list:
- Processing purpose and duration
- Data categories (prompts may contain PII)
- Residency and transfer mechanism (SCCs, UK IDTA)
- Subprocessor notification and objection rights
- Deletion on termination — including vector indexes and logs
- Audit and penetration test evidence
Security schedule should require: encryption in transit/at rest, RBAC, secrets management, logging without raw PII retention beyond policy, and prompt injection controls for agent tools.
Change control — model swap, corpus, autonomy
Any of these triggers a CR and often re-acceptance:
- Foundation model version change
- New data source or jurisdiction
- Autonomy increase (recommend → auto-send)
- New tool integrations with write access
- SLA or liability tier change
Embed change categories in SOW: Class A (no re-UAT), Class B (eval re-run), Class C (legal/risk re-approval).
Frameworks and methods
Three-layer contract stack for AI
Layer 1: MSA — legal frame (liability, IP, term)
Layer 2: SOW — delivery, acceptance, milestones
Layer 3: Technical schedules — DPA, security, SLA, subprocessor list, AI use policy
No production launch without Layer 3 complete.
MEAT acceptance framework
Measurable — numeric threshold or binary test
Evidence-based — named eval set and method
Achievable — baseline measured in discovery
Time-bound — test window and retest rules
Example: "MEAT: Faithfulness ≥88% on Claims-Eval-v3 (500 queries) within 10 business days of UAT start; retest if corpus adds >15% new documents."
Risk-tiered contracting
| Tier | Example use case | Contract emphasis |
|---|---|---|
| Low | Internal FAQ, no PII | Standard SLA, limited acceptance |
| Medium | Employee copilot with internal data | DPA, eval acceptance, audit logs |
| High | Customer-facing advice, credit, health | Human review, liability carve-outs, kill-switch, regulatory schedules |
Align tier with topic 19 governance and topic 18 privacy artefacts.
Commercial linkage to FinOps
Contract should reference:
- Consumption model — per token, per seat, per transaction
- Budget caps and alerts — client-side FinOps ownership
- Rate card for overages — pre-agreed, not list price surprise
- Model routing policy — small vs large model cost split documented
See AI FinOps and Commercial Design for TCO templates.
RFP red-line playbook — common unsafe clauses
| Clause | Problem | Counter-proposal |
|---|---|---|
| "100% accurate responses" | Unmeasurable | Faithfulness + citation thresholds |
| "Vendor warrants no hallucination" | Impossible | Harm rate ceiling + escalation |
| "Acceptance on go-live date" | Bypasses UAT | Milestone acceptance with eval annex |
| "Unlimited scope changes" | Margin and risk | CR process with eval re-trigger |
| "Client indemnifies all AI output" | One-sided | Tiered responsibility + HITL |
Procurement timeline — typical gates
| Week | Activity | AI Solution Engineer role |
|---|---|---|
| 1–2 | RFP issue / bidder conference | Clarify eval and data assumptions |
| 3–5 | Written questions | Submit technical Q&A on acceptance |
| 6–8 | Proposal submission | Architecture, acceptance annex, TCO |
| 9–10 | Orals / demos | Live eval, not scripted only |
| 11–12 | BAFO / negotiation | Technical schedule red lines |
| 13+ | Contracting | SOW test scripts, DPA subprocessor review |
Real-world scenarios
Scenario A — UK insurer: RFP for claims policy copilot
Context: Top-5 UK insurer; 2,800 claims handlers; RFP for internal RAG over policy manuals and precedents. Budget envelope £1.2m year one (build + run). Incumbent SI and two specialists compete. RFP draft requires "100% policy compliance in all answers."
Stakeholders: Procurement (weighted 30% cost), Claims COO (buyer), CISO (security schedule veto), Legal (liability), Model Risk (acceptance).
Intervention:
- Bidder conference: propose replacing absolute accuracy with faithfulness ≥91% on Claims-Eval-500 plus mandatory citation for coverage limits and exclusions
- Acceptance annex: 30-day UAT; failures routed to human review queue; no auto-deny of claims in phase 1
- Commercial: £680k build, £420k run-rate (private Azure OpenAI + search); FinOps dashboard in SOW deliverable
- Liability: cap at 12 months fees; carve-out for gross negligence; client retains decision on claim outcomes
Outcome: Preferred bidder selected; contract signed with MEAT acceptance. UAT achieves 93.2% faithfulness; go-live with 100% human approval on outputs touching payout calculation.
Numbers: Handle-time reduction pilot −14% (baseline 11.2 min → 9.6 min); projected annual benefit £3.1m vs £1.1m total cost — ROI clears hurdle 2.1x.
Scenario B — US regional bank: MSA negotiation with SaaS copilot vendor
Context: $850m assets regional bank; vendor MSA limits liability to $500k; bank's model risk policy requires $5m minimum for customer-facing tools. Vendor refuses uncapped; bank needs vendor for speed (16-week regulatory exam in 11 months).
Conflict: Legal deadlock; procurement wants signature before fiscal year end.
Intervention:
- Risk tier reclassification: phase 1 internal only (loan ops), not customer-facing — lowers required cap to $2m
- SOW adds insurance certificate — cyber $10m, tech E&O $5m
- IP: bank owns prompts, eval sets, and fine-tune exports; vendor background IP licensed
- SLA: 99.5% availability; Sev 1 harmful output response 4 hours; model version change requires 10-day notice and regression eval
- Termination: 90-day data export including vectors; $120k migration assistance fixed fee in SOW
Outcome: MSA signed with amended cap schedule by tier; customer-facing phase 2 requires new SOW and $5m cap renegotiation.
Numbers: Internal phase saves ~420 hours/month loan doc review; vendor run-rate $38k/month vs $52k build alternative — 27% lower year-one TCO with exit clause valued by CRO.
Scenario C — EU manufacturer: RFQ for document extraction agent
Context: German automotive supplier; 12,000 supplier invoices/month; RFQ for agent extracting PO matching fields. Procurement uses lowest price award rule. Bidder A: €180k fixed; Bidder B: €290k with eval guarantees.
Technical trap: Bidder A promises 98% field accuracy without defining test set or invoice types.
Intervention:
- Evaluation criteria reweight: technical 45%, cost 35%, delivery **20%
- Mandatory PoC on 500 representative invoices (scanned + EDI)
- Acceptance: ≥96% F1 on mandatory fields; human review for confidence <0.85
- DPA: data stays EU; no training on client data; deletion 30 days post-termination
Outcome: Bidder B wins at €290k; go-live F1 97.1%; Bidder A post-mortem would have failed UAT at 89% on handwritten invoices.
Numbers: Straight-through processing 34% of volume (target 30%); €1.8m/year labour redeployment value; payback 5.2 months.
Scenario D — Public sector framework call-off — CR for model change
Context: UK council; framework call-off for resident-services chatbot; £240k over 24 months. Vendor announces foundation model deprecation in 90 days; replacement model changes tone and citation behaviour.
Contract gap: MSA silent on model migration obligations.
Intervention:
- CR Class B: vendor runs regression on Resident-Eval-200; council joint UAT 15 days
- Cost: £0 migration fee (negotiated from MSA gap); council absorbs 40 hours internal test
- Updated SLA: vendor notifies model changes 60 days minimum; acceptance retest mandatory
Outcome: Migration completes with faithfulness drop 94% → 91% then prompt tuning recovery to 93.5%; no breach notice.
Lesson: Always embed model change notification and retest in technical schedule — framework terms are rarely AI-specific.
Indemnity and liability — scenario tables for negotiation prep
Scenario 1 — Hallucinated policy limit: Copilot cites £50k limit; actual £25k; customer overcommits.
| Party | Typical ask | Engineering input |
|---|---|---|
| Client legal | Vendor indemnifies all output loss | Impossible; propose HITL on limits |
| Vendor legal | Cap at 12-month fees | Accept if tier-medium internal |
| Risk | Human approval on financial figures | Product requirement |
Scenario 2 — Training data leak: Client confidential docs appear in another tenant's responses (vendor fault).
| Party | Ask | Contract clause |
|---|---|---|
| Client | Uncapped indemnity | Insurance + vendor breach clause |
| Vendor | Security incident process | 72h notification |
Use tables like these in negotiation prep—legal owns outcome; you supply operability facts.
Security schedule cross-reference — AI control mapping
Map contract security schedule to architecture controls:
| Schedule line | Control | Test in UAT |
|---|---|---|
| Encryption at rest | AES-256 on index | Config audit |
| RBAC | Role matrix | Pen test sample |
| Logging | No raw PII >30d | Log review |
| Subprocessor approval | 30-day notice | Process drill |
| Prompt injection | Tool authZ deny | Red team case |
Incomplete mapping → acceptance risk at security gate.
CR worked example — autonomy increase
Trigger: Business requests auto-send email responses without human approval.
Class: C — requires legal, risk, re-UAT.
SOW CR text:
- Scope: Enable auto-send for tier-1 FAQ intents only; intent classifier version IC-2.1
- Acceptance: Harmful rate <0.02% on AutoSend-Eval-300; human queue for confidence <0.9
- Fee: £45k; timeline +6 weeks
- Liability: Cap increase to £2m aggregate for phase
Without CR, vendor and delivery team operate out of contract.
Procurement interview questions — for AI Solution Engineer role on deal team
When joining bid team, ask procurement:
- What is lowest price vs best value award rule?
- Can we challenge requirements in writing without disqualification?
- Who signs technical acceptance—procurement or business?
- Is incumbent getting unequal extension time?
- Where does FinOps sit in evaluation weight?
Answers shape response strategy and red-line priority.
Practice exercises
Primary exercise — Draft acceptance annex for RAG pilot (90 minutes)
Brief: Retail bank internal HR policy copilot; 4,500 employees; corpus HR-Policy-v4 (~820 documents); private deployment; pilot 90 days.
Tasks:
- Write five MEAT acceptance criteria (faithfulness, citation, refusal, latency, availability).
- Define UAT process: who runs eval, duration, pass/fail, remediation loop.
- List three explicit exclusions (what the system is not required to do).
- Draft one CR trigger table (Class A/B/C) for corpus update, model change, new country policy.
- Write two RFP red-line responses to unsafe client requirements.
Acceptance criteria:
- No absolute accuracy language
- Thresholds plausible for HR domain (cite industry eval ranges 85–95%)
- Human escalation for compensation and disciplinary topics named
- ≤600 words for acceptance annex body
Stretch exercise — Full SOW outline with commercial linkage (half day)
Brief: Healthcare provider member-services assistant; 1.1m members; phase 1 FAQ (no PHI in prompts); phase 2 authenticated claims status; $2.4m budget cap over 18 months.
Tasks:
- SOW structure: scope, milestones, deliverables, assumptions, dependencies.
- IP schedule: prompts, eval sets, fine-tunes, outputs.
- SLA and severity matrix including harmful medical advice scenario.
- Liability and risk-tier table for phase 1 vs 2.
- FinOps appendix: consumption assumptions, monthly cap, alert thresholds — cross-reference FinOps guide.
- Subprocessor table with model API, cloud, observability vendors.
- Exit plan: data export, 90-day transition assistance, cost estimate.
Acceptance criteria:
- Phase 2 gated on DPIA and model risk sign-off referenced in SOW
- MEAT acceptance for phase 1 only
- TCO table 3 years with low/base/high token scenarios
- Named roles: client product owner, vendor tech lead, legal approver
Steering pack insert — procurement status (template)
For client steercos during long RFP/contract cycles:
BLUF: Contract on track for DD/MM signature / at risk because acceptance language.
Decisions needed: Approve MEAT thresholds; approve liability tier; nominate UAT owner.
RAG: Commercial G/A/R with reason.
Next gate: Legal red-line return DD/MM.
Reduces surprise when procurement timeline slips build.
FAQ — RFP and contracting for AI engineers
Q: Can we ever accept "100% accuracy" in contract?
A: No for probabilistic systems—propose MEAT faithfulness + human escalation; document in clarification or red-line.
Q: Who signs UAT—procurement or business?
A: Usually business process owner with technical evidence; clarify in SOW before build completes.
Q: Does fine-tune IP default to vendor?
A: Never assume—negotiate export and ownership explicitly in IP schedule.
Q: What if model provider changes terms mid-contract?
A: CR or MSA change-in-law clause; trigger eval retest and exit option review.
Q: How link SLA credits to harmful output?
A: Technical schedule—Sev1 harmful output triggers service credit and root cause RCA, not only uptime credits.
Questions you should be able to answer
- What is the difference between RFI, RFP and RFQ in your current procurement, and which fits a novel AI use case?
- What belongs in the MSA vs the SOW vs technical schedules?
- How do you write acceptance criteria for a probabilistic system without promising 100% accuracy?
- Who owns IP in prompts, eval datasets, fine-tuned weights, and generated outputs?
- What subprocessors must appear in the DPA for a typical RAG deployment?
- What SLA metrics matter beyond API uptime for AI services?
- How do service credits align with harmful output incidents vs latency breaches?
- What liability cap is proportionate to the use-case risk tier?
- What triggers a change request and re-acceptance for model or corpus changes?
- How do you red-line "vendor warrants no hallucination" constructively?
- What FinOps assumptions should be attached to consumption-based pricing?
- How do you test acceptance — who owns the eval set and UAT window?
- What happens on termination — vector index, logs, fine-tune export?
- How do evaluation criteria weights avoid rewarding demo theatre?
- What explicit exclusions prevent scope creep into autonomous decisions?
RFP response writing — section-by-section craft
Executive summary (1 page max): Lead with fit and risk honesty. Executives read this only—if you bury acceptance philosophy here, procurement may never see it. Structure: (1) we understand your outcome; (2) our differentiated approach; (3) risks we will not pretend away; (4) why us now.
Requirements traceability matrix: Attach as appendix—each RFP line item mapped to response section, approach, and acceptance test. Evaluators score completeness; gaps become clarification questions or lost points. For AI requirements like "accurate answers," map to faithfulness metric + eval set ID—never leave orphan requirements.
Pricing narrative: Separate build, run-rate, change budget, and optional scale. Tie consumption to FinOps scenarios from AI FinOps and Commercial Design. Example table:
| Year | Build | Run (base) | Run (high) | Total base |
|---|---|---|---|---|
| 1 | £680k | £420k | £580k | £1.10m |
| 2 | £120k | £440k | £680k | £560k |
| 3 | £80k | £460k | £720k | £540k |
Procurement compares TCO, not year-one licence alone.
Negotiation dynamics — procurement, legal, and technical alignment
Procurement optimises competition, fairness, and audit trail. Legal optimises risk allocation and enforceability. Technical optimises operability. When these three talk past each other, contracts fail in production.
Joint working session (recommended before red-line exchange):
- Technical walks through acceptance tests—what pass/fail looks like
- Legal maps liability to risk tier and human oversight design
- Procurement confirms evaluation criteria still match negotiated acceptance
- FinOps validates consumption caps in order form
Concession trading: Never give liability cap without getting model change notification, export rights, or eval re-test clause. Document trades in deal memo.
Warranty, insurance and indemnity schedules
Beyond limitation of liability, enterprise AI deals may require:
| Instrument | Purpose | AI nuance |
|---|---|---|
| Tech E&O insurance | Professional negligence | Verify AI/automation covered |
| Cyber policy | Breach, ransomware | Prompt injection exfiltration |
| IP indemnity | Third-party IP claims | Training data provenance |
| Performance warranty | Short-period defect fix | Define "defect" as eval fail not subjective dislike |
Warranty period: 90 days post-acceptance common; exclude defects from client corpus changes without CR.
Insurance certificates — verification checklist
- Named insured matches contracting entity
- Policy period covers implementation + 12 months run
- Limits meet client minimums by tier
- AI/automation language not excluded in schedule
- Certificate received before production go-live if contract requires
Worked example — MEAT acceptance annex excerpt
Use case: Internal legal contract review assistant (tier-medium).
Criterion 1 — Faithfulness: On Legal-Eval-400 (held-out clause questions), groundedness score ≥ 89% using RAGAS or client-agreed equivalent, measured over 10 consecutive business days of UAT.
Criterion 2 — Citation: For answers addressing liability caps, indemnity, or termination, 100% include link to source clause ID in corpus metadata.
Criterion 3 — Refusal: On Legal-Abstain-75 (questions outside corpus), appropriate refusal or escalation rate ≥ 94%.
Criterion 4 — Latency: p95 end-to-end ≤ 5.0s at 200 simulated concurrent sessions.
Criterion 5 — Availability: Orchestration layer ≥ 99.5% during UAT window excluding agreed maintenance.
Exclusions: System shall not provide binding legal advice; shall not auto-redline without human publish; shall not process privilege-marked docs until Class C CR complete.
Retest triggers: Foundation model change; corpus delta > 20% by token count; new jurisdiction pack.
Bidder conference and clarification questions — AI-specific
Strong clarification questions (submitted in writing):
- "Requirement 4.2.1 specifies 100% accuracy—may we propose faithfulness on Eval-XX with human review for failures?"
- "Who owns UAT eval set creation and maintenance post-launch?"
- "Are prompts and system instructions client IP upon payment?"
- "What is required subprocessor notification lead time for new model providers?"
- "Is phase 2 autonomy (auto-send) in scope for this RFP or separate gate?"
Questions educate procurement on record—reduces unsafe language in final RFP for future cycles.
Contract lifecycle — post-signature obligations
| Phase | Contract mechanism | AI action |
|---|---|---|
| Mobilisation | SOW kick-off | Baseline eval; config management |
| Build | Milestone acceptance | Partial MEAT where applicable |
| UAT | Acceptance annex | Full eval; issue log |
| Hypercare | SLA begins | Quality SLO monitoring |
| Steady state | CR process | Model/corpus change control |
| Renewal | Price hold / re-bid | TCO actuals vs forecast |
| Termination | Exit schedule | Export, delete, transition |
Cross-border and framework agreements
Framework call-offs (public sector, GPO): master terms may lack AI schedules—attach call-off-specific technical schedule every time. Never assume framework covers eval acceptance.
Multi-jurisdiction: Data residency clause must list each country corpus; model routing per region; subprocessors per region in DPA appendix.
FinOps commercial linkage — contract clauses to request
From AI FinOps and Commercial Design:
- Monthly usage report with token breakdown by model and feature
- Budget alert integration to client FinOps tooling (webhook/API)
- Rate card for overages pre-agreed—not list price at invoice
- Right to route models for cost optimisation without vendor penalty if eval maintained
- Annual true-up cap at ±15% of forecast unless CR
Monthly procurement–delivery rhythm (during implementation)
| Week | Activity |
|---|---|
| 1 | Milestone status vs SOW; CR log review |
| 2 | Eval results vs acceptance thresholds |
| 3 | FinOps actuals vs order form forecast |
| 4 | Risk/legal touch if model or data change pending |
Prevents acceptance surprise at UAT boundary.
Oral presentation and demo — contractual implications
Orals often create implied promises. Discipline:
- Demo labels: "PoC environment; metrics from Eval-v2 dated DD/MM"
- No ad hoc accuracy claims; slide footer with measurement method
- Handout matches written proposal—verbal-only commitments become CR or lost
Record orals where client policy allows; send summary email within 24h confirming what was not promised.
Glossary — procurement terms for AI engineers
| Term | Meaning |
|---|---|
| BAFO | Best and Final Offer — last negotiation round |
| Indemnity | Compensation for specified losses |
| Limitation of liability | Cap on damages recoverable |
| Order form | Commercial SKU attachment to MSA |
| Schedule | Technical/legal annex (security, DPA, SLA) |
| Subprocessor | Vendor's downstream processor (model API, cloud) |
| UAT | User acceptance testing against contract criteria |
| Warranty | Defect repair obligation for defined period |
Fluency in terms speeds joint sessions with legal—reduces translation errors.
Final readiness — contract signature checklist
- MEAT acceptance annex signed or incorporated
- IP schedule complete
- DPA subprocessors current
- SLA severity includes harmful output
- CR classes defined
- FinOps reporting in SOW
- Exit/export clause tested in PoC
- Insurance certs verified
- Technical evaluator named for UAT sign-off
Negative cases — when contracting fails
The 100% accuracy RFP
Symptom: RFP embeds vendor marketing language; bidders fail to challenge; UAT impossible.
Impact: Perpetual dispute; blame between SI and client; programme pause.
Fix: MEAT acceptance; educate procurement before RFP issue; bidder conference Q&A on record.
Acceptance on calendar date
Symptom: "Go-live 1 June" with no eval gate; pressure to sign acceptance regardless of quality.
Impact: Production incidents; regulatory exposure; warranty claims.
Fix: Milestone acceptance tied to eval results; steering holds date until pass or documented risk acceptance.
IP vacuum
Symptom: SOW silent on prompts and fine-tunes; vendor claims background IP on all deliverables.
Impact: Cannot migrate; re-procurement from scratch; loss of eval assets.
Fix: IP schedule in every AI SOW; export format specified pre-signature.
SLA without quality
Symptom: 99.9% uptime SLA; faithfulness drops 40% after corpus update; no breach.
Impact: Users lose trust; support tickets rise; no contractual lever.
Fix: Quality SLOs in technical schedule; regression eval on corpus change; optional service credits for eval failure.
Liability mismatch
Symptom: $500k cap on customer-facing credit decision tool.
Impact: Uninsurable residual risk; board refuses launch; emergency renegotiation under time pressure.
Fix: Risk-tiered caps; insurance certificates; phase internal before external.
Subprocessor surprise
Symptom: Model provider not listed; data processed in non-contracted region.
Impact: DPA breach; regulator notification; contract termination rights invoked.
Fix: Subprocessor notification workflow; residency clause; audit right.
Practice checklist
- I can explain RFI/RFP/RFQ selection for a novel AI use case
- I drafted MEAT acceptance without absolute accuracy promises
- I mapped IP ownership for prompts, evals and fine-tunes
- I linked SLAs to quality SLOs, not only uptime
- I completed the primary acceptance annex exercise
- I identified three negative cases relevant to my context
- I cross-referenced FinOps assumptions with commercial terms
Related playbook content
- AI FinOps and Commercial Design — TCO, consumption models and chargeback
- Commercial and Financial Modelling — Business case and pricing logic
- Performance Engineering and AI FinOps — Run-rate and model routing cost
- Privacy, Legal and Compliance — DPA and lawful basis inputs
- Responsible AI and Governance — Risk tier and approval gates
- Vendor and Technology Evaluation — Scorecards and PoC before contract
- Presales and Solution Shaping — Proposal narrative alignment
- Delivery and Programme Management — Milestones and change control
- How to use this Learning Map — Reference-depth study method
- 8D Framework — Lifecycle gates
Discussion
Comments
Share feedback or questions about this page. No account required.
Loading comments…