Skip to main content

RFP, Procurement and Contracting

Executive view

Focus on whether proposals commit to measurable outcomes (faithfulness, latency, availability, human escalation)—not marketing superlatives. Ask: who signs acceptance, what happens when the model provider changes terms, and whether liability matches the risk tier.

Decision required: Are evaluation criteria weighted toward client outcomes and exit portability—not vendor demo polish?

Technical view

Translate architecture into contract language: data flows, subprocessors, IP in prompts and fine-tunes, rollback obligations, observability access, and SLAs tied to SLOs you can actually monitor.

Never sign technical schedules you cannot test. Draft acceptance tests alongside the SOW, not after build.

Why this matters

AI procurement fails in predictable ways. Vendors promise "enterprise-grade accuracy." Business sponsors paste those phrases into RFPs. Delivery teams inherit commitments that no eval suite can prove. Legal teams insert liability caps that leave the client exposed when a model hallucinates in a regulated workflow. Procurement optimises for licence cost while ignoring retraining, egress, and exit migration.

An AI Solution Engineer sits between these functions. You are not a lawyer, but you must translate probabilistic systems into contract-grade language—thresholds, test methods, escalation paths, and explicit exclusions. Weak contracting produces:

  • Acceptance deadlock — UAT fails but contract says "substantially complete"
  • Scope creep without CR — "just add another corpus" with no change control
  • IP disputes — who owns prompts, eval sets, fine-tunes, and synthetic training data
  • SLA theatre — 99.9% uptime on the API while retrieval quality collapses
  • Vendor lock-in — no exit clause when the model provider deprecates a deployment

Strong contracting produces testable acceptance, aligned incentives, clear IP and data rights, and commercial terms linked to FinOps so run-rate cost does not surprise finance after go-live.

Learn

RFI, RFP and RFQ — when each applies

InstrumentPurposeTypical durationAI relevance
RFI (Request for Information)Market scan, capability discovery2–4 weeksModel hosting options, residency, agent platforms
RFP (Request for Proposal)Competitive selection with solution design6–12 weeksEnd-to-end SI + platform + change
RFQ (Request for Quotation)Price for defined spec2–6 weeksKnown SKU: API seats, vector DB capacity

Rule: Do not skip RFI when the client has never procured GenAI at scale—otherwise RFP questions embed vendor marketing as requirements.

SOW, MSA, CR and order forms

MSA (Master Services Agreement) — umbrella legal frame: liability, indemnity, IP, confidentiality, termination, governing law.

SOW (Statement of Work) — specific delivery: scope, milestones, acceptance, fees, assumptions, dependencies.

CR (Change Request) — formal amendment when scope, data, model tier, or autonomy changes.

Order form / schedule — commercial SKU, term, renewal, price holds.

For AI, the SOW must reference technical schedules: data processing, security, model use restrictions, subprocessors, and acceptance annex.

Evaluation criteria — functional, technical, commercial

Procurement scorecards typically weight:

CategoryExample sub-criteriaAI-specific nuance
Functional fitUse-case coverage, workflow integrationHuman-in-loop design, not feature count
Technical fitArchitecture, security, scalabilityRAG eval methodology, not demo latency alone
Security & complianceSOC2, ISO 27001, DPAModel training opt-out, residency, audit logs
DeliveryMethodology, team, referencesPrompt/version control, MLOps maturity
SupportSLAs, escalation, trainingIncident playbooks for harmful output
CommercialTCO 3–5 years, exit costToken/run-rate modelled, not licence line only
Innovation / ESGRoadmap, sustainabilityModel routing for energy; inclusion

Anti-pattern: 40% weight on "presentation quality" — rewards storytelling over operability.

Writing an AI RFP response — structure

  1. Executive summary — BLUF: fit, differentiators, risks acknowledged
  2. Understanding of requirements — mirror client language; flag ambiguities
  3. Solution overview — architecture diagram, data flows, human oversight
  4. Delivery approach — phases, gates, eval cadence, change control
  5. Team and governance — named roles, RACI, steering cadence
  6. Security and compliance — control mapping to client schedule
  7. Commercial model — capex/opex, consumption assumptions, FinOps linkage
  8. Acceptance and success criteria — measurable, testable, tiered
  9. Assumptions and dependencies — data readiness, API access, sponsor time
  10. References and case studies — numbers, not adjectives

Responses that negotiate unsafe requirements in writing ("we propose faithfulness ≥92% on held-out set X with citation requirement, not 100% accuracy") score higher with risk-aware evaluators.

Acceptance criteria for probabilistic AI

Absolute accuracy guarantees are scientifically and legally unsafe. Replace with:

MeasureDefinitionExample threshold
Faithfulness / groundednessAnswer supported by retrieved or cited source≥90% on agreed eval set
Citation coverageMaterial claims linked to source span100% for regulated facts
Refusal appropriatenessCorrect "I don't know" when evidence missing≥95% on abstention set
Harmful output ratePolicy violations per 10k queriesBelow agreed ceiling
Latency p95End-to-end under load≤4s at 500 concurrent
AvailabilityAPI + orchestration layer99.5% monthly

Acceptance should specify: eval dataset ownership, who runs UAT, retest trigger (model version change, corpus >20% delta), and human review queue for failures.

SLAs — what to measure and what to exclude

Standard IT SLAs (uptime, incident response) are necessary but insufficient for AI.

Include:

  • API/orchestration availability and p95 latency
  • Incident severity matrix (harmful output = Sev 1)
  • Mean time to contain (disable feature, rollback model)
  • Support hours and escalation path to engineering

Exclude or carve out:

  • Third-party foundation model outages beyond failover design
  • Quality degradation caused by client corpus pollution without change control
  • Usage spikes beyond contracted rate limits without CR

Service credits should align with business impact—not token refunds alone when a copilot gave wrong policy advice.

Liability, indemnity and limitation of liability

Client concerns: Who pays when AI output causes financial loss, regulatory breach, or reputational harm?

Vendor concerns: Cap liability at fees paid; exclude consequential damages; indemnity only for IP infringement of vendor code.

AI Solution Engineer input:

  • Map use case risk tier to liability discussion (informational vs decision support vs autonomous action)
  • Document human-in-loop and disclaimers — they affect but do not eliminate duty of care
  • Flag indemnity for training data — client data must not indemnify vendor for misuse in foundation training if prohibited
  • Ensure subprocessor flow-down — model provider terms do not void client protections

Never advise on legal conclusions; surface trade-offs for legal and procurement.

IP — prompts, outputs, models and fine-tunes

AssetTypical client positionTypical vendor positionNegotiation note
Client documents / corpusClient ownsLicence to deliverScope: processing only, no training
Prompts and system instructionsClient ownsVendor retains templatesProject-specific prompts → client
Eval sets and rubricsClient ownsJoint if vendor-builtCritical for regression
Fine-tuned weightsClient owns if paidVendor licenceExport format in contract
Generated outputsClient ownsVariesClarify for regulated records
Pre-existing vendor IPLicenceVendor ownsBackground IP schedule

Fine-tune clause must address: export on termination, portability format, and prohibition on vendor reuse of client-specific weights in other accounts.

DPA, security schedules and subprocessors

AI almost always involves subprocessors: cloud host, model API, observability SaaS, annotation vendor.

DPA must list:

  • Processing purpose and duration
  • Data categories (prompts may contain PII)
  • Residency and transfer mechanism (SCCs, UK IDTA)
  • Subprocessor notification and objection rights
  • Deletion on termination — including vector indexes and logs
  • Audit and penetration test evidence

Security schedule should require: encryption in transit/at rest, RBAC, secrets management, logging without raw PII retention beyond policy, and prompt injection controls for agent tools.

Change control — model swap, corpus, autonomy

Any of these triggers a CR and often re-acceptance:

  • Foundation model version change
  • New data source or jurisdiction
  • Autonomy increase (recommend → auto-send)
  • New tool integrations with write access
  • SLA or liability tier change

Embed change categories in SOW: Class A (no re-UAT), Class B (eval re-run), Class C (legal/risk re-approval).

Frameworks and methods

Three-layer contract stack for AI

Layer 1: MSA — legal frame (liability, IP, term)
Layer 2: SOW — delivery, acceptance, milestones
Layer 3: Technical schedules — DPA, security, SLA, subprocessor list, AI use policy

No production launch without Layer 3 complete.

MEAT acceptance framework

Measurable — numeric threshold or binary test
Evidence-based — named eval set and method
Achievable — baseline measured in discovery
Time-bound — test window and retest rules

Example: "MEAT: Faithfulness ≥88% on Claims-Eval-v3 (500 queries) within 10 business days of UAT start; retest if corpus adds >15% new documents."

Risk-tiered contracting

TierExample use caseContract emphasis
LowInternal FAQ, no PIIStandard SLA, limited acceptance
MediumEmployee copilot with internal dataDPA, eval acceptance, audit logs
HighCustomer-facing advice, credit, healthHuman review, liability carve-outs, kill-switch, regulatory schedules

Align tier with topic 19 governance and topic 18 privacy artefacts.

Commercial linkage to FinOps

Contract should reference:

  • Consumption model — per token, per seat, per transaction
  • Budget caps and alerts — client-side FinOps ownership
  • Rate card for overages — pre-agreed, not list price surprise
  • Model routing policy — small vs large model cost split documented

See AI FinOps and Commercial Design for TCO templates.

RFP red-line playbook — common unsafe clauses

ClauseProblemCounter-proposal
"100% accurate responses"UnmeasurableFaithfulness + citation thresholds
"Vendor warrants no hallucination"ImpossibleHarm rate ceiling + escalation
"Acceptance on go-live date"Bypasses UATMilestone acceptance with eval annex
"Unlimited scope changes"Margin and riskCR process with eval re-trigger
"Client indemnifies all AI output"One-sidedTiered responsibility + HITL

Procurement timeline — typical gates

WeekActivityAI Solution Engineer role
1–2RFP issue / bidder conferenceClarify eval and data assumptions
3–5Written questionsSubmit technical Q&A on acceptance
6–8Proposal submissionArchitecture, acceptance annex, TCO
9–10Orals / demosLive eval, not scripted only
11–12BAFO / negotiationTechnical schedule red lines
13+ContractingSOW test scripts, DPA subprocessor review

Real-world scenarios

Scenario A — UK insurer: RFP for claims policy copilot

Context: Top-5 UK insurer; 2,800 claims handlers; RFP for internal RAG over policy manuals and precedents. Budget envelope £1.2m year one (build + run). Incumbent SI and two specialists compete. RFP draft requires "100% policy compliance in all answers."

Stakeholders: Procurement (weighted 30% cost), Claims COO (buyer), CISO (security schedule veto), Legal (liability), Model Risk (acceptance).

Intervention:

  • Bidder conference: propose replacing absolute accuracy with faithfulness ≥91% on Claims-Eval-500 plus mandatory citation for coverage limits and exclusions
  • Acceptance annex: 30-day UAT; failures routed to human review queue; no auto-deny of claims in phase 1
  • Commercial: £680k build, £420k run-rate (private Azure OpenAI + search); FinOps dashboard in SOW deliverable
  • Liability: cap at 12 months fees; carve-out for gross negligence; client retains decision on claim outcomes

Outcome: Preferred bidder selected; contract signed with MEAT acceptance. UAT achieves 93.2% faithfulness; go-live with 100% human approval on outputs touching payout calculation.

Numbers: Handle-time reduction pilot −14% (baseline 11.2 min → 9.6 min); projected annual benefit £3.1m vs £1.1m total cost — ROI clears hurdle 2.1x.

Scenario B — US regional bank: MSA negotiation with SaaS copilot vendor

Context: $850m assets regional bank; vendor MSA limits liability to $500k; bank's model risk policy requires $5m minimum for customer-facing tools. Vendor refuses uncapped; bank needs vendor for speed (16-week regulatory exam in 11 months).

Conflict: Legal deadlock; procurement wants signature before fiscal year end.

Intervention:

  • Risk tier reclassification: phase 1 internal only (loan ops), not customer-facing — lowers required cap to $2m
  • SOW adds insurance certificate — cyber $10m, tech E&O $5m
  • IP: bank owns prompts, eval sets, and fine-tune exports; vendor background IP licensed
  • SLA: 99.5% availability; Sev 1 harmful output response 4 hours; model version change requires 10-day notice and regression eval
  • Termination: 90-day data export including vectors; $120k migration assistance fixed fee in SOW

Outcome: MSA signed with amended cap schedule by tier; customer-facing phase 2 requires new SOW and $5m cap renegotiation.

Numbers: Internal phase saves ~420 hours/month loan doc review; vendor run-rate $38k/month vs $52k build alternative — 27% lower year-one TCO with exit clause valued by CRO.

Scenario C — EU manufacturer: RFQ for document extraction agent

Context: German automotive supplier; 12,000 supplier invoices/month; RFQ for agent extracting PO matching fields. Procurement uses lowest price award rule. Bidder A: €180k fixed; Bidder B: €290k with eval guarantees.

Technical trap: Bidder A promises 98% field accuracy without defining test set or invoice types.

Intervention:

  • Evaluation criteria reweight: technical 45%, cost 35%, delivery **20%
  • Mandatory PoC on 500 representative invoices (scanned + EDI)
  • Acceptance: ≥96% F1 on mandatory fields; human review for confidence <0.85
  • DPA: data stays EU; no training on client data; deletion 30 days post-termination

Outcome: Bidder B wins at €290k; go-live F1 97.1%; Bidder A post-mortem would have failed UAT at 89% on handwritten invoices.

Numbers: Straight-through processing 34% of volume (target 30%); €1.8m/year labour redeployment value; payback 5.2 months.

Scenario D — Public sector framework call-off — CR for model change

Context: UK council; framework call-off for resident-services chatbot; £240k over 24 months. Vendor announces foundation model deprecation in 90 days; replacement model changes tone and citation behaviour.

Contract gap: MSA silent on model migration obligations.

Intervention:

  • CR Class B: vendor runs regression on Resident-Eval-200; council joint UAT 15 days
  • Cost: £0 migration fee (negotiated from MSA gap); council absorbs 40 hours internal test
  • Updated SLA: vendor notifies model changes 60 days minimum; acceptance retest mandatory

Outcome: Migration completes with faithfulness drop 94% → 91% then prompt tuning recovery to 93.5%; no breach notice.

Lesson: Always embed model change notification and retest in technical schedule — framework terms are rarely AI-specific.

Indemnity and liability — scenario tables for negotiation prep

Scenario 1 — Hallucinated policy limit: Copilot cites £50k limit; actual £25k; customer overcommits.

PartyTypical askEngineering input
Client legalVendor indemnifies all output lossImpossible; propose HITL on limits
Vendor legalCap at 12-month feesAccept if tier-medium internal
RiskHuman approval on financial figuresProduct requirement

Scenario 2 — Training data leak: Client confidential docs appear in another tenant's responses (vendor fault).

PartyAskContract clause
ClientUncapped indemnityInsurance + vendor breach clause
VendorSecurity incident process72h notification

Use tables like these in negotiation prep—legal owns outcome; you supply operability facts.

Security schedule cross-reference — AI control mapping

Map contract security schedule to architecture controls:

Schedule lineControlTest in UAT
Encryption at restAES-256 on indexConfig audit
RBACRole matrixPen test sample
LoggingNo raw PII >30dLog review
Subprocessor approval30-day noticeProcess drill
Prompt injectionTool authZ denyRed team case

Incomplete mapping → acceptance risk at security gate.

CR worked example — autonomy increase

Trigger: Business requests auto-send email responses without human approval.

Class: C — requires legal, risk, re-UAT.

SOW CR text:

  • Scope: Enable auto-send for tier-1 FAQ intents only; intent classifier version IC-2.1
  • Acceptance: Harmful rate <0.02% on AutoSend-Eval-300; human queue for confidence <0.9
  • Fee: £45k; timeline +6 weeks
  • Liability: Cap increase to £2m aggregate for phase

Without CR, vendor and delivery team operate out of contract.

Procurement interview questions — for AI Solution Engineer role on deal team

When joining bid team, ask procurement:

  1. What is lowest price vs best value award rule?
  2. Can we challenge requirements in writing without disqualification?
  3. Who signs technical acceptance—procurement or business?
  4. Is incumbent getting unequal extension time?
  5. Where does FinOps sit in evaluation weight?

Answers shape response strategy and red-line priority.

Practice exercises

Primary exercise — Draft acceptance annex for RAG pilot (90 minutes)

Brief: Retail bank internal HR policy copilot; 4,500 employees; corpus HR-Policy-v4 (~820 documents); private deployment; pilot 90 days.

Tasks:

  1. Write five MEAT acceptance criteria (faithfulness, citation, refusal, latency, availability).
  2. Define UAT process: who runs eval, duration, pass/fail, remediation loop.
  3. List three explicit exclusions (what the system is not required to do).
  4. Draft one CR trigger table (Class A/B/C) for corpus update, model change, new country policy.
  5. Write two RFP red-line responses to unsafe client requirements.

Acceptance criteria:

  • No absolute accuracy language
  • Thresholds plausible for HR domain (cite industry eval ranges 85–95%)
  • Human escalation for compensation and disciplinary topics named
  • ≤600 words for acceptance annex body

Stretch exercise — Full SOW outline with commercial linkage (half day)

Brief: Healthcare provider member-services assistant; 1.1m members; phase 1 FAQ (no PHI in prompts); phase 2 authenticated claims status; $2.4m budget cap over 18 months.

Tasks:

  1. SOW structure: scope, milestones, deliverables, assumptions, dependencies.
  2. IP schedule: prompts, eval sets, fine-tunes, outputs.
  3. SLA and severity matrix including harmful medical advice scenario.
  4. Liability and risk-tier table for phase 1 vs 2.
  5. FinOps appendix: consumption assumptions, monthly cap, alert thresholds — cross-reference FinOps guide.
  6. Subprocessor table with model API, cloud, observability vendors.
  7. Exit plan: data export, 90-day transition assistance, cost estimate.

Acceptance criteria:

  • Phase 2 gated on DPIA and model risk sign-off referenced in SOW
  • MEAT acceptance for phase 1 only
  • TCO table 3 years with low/base/high token scenarios
  • Named roles: client product owner, vendor tech lead, legal approver

Steering pack insert — procurement status (template)

For client steercos during long RFP/contract cycles:

BLUF: Contract on track for DD/MM signature / at risk because acceptance language.

Decisions needed: Approve MEAT thresholds; approve liability tier; nominate UAT owner.

RAG: Commercial G/A/R with reason.

Next gate: Legal red-line return DD/MM.

Reduces surprise when procurement timeline slips build.

FAQ — RFP and contracting for AI engineers

Q: Can we ever accept "100% accuracy" in contract?
A: No for probabilistic systems—propose MEAT faithfulness + human escalation; document in clarification or red-line.

Q: Who signs UAT—procurement or business?
A: Usually business process owner with technical evidence; clarify in SOW before build completes.

Q: Does fine-tune IP default to vendor?
A: Never assume—negotiate export and ownership explicitly in IP schedule.

Q: What if model provider changes terms mid-contract?
A: CR or MSA change-in-law clause; trigger eval retest and exit option review.

Q: How link SLA credits to harmful output?
A: Technical schedule—Sev1 harmful output triggers service credit and root cause RCA, not only uptime credits.

Questions you should be able to answer

  1. What is the difference between RFI, RFP and RFQ in your current procurement, and which fits a novel AI use case?
  2. What belongs in the MSA vs the SOW vs technical schedules?
  3. How do you write acceptance criteria for a probabilistic system without promising 100% accuracy?
  4. Who owns IP in prompts, eval datasets, fine-tuned weights, and generated outputs?
  5. What subprocessors must appear in the DPA for a typical RAG deployment?
  6. What SLA metrics matter beyond API uptime for AI services?
  7. How do service credits align with harmful output incidents vs latency breaches?
  8. What liability cap is proportionate to the use-case risk tier?
  9. What triggers a change request and re-acceptance for model or corpus changes?
  10. How do you red-line "vendor warrants no hallucination" constructively?
  11. What FinOps assumptions should be attached to consumption-based pricing?
  12. How do you test acceptance — who owns the eval set and UAT window?
  13. What happens on termination — vector index, logs, fine-tune export?
  14. How do evaluation criteria weights avoid rewarding demo theatre?
  15. What explicit exclusions prevent scope creep into autonomous decisions?

RFP response writing — section-by-section craft

Executive summary (1 page max): Lead with fit and risk honesty. Executives read this only—if you bury acceptance philosophy here, procurement may never see it. Structure: (1) we understand your outcome; (2) our differentiated approach; (3) risks we will not pretend away; (4) why us now.

Requirements traceability matrix: Attach as appendix—each RFP line item mapped to response section, approach, and acceptance test. Evaluators score completeness; gaps become clarification questions or lost points. For AI requirements like "accurate answers," map to faithfulness metric + eval set ID—never leave orphan requirements.

Pricing narrative: Separate build, run-rate, change budget, and optional scale. Tie consumption to FinOps scenarios from AI FinOps and Commercial Design. Example table:

YearBuildRun (base)Run (high)Total base
1£680k£420k£580k£1.10m
2£120k£440k£680k£560k
3£80k£460k£720k£540k

Procurement compares TCO, not year-one licence alone.

Procurement optimises competition, fairness, and audit trail. Legal optimises risk allocation and enforceability. Technical optimises operability. When these three talk past each other, contracts fail in production.

Joint working session (recommended before red-line exchange):

  1. Technical walks through acceptance tests—what pass/fail looks like
  2. Legal maps liability to risk tier and human oversight design
  3. Procurement confirms evaluation criteria still match negotiated acceptance
  4. FinOps validates consumption caps in order form

Concession trading: Never give liability cap without getting model change notification, export rights, or eval re-test clause. Document trades in deal memo.

Warranty, insurance and indemnity schedules

Beyond limitation of liability, enterprise AI deals may require:

InstrumentPurposeAI nuance
Tech E&O insuranceProfessional negligenceVerify AI/automation covered
Cyber policyBreach, ransomwarePrompt injection exfiltration
IP indemnityThird-party IP claimsTraining data provenance
Performance warrantyShort-period defect fixDefine "defect" as eval fail not subjective dislike

Warranty period: 90 days post-acceptance common; exclude defects from client corpus changes without CR.

Insurance certificates — verification checklist

  • Named insured matches contracting entity
  • Policy period covers implementation + 12 months run
  • Limits meet client minimums by tier
  • AI/automation language not excluded in schedule
  • Certificate received before production go-live if contract requires

Worked example — MEAT acceptance annex excerpt

Use case: Internal legal contract review assistant (tier-medium).

Criterion 1 — Faithfulness: On Legal-Eval-400 (held-out clause questions), groundedness score ≥ 89% using RAGAS or client-agreed equivalent, measured over 10 consecutive business days of UAT.

Criterion 2 — Citation: For answers addressing liability caps, indemnity, or termination, 100% include link to source clause ID in corpus metadata.

Criterion 3 — Refusal: On Legal-Abstain-75 (questions outside corpus), appropriate refusal or escalation rate ≥ 94%.

Criterion 4 — Latency: p95 end-to-end ≤ 5.0s at 200 simulated concurrent sessions.

Criterion 5 — Availability: Orchestration layer ≥ 99.5% during UAT window excluding agreed maintenance.

Exclusions: System shall not provide binding legal advice; shall not auto-redline without human publish; shall not process privilege-marked docs until Class C CR complete.

Retest triggers: Foundation model change; corpus delta > 20% by token count; new jurisdiction pack.

Bidder conference and clarification questions — AI-specific

Strong clarification questions (submitted in writing):

  • "Requirement 4.2.1 specifies 100% accuracy—may we propose faithfulness on Eval-XX with human review for failures?"
  • "Who owns UAT eval set creation and maintenance post-launch?"
  • "Are prompts and system instructions client IP upon payment?"
  • "What is required subprocessor notification lead time for new model providers?"
  • "Is phase 2 autonomy (auto-send) in scope for this RFP or separate gate?"

Questions educate procurement on record—reduces unsafe language in final RFP for future cycles.

Contract lifecycle — post-signature obligations

PhaseContract mechanismAI action
MobilisationSOW kick-offBaseline eval; config management
BuildMilestone acceptancePartial MEAT where applicable
UATAcceptance annexFull eval; issue log
HypercareSLA beginsQuality SLO monitoring
Steady stateCR processModel/corpus change control
RenewalPrice hold / re-bidTCO actuals vs forecast
TerminationExit scheduleExport, delete, transition

Cross-border and framework agreements

Framework call-offs (public sector, GPO): master terms may lack AI schedules—attach call-off-specific technical schedule every time. Never assume framework covers eval acceptance.

Multi-jurisdiction: Data residency clause must list each country corpus; model routing per region; subprocessors per region in DPA appendix.

FinOps commercial linkage — contract clauses to request

From AI FinOps and Commercial Design:

  • Monthly usage report with token breakdown by model and feature
  • Budget alert integration to client FinOps tooling (webhook/API)
  • Rate card for overages pre-agreed—not list price at invoice
  • Right to route models for cost optimisation without vendor penalty if eval maintained
  • Annual true-up cap at ±15% of forecast unless CR

Monthly procurement–delivery rhythm (during implementation)

WeekActivity
1Milestone status vs SOW; CR log review
2Eval results vs acceptance thresholds
3FinOps actuals vs order form forecast
4Risk/legal touch if model or data change pending

Prevents acceptance surprise at UAT boundary.

Oral presentation and demo — contractual implications

Orals often create implied promises. Discipline:

  • Demo labels: "PoC environment; metrics from Eval-v2 dated DD/MM"
  • No ad hoc accuracy claims; slide footer with measurement method
  • Handout matches written proposal—verbal-only commitments become CR or lost

Record orals where client policy allows; send summary email within 24h confirming what was not promised.

Glossary — procurement terms for AI engineers

TermMeaning
BAFOBest and Final Offer — last negotiation round
IndemnityCompensation for specified losses
Limitation of liabilityCap on damages recoverable
Order formCommercial SKU attachment to MSA
ScheduleTechnical/legal annex (security, DPA, SLA)
SubprocessorVendor's downstream processor (model API, cloud)
UATUser acceptance testing against contract criteria
WarrantyDefect repair obligation for defined period

Fluency in terms speeds joint sessions with legal—reduces translation errors.

Final readiness — contract signature checklist

  • MEAT acceptance annex signed or incorporated
  • IP schedule complete
  • DPA subprocessors current
  • SLA severity includes harmful output
  • CR classes defined
  • FinOps reporting in SOW
  • Exit/export clause tested in PoC
  • Insurance certs verified
  • Technical evaluator named for UAT sign-off

Negative cases — when contracting fails

The 100% accuracy RFP

Symptom: RFP embeds vendor marketing language; bidders fail to challenge; UAT impossible.

Impact: Perpetual dispute; blame between SI and client; programme pause.

Fix: MEAT acceptance; educate procurement before RFP issue; bidder conference Q&A on record.

Acceptance on calendar date

Symptom: "Go-live 1 June" with no eval gate; pressure to sign acceptance regardless of quality.

Impact: Production incidents; regulatory exposure; warranty claims.

Fix: Milestone acceptance tied to eval results; steering holds date until pass or documented risk acceptance.

IP vacuum

Symptom: SOW silent on prompts and fine-tunes; vendor claims background IP on all deliverables.

Impact: Cannot migrate; re-procurement from scratch; loss of eval assets.

Fix: IP schedule in every AI SOW; export format specified pre-signature.

SLA without quality

Symptom: 99.9% uptime SLA; faithfulness drops 40% after corpus update; no breach.

Impact: Users lose trust; support tickets rise; no contractual lever.

Fix: Quality SLOs in technical schedule; regression eval on corpus change; optional service credits for eval failure.

Liability mismatch

Symptom: $500k cap on customer-facing credit decision tool.

Impact: Uninsurable residual risk; board refuses launch; emergency renegotiation under time pressure.

Fix: Risk-tiered caps; insurance certificates; phase internal before external.

Subprocessor surprise

Symptom: Model provider not listed; data processed in non-contracted region.

Impact: DPA breach; regulator notification; contract termination rights invoked.

Fix: Subprocessor notification workflow; residency clause; audit right.

Practice checklist

  • I can explain RFI/RFP/RFQ selection for a novel AI use case
  • I drafted MEAT acceptance without absolute accuracy promises
  • I mapped IP ownership for prompts, evals and fine-tunes
  • I linked SLAs to quality SLOs, not only uptime
  • I completed the primary acceptance annex exercise
  • I identified three negative cases relevant to my context
  • I cross-referenced FinOps assumptions with commercial terms

Discussion

Comments

Share feedback or questions about this page. No account required.

Loading comments…