Buying AI in the public sector runs into a specific problem: procurement rules assume you can specify what you are buying and verify you received it. Language models are probabilistic, they change under you when the vendor updates them, and "it works well" is not a measurable acceptance criterion.
None of that makes AI unbuyable. It does mean the solicitation has to be written differently.
Start by classifying what you are actually buying
Three quite different purchases hide behind the word "AI," and they belong on different vehicles:
- A commercial product with AI features. A case management system that now summarizes documents. Buy it like software: Schedule or cooperative, ordinary evaluation.
- Model access. API capacity from a provider, consumed by systems you already run. Closer to a cloud services purchase, priced by usage.
- A custom build. Development services producing a system that uses AI. This is a professional services buy — IDIQ task order or Schedule labour categories — and should be scoped in phases.
Conflating these is the most common early mistake. A custom build competed as a product purchase produces bids that cannot be compared.
What belongs in the statement of work
Standard IT language leaves gaps here. Add these explicitly:
- Data handling. State plainly that agency data must not be used to train the vendor's or any third party's models. Require it in writing, not in a marketing FAQ.
- Data residency and retention. Where inputs and outputs are stored, for how long, and how deletion is verified.
- A measurable accuracy target against a defined test set, with a stated method. "High accuracy" is unenforceable. "≥ 90% on the 200-item evaluation set in Appendix B, measured monthly" is.
- Human review requirements. Which decisions may never be fully automated, and what the reviewer sees.
- Explainability and citation. For anything advising a decision about a person, require the system to surface its source material.
- Model change notification. The vendor must tell you before swapping or deprecating the underlying model, because it can change behaviour overnight.
- Ownership. Who owns the prompts, the fine-tuned artifacts, the evaluation set, and the output. Say so.
- Accessibility. Section 508 applies to AI interfaces exactly as it does to everything else.
Terms to refuse
- Any right to use agency data for model training or product improvement.
- Unilateral changes to the model or its behaviour without notice.
- Blanket disclaimers of all liability for output, on a system the agency will act on.
- Auto-renewal, which many public entities cannot lawfully accept anyway.
- Data export only in a proprietary format, or at a punitive price.
Buy it in phases
The strongest structure for a custom AI build is three separately funded phases with a real decision point between each:
- Discovery. Fixed price. Produces the data assessment, the evaluation set, and a firm estimate. The agency owns the output regardless of what happens next.
- Pilot. Limited scope, real data, measured against the evaluation set. Ends with a go/no-go supported by numbers.
- Production. Only funded if the pilot cleared its threshold.
This is not caution for its own sake. It converts an unpredictable purchase into three predictable ones, and gives the contracting officer a defensible file at every step.
Security and authorization
An AI system handling agency data is a system, with all the usual obligations. It needs a boundary, a control set, and an authorization path — see our walkthrough of what an ATO involves for realistic timelines. Where the model is consumed as an external API, the data flow across that boundary is the thing assessors will focus on hardest.
Inheriting from already-authorized infrastructure matters even more here than usual, and it is worth choosing a deployment pattern with that in mind before the architecture is settled.
The honest framing
Most agency AI value right now is unglamorous: summarizing documents, extracting data from forms, routing correspondence, and helping staff find answers in material they already own. Those are tractable, measurable, and low-risk relative to anything that decides about a person.
Start there, write the accuracy target into the contract, and keep a human in the loop on anything consequential. For the broader governance framing, the guide to buying AI responsibly in government covers the policy side, and contract vehicles covers the mechanics.
Scoping an AI purchase? We will help you write a statement of work that can actually be evaluated.

