Back to all articlesai implementation

AI Vendor Selection for Enterprises: The Evidence to Demand Before a Pilot, Not After

An evidence checklist rather than a criteria list: the five artefacts to require a vendor to produce on your own data before you sign, who has to produce each one, what a refusal tells you, and the pilot contract terms that decide who carries the cost when a pilot does not reach production.

AI Vendor Selection for Enterprises: The Evidence to Demand Before a Pilot, Not After

Listen to this article (2 min)
0:00--:--

You have budget, a shortlist of three, and a problem: all three demo well. That is not a coincidence and it is not a sign that all three are good. A demo is a run on data the vendor chose, at a difficulty the vendor set, with the failure cases filtered out by whoever built the deck. Every vendor who reaches a shortlist can pass it.

Which is why selection frameworks that score vendors on criteria do not separate them. "Strong production controls." "Proven track record." "Clear governance model." Every serious vendor claims all of it, most of them believe it, and none of the claims can be checked at the point where you have to decide.

So stop scoring claims. Score artefacts. For each thing you need to be true, name the specific document, run, or reference that would demonstrate it, name who has to produce it, and treat a refusal to produce it as the answer to the question. A vendor who cannot show you a per-field error breakdown does not have one, and a vendor who does not have one has not looked.

Everything below is designed to be produced before you sign a pilot, not discovered in month four.

The five artefacts

What you need to knowArtefact to demandWho produces itWhat a refusal tells you
Does it work on our documents, not theirsA run on a sample you selected, including the awkward casesVendor, on data you supplyThey have only tested on clean, in-distribution inputs
Where does it fail, specificallyPer-field or per-task error breakdown, not one accuracy numberVendor, from that same runThey measure at the document level and cannot see assignment errors
Has this survived production somewhereA named reference in your industry with volumes attachedVendor, and the reference speaks to you unaccompaniedTheir production deployments are pilots that never graduated
What happens when we leaveWritten statement on model, weights, prompts, fine-tunes and data at terminationVendor's counsel, in the contractExit was never designed, and will be negotiated from a weak position
What happens when it is wrong at 2amThe escalation path, with owner and time bound per severityVendor, and your own operations leadNobody owns the output once it is live

Five artefacts. Each takes a vendor between a day and a week to produce if they already run production systems, and is impossible to produce convincingly if they do not. That asymmetry is the whole value of the list.

1. A run on your documents, not their samples

Send each vendor the same set of your own inputs and require the raw output back. Not a summary, not a slide of results — the actual output, in the format their system produces, for every item you sent.

Choose the sample yourself and choose it badly on purpose. If a hundred items is what you can prepare, spend seventy on ordinary volume and thirty on the cases your team already knows are hard: the supplier whose format changed last year, the scanned page that came through a fax, the document that spans a page break, the one in a second language, the one where two fields legitimately contain the same value. Include three items where the correct answer is no answer — the field genuinely is not on the page. What a system does when the answer is absent is the single most informative thing in the run, and it is never in a demo.

Score the returns yourself against a key you wrote before you saw any of them.

Who produces it: the vendor, on your data, at their cost. If they ask you to pay for a proof of concept before showing anything on your inputs, you are paying to find out whether the product works, which is a cost that belongs on their side of the table at this stage.

What a refusal tells you: the honest version of "we cannot run on your data yet" is a product that has only ever been tested in distribution. That is a real answer and sometimes an acceptable one for an early vendor — but it should change the price, the pilot terms and the autonomy you grant, and it should be written down rather than absorbed.

2. A per-field error breakdown, not an accuracy figure

A single accuracy number is the least useful measurement in this category, because it averages together two failures that behave nothing alike.

A model can misread characters — the recognition fails, the value comes back wrong-looking, confidence drops, and a threshold catches it. Or it can read the characters perfectly and put them in the wrong field. The second produces a well-formed, plausible value with high confidence, and no threshold anywhere catches it, because nothing about the value looks wrong. We wrote that distinction up in full in what a document extraction confidence score actually means; the short version is that vendor documentation generally defines confidence as certainty about the recognised text, not as the probability that the value belongs in the field it landed in.

So the number to demand is per field, and it separates absent, wrong, and wrong field. Ninety-four percent overall can be a system that is near-perfect on invoice number and date and unusable on tax, which matters enormously if tax is the reason you are buying.

Ask for it in this shape:

FieldPresentCorrect valueCorrect fieldFailure mode when wrong
Invoice number100%99%99%Confuses with PO number on two layouts
Total100%98%98%Picks amount due on part-paid invoices
Tax amount71%96%88%Single tax field, collapses multi-rate invoices
Line item unit price92%91%84%Row misalignment on continuation pages

Who produces it: the vendor, from the run in point one. If they produce it from their own benchmark instead, you have learned something about their benchmark and nothing about your documents.

What a refusal tells you: they measure at the document level. Document-level accuracy is a sales metric — it tells you how often nothing went wrong, which is not the same as knowing what goes wrong when something does.

There is a second, sharper reason to insist on the field-level view: "trained for this task" is not one capability. On invoices specifically, the three major cloud parsers disagree in their published schemas about whether line-level tax exists at all — we walked through exactly where they diverge. A vendor building on one of them inherits that shape whether or not they mention it, and the per-field table is where it becomes visible.

3. A named production reference in your industry, with volumes

Not a logo. Not a case study the vendor wrote. A named company, in your industry, running the same class of decision, with a number attached to it: items per month, users, how long it has been live.

Then ask to speak to them without the vendor on the call, and ask these five:

  1. What changed between the demo and the first month in production?
  2. Which decisions stayed human longer than you expected, and why?
  3. Where did the review queue become the bottleneck?
  4. What happened the first time the output was wrong in a way that reached a customer?
  5. If you ran the selection again, what would you write differently into the contract?

The fifth question is the one that pays. Nobody answers it with "nothing".

Who produces it: the vendor introduces, the reference speaks freely. A reference call the vendor sits on is a sales call with extra steps.

What a refusal tells you: if every reference is under an NDA that prevents them naming volumes, the volumes are small. If every reference is a pilot, nothing has graduated. Both are survivable facts about an early vendor; neither is survivable as a surprise.

4. What happens to the model and the data at termination

This is the artefact buyers most often skip and most often regret, because it is the only one whose value is entirely in the future and the only one that cannot be renegotiated later from a position of strength.

Require it in writing, covering:

  • Your data. Deletion on what timeline, verified how, including copies in logs, backups, evaluation sets and any vendor-side monitoring.
  • Derived artefacts. Fine-tuned weights, adapters, prompt libraries, extraction templates, labelled data your team produced during the engagement. Who owns each. "The model improves for everyone" is a legitimate business model; it just needs to be a sentence in the contract rather than a discovery.
  • Portability. In what format, and with what notice, you can extract the configuration that encodes your process — the rules, thresholds, routing logic and approval boundaries. This is usually months of your operators' work, and it is the thing that makes switching cost real.
  • Continuity. What runs, and for how long, between notice and cutover.

Who produces it: the vendor's counsel, into the contract. Not a support-page link, not a security questionnaire response.

What a refusal tells you: exit was never designed. That is common and it is not automatically disqualifying — but the cost of it lands on you, and it should be priced in now rather than in the year you want to move.

5. The escalation path when the output is wrong in production

Every system of this kind is wrong sometimes. That is not the risk. The risk is being wrong at 2am with nobody named.

Demand a single page, per severity level: what triggers it, who is paged, what the response time is, what the system does in the meantime, and who decides to switch the automation off. Then check the last line especially — the ability to fall back to the previous process, immediately, without the vendor's involvement, is what makes the whole engagement reversible.

Ask for one real incident. Not a hypothetical: an actual case where their system produced a wrong output that reached a customer or a ledger, what the detection latency was, and what changed afterwards. A vendor with production deployments has this story. A vendor without one will tell you their system has not been wrong yet, which tells you they are not measuring.

Who produces it: the vendor for the mechanism, and your own operations lead for whether the path is one your team can actually run at 2am. The second half is the one buyers forget, and it is the half that fails.

What a refusal tells you: nobody owns the output once it is live. Ownership that is not named before signature does not appear afterwards.

The pilot contract is where the evidence gets teeth

Most selection processes treat the contract as paperwork that follows the decision. For anything in this category the contract is part of the evidence, because it is where the vendor's confidence in their own claims becomes financial.

Three terms decide most of it.

Who pays if the pilot does not reach production. Someone will. Right now it is you by default: you will have spent your team's time, your data preparation, your integration work, and you will have paid for the pilot. A vendor who believes their own evidence will share some of that exposure — a reduced fee, a fee contingent on the agreed threshold, a credit against the first production term. The specific structure matters far less than the fact of the conversation, because how a vendor responds to the question is itself an artefact.

What success means, in numbers, written before the pilot starts. Not "improved accuracy" or "demonstrated value". A threshold, per decision class, on the metrics from artefact two, measured on a sample defined before the run: straight-through processing above X% on standard-layout invoices, with assignment errors on tax below Y%, measured on the 300 documents in the agreed test set. One number per decision class, and a stated consequence for missing it.

Why a definition written afterwards is not one. This is the trap, and it is almost always set with good intentions. The pilot runs, the results come in mixed, and everyone in the room — including your own team, who now have sunk months in it — reads the results in the most favourable available light. Not dishonestly; that is simply what an undefined threshold does. A success criterion agreed after the evidence exists is unfalsifiable by construction, which means the pilot proved nothing and you are making the original decision again with less budget and more commitment.

Write the numbers down while you still do not know the answer. That is the only moment they can be honest.

One more clause worth naming: what happens to autonomy scope if the evidence comes back worse than expected. The right answer is that the scope narrows and the engagement continues at a smaller boundary — not that it stops, and not that it proceeds as planned. A contract that only has "proceed" and "terminate" forces a bad choice on a mediocre result, which is the most likely result.

Which lane you are buying decides how much evidence you need

Not every decision needs all five artefacts at full strength. What sets the bar is where the decision sits:

  • Delegate — the agent decides; your team audits. Everything above, in full, plus the per-field breakdown on your own data as a gating threshold. You are removing a human from the loop, so the evidence has to stand in for their judgement.
  • Surface — the agent prepares; a person approves. The bar moves. Here the question is not "is it right" but "does it make the reviewer faster without making them credulous", and the artefact to demand is a measured acceptance rate from a live deployment, plus what happens to review time. A recommendation nobody has time to check is a decision that has been quietly delegated.
  • Hold — stays human. You are not buying a decision system at all; you are buying assistance. The evidence bar is low and so should the price and the autonomy be.

A vendor who cannot have this conversation at the level of individual decisions — who answers "our agent handles accounts payable" when you asked which specific decisions it makes alone — is describing a product, not a deployment. That gap is usually the whole story.

What to do before the next vendor call

  1. Assemble the sample first. A hundred of your own items, chosen to include the ugly cases, with an answer key your team wrote. This one artefact is more discriminating than any scorecard, and it is yours to reuse across all three vendors.
  2. Write the success numbers before you talk to anyone. Per decision class, with a threshold. If you cannot write them, you do not yet know what you are buying, and no vendor process will tell you.
  3. Send the five artefacts as a list, to all three, at the same time. How long each takes to respond, and what they push back on, is a signal you get for free.
  4. Put the termination clause in the first draft, not the last. Its cost is lowest before you have chosen.

Where we sit

We run a paid diagnostic: we sit inside one operation, classify its decisions into Delegate, Surface and Hold, and hand you a written decision ledger for it. It is short, and if AI is not the right answer for that operation we say so and stop there. The reason the diagnostic is paid is the same reason the artefacts above are worth demanding — evidence that costs the person producing it nothing is not evidence.

If you are running a selection now, the ledger tells you which lane each decision belongs in, which in turn tells you how hard to push on each of the five artefacts. That is usually the difference between a shortlist you can decide between and three vendors who all demo well.

FAQ

What should an enterprise ask an AI vendor for before signing a pilot?

Five artefacts. A run on a sample of your own inputs that you selected, including the difficult cases and some where the correct answer is that no answer exists. A per-field or per-task error breakdown from that run, separating absent values, wrong values and correct values in the wrong field. A named production reference in your industry with volumes attached, who will speak to you without the vendor on the call. A written statement of what happens to your data, any fine-tuned weights, and your process configuration when the contract ends. And the escalation path when the output is wrong in production, with a named owner and a time bound per severity level.

Why is a single accuracy figure not enough to evaluate an AI vendor?

Because it averages two failures that behave differently. Recognition errors produce wrong-looking values with low confidence, so a threshold catches them. Assignment errors produce correct values placed in the wrong field, with high confidence, so no threshold catches them and they reach systems of record looking perfect. A single number hides the split, and it also hides which fields fail: a system reported at ninety-four percent can be near-perfect on invoice number and unusable on tax, which decides the purchase if tax is why you are buying.

Who should pay if an AI pilot does not reach production?

That is a term to negotiate before the pilot, not a question to answer afterwards. By default the buyer absorbs the whole cost, including their own team's time and data preparation, which gives the vendor no exposure to the outcome. A vendor confident in their evidence will usually share some of it — a contingent fee, a reduced rate, or a credit against a production term. How they respond to the question is itself informative, regardless of where it lands.

Why must pilot success criteria be defined before the pilot starts?

Because a criterion written after the results exist is unfalsifiable. Once a team has spent months on a pilot, mixed results get read in the most favourable available light — not dishonestly, but because that is what an undefined threshold does. Define one number per decision class, on a test set agreed in advance, with a stated consequence for missing it. Also define what happens on a mediocre result: the scope should be able to narrow rather than forcing a choice between proceeding as planned and terminating.

Need help with AI implementation?

We build production AI systems that actually ship. Not demos, not POCs—real systems that run your business.

Get in Touch

Trusted by operators at

DMartSwiggyFreightTigerBoldfitZapkeyDPDZeroPopular Motor Ventures