Four quotes.
No way to compare them.
One vendor wants $40,000, another wants $600,000, and both demos looked convincing. The gap is rarely capability. It's what each one is quietly assuming about your data, your volume and who cleans up when the output is wrong. I ask those questions for you.

The demo is designed to work.
Every AI demo you will ever see runs on data the vendor chose. That isn't dishonest, it's how demos work. The problem is that your buying decision depends entirely on the cases the demo left out: the malformed documents, the ambiguous ones, the fifteen percent where the right answer requires knowing something about your business that isn't in the file.
Knowing how to evaluate AI vendors comes down to testing that gap before you sign. I take your real historical cases, including the ugly ones, and I make the vendor run them. I ask how output quality is measured and ask to see the numbers. I price the run cost at your volume rather than the volume in their pricing page. And I read the contract for the clauses that decide what happens when the system is confidently wrong, which is the scenario nobody negotiates until it happens.
I take no referral fees and no vendor relationships, which is the only reason this advice is worth anything. Sometimes the recommendation is to buy from the cheapest one on the list. Occasionally it's that all four should be declined and the problem solved another way. I serve US and Canadian companies, and I've sat on both sides of this table: I've bought these systems and I've built them.
Also known as: AI vendor evaluation, AI due diligence checklist, AI vendor selection, AI procurement support, AI RFP evaluation, AI platform assessment.
Four ways a good-looking
vendor disappoints.
None of these are visible in a sales cycle. All of them are findable in a two-week evaluation.
The thin wrapper
A prompt, a model API and a nice interface, priced like proprietary technology. Legitimate as a product, badly overpriced as one.
Accuracy with no denominator
"Ninety-four percent accurate." On what test set, scored by whom, and what happens in the other six percent?
The pilot that can't scale
It works at a hundred documents a day. At four thousand the latency, the cost curve or the human review queue makes it unusable.
Your data, their moat
The contract lets them train on your inputs and gives you no export path. Two years in, switching costs more than the original build.
Five areas, and the
questions that separate them.
The full checklist goes into the report. These are the areas it covers and the questions that do the most work.
Capability, tested
Whether it works on your cases, not on theirs.
- →Run fifty of our real historical cases, including the failures
- →How do you measure output quality, and can we see the test set?
- →What does the system do when it isn't confident?
- →Show us a customer at our volume and let us call them
Architecture & substance
What is actually theirs, and what is a model API with a subscription attached.
- →Which model providers do you depend on, and what happens if one changes pricing?
- →What have you built that a competent team couldn't rebuild in a quarter?
- →How do you version and evaluate changes to prompts and models?
- →Where does our data go, and which subprocessors touch it?
Economics at your volume
The number that decides whether you keep this in year two.
- →Total cost per transaction at our actual monthly volume
- →What happens to the price when volume doubles, and when it halves?
- →How much human review does the workflow still require?
- →What's the all-in cost of the integration work on our side?
Risk, security & compliance
The part legal will ask about after the build is done, so ask it first.
- →Data residency, retention and whether inputs are used for training
- →Security posture, penetration testing and incident history
- →How the system's decisions are logged for audit
- →Alignment with the obligations you already carry to your own customers
Contract & exit
Negotiated before signature, because afterward you have no leverage at all.
- →Performance commitments with a remedy attached, not just an SLA on uptime
- →Export of your data and your configuration in a usable format
- →Price protection at renewal and on volume growth
- →Who is liable when a wrong output causes a real loss
Two weeks,
before you sign.
Longer for a platform decision with four or more vendors in scope.
Frame the decision
- Agree what the system has to do and how you'll know it did
- Assemble fifty real cases, weighted toward the hard ones
- Set the comparison criteria before anyone sees another demo
- Establish the run-cost model at your real volume
Put the vendors through it
- Technical sessions with each vendor's engineers, not the account team
- Run the test cases and score the results the same way for everyone
- Reference calls with customers at comparable scale
- Security and data-handling review
Score and negotiate
- A single comparison table with evidence behind each cell
- Contract review focused on performance, exit and liability
- Negotiation support, including the questions that move price
- A written recommendation with the reasoning and the risks
Make every vendor run the same fifty of your ugliest historical cases. The scoreboard writes itself, and it rarely matches the demo.
About evaluating vendors.
How do I evaluate AI vendors and AI companies?
How to evaluate AI companies selling into this space, in one sentence: test them on your data before you compare their pricing. Assemble fifty real historical cases weighted toward the difficult ones, make every vendor run the same set, and score the results yourself against criteria you set before the demos started. Then price the run cost at your actual volume, ask how they measure output quality and to see the test set, and call a reference customer at comparable scale. The vendor who welcomes that process is usually the one worth buying from.
What questions should I ask an AI vendor?
Five do most of the work. What is your accuracy measured against, and who assembled the test set? Where does a case go when the system isn't confident? Total cost per transaction at our volume, including the human review we'd still be doing? Which model providers are you dependent on, and what happens when one of them changes pricing? And if we leave, can we export our data and configuration in a usable format? Vague answers to the last two are the most reliable warning sign in the whole process.
Is an AI automation agency legit?
Plenty are. The category also has a low barrier to entry, so it attracts people whose entire capability is assembling a low-code workflow. The distinction is not whether they build on top of an existing model, since almost everyone does and there's nothing wrong with it. The distinction is whether they can tell you how quality is measured, what it costs to run at your volume, and who owns and debugs the system after launch. An agency with an evaluation process and a named handover plan is legitimate. One that judges quality by whether customers complain is selling you a demo.
What belongs on an AI due diligence checklist?
Five areas. Capability tested on your own cases rather than the vendor's. Architectural substance, meaning what they actually built versus what they subscribe to. Economics at your volume, including the human review the workflow still needs. Risk, covering data residency, training rights, logging and security. And the contract terms for performance remedies, data export and liability when an output causes a loss. The report I produce works through all five with evidence rather than assertions.
Do you take referral fees from vendors?
No, and I don't hold vendor partnerships. It's the only structure under which this advice is worth paying for. I'm regularly the person telling a client that the expensive option is right, and just as regularly the person telling them that none of the four should be signed and the problem is better solved another way.
What does a vendor evaluation cost?
It's quoted to scope and depends on how many vendors are in the running and how complex the decision is. A two-vendor comparison is usually one to two weeks of work; a platform decision with four vendors and a security review runs closer to three. Against a contract in the six figures, the cost of the evaluation is small next to the cost of picking wrong. That is the argument, rather than anything above it.
Related paths.
How to hire an AI consultant
The brief to write first, the nine questions, and the answers that should worry you.
Technical Due Diligence
The same discipline applied to an acquisition target rather than a vendor.
AI Governance & Risk
What happens after you sign: policy, review gates and who answers when it's wrong.
Got quotes you
can't compare?
Send me the shortlist and what you're trying to achieve. I'll tell you which questions would separate them fastest, whether or not you hire me to ask them.