Thinklytics

Buyer Guide · 8 min read · September 2026

What to Ask Before You Choose an AI Model or Vendor

By Thinklytics Partners, Buyer Guide

Benchmarks, demos and token price are close to useless as predictors. Where the data goes, what it costs to leave, and whether you can tell when it degrades are the questions that decide it. Plus the build, buy or configure decision, which has three answers rather than two.

Model choice matters far less than most procurement processes assume, and the things that do matter, where your data goes, what the second year costs, how you leave, are usually not on the evaluation scorecard at all. Choose for exit cost and data boundary, not for benchmark scores.

Why the usual evaluation goes wrong

What predicts a working deployment, and what does not

The three unticked items are what a typical selection process compares. All three are close to useless as predictors.

  • Published benchmark rankings. They measure general capability on public tasks. A model that ranks third on a public leaderboard routinely outperforms the first on a particular extraction job. The ranking is not about you.
  • Demo quality. Every demo works, because demos are built on curated data. Ask for it to be run on twenty of your own worst examples, chosen by you, live.
  • Price per million tokens. A fraction of total cost. Retrieval infrastructure, evaluation, monitoring, the integration work and the person who keeps the content current will usually exceed the model bill.
  • Where the data goes, in writing. Training use by default or under any circumstance, jurisdiction of processing, retention including logs and abuse monitoring, and who at the vendor can read it under what process.
  • What it costs to leave. Assume in two years you want to move. Are the prompts, evaluation sets and retrieval index portable, does the integration work transfer, and is the semantic layer yours?
  • Whether you can tell if it got worse. If the answer involves users reporting it, that is not monitoring, that is your customers doing quality assurance.
  • Fit to the actual task. Run the shortlist on the same fifty real examples, scored by the people who do the work today, blind to which model produced which output. It takes about a week.

Choose for exit cost and data boundary, not for benchmark scores.

Source: Thinklytics AI readiness practice, 2026.

A typical selection compares models on published benchmarks, demo quality, and price per million tokens. All three are close to useless as predictors of whether the deployment will work.

Benchmarks measure general capability on public tasks. Your task is narrow, specific, and uses your vocabulary. A model that ranks third on a public leaderboard routinely outperforms the first on a particular extraction job. The ranking is not about you.

Demos are built on curated data. Every demo works. The question is what happens to the invoice with the handwritten annotation, the contract with the scanned amendment stapled to the back, the ticket written in three languages. Ask for the demo to be run on twenty of your own worst examples, chosen by you, live. The response to that request tells you more than the result does.

Token price is a fraction of total cost. Retrieval infrastructure, evaluation, monitoring, the integration work, and the person who keeps the content current will usually exceed the model bill.

What to actually evaluate

Where the data goes, in writing

Not we take security seriously. The specific answers:

  • Is our content used for training, by default or under any circumstance, and is that contractual or just current policy?
  • Where is it processed, in which jurisdiction, and can we pin it?
  • How long is it retained, including in logs and abuse-monitoring systems, and can that be reduced to zero?
  • Who at the vendor can read it, under what process?

If you are in healthcare, financial services, government or education, these answers determine whether the deployment is possible at all, which makes them the first questions rather than the compliance annexe.

What it costs to leave

The question that predicts regret. Assume in two years you want to move.

  • Are the prompts, the evaluation sets and the retrieval index portable, or expressed in a proprietary format?
  • Does the integration work you are about to pay for transfer, or is it written against one vendor's SDK?
  • Is the semantic layer and the data model yours, held in your systems?

An architecture where the model is a replaceable component costs slightly more to build and is worth it the first time pricing changes or a better model appears, both of which will happen inside the payback period.

Whether you can tell if it got worse

Ask how you will know the system has degraded. If the answer involves users reporting it, that is not monitoring, that is your customers doing quality assurance.

You need an evaluation set, a fixed collection of real inputs with known-correct outputs, held by you, that runs on every change. Without one there is no way to tell an upgrade from a regression, and models change under you whether or not you asked.

Fit to the actual task

Run the shortlist on the same fifty real examples, scored by the people who do the work today, blind to which model produced which output. It takes about a week and it settles the question that no amount of vendor material will.

The build, buy or configure question

Build, buy or configure

Three options rather than two, and the middle one is where most organisations belong.

OptionWhat you getWhen it is right
Buy the productA finished application for a common problem: contract review, invoice capture, service desk deflection. You accept someone else's definitions and workflow.When your process is standard, which it is more often than people like to admit. Fast, and cheapest to start.
Configure a platformYour data, your definitions, your workflow, on infrastructure someone else maintains.Where the majority of sensible enterprise AI work sits. Underrepresented in vendor material because it is less profitable to sell than either alternative.
BuildA system you own end to end, and staff forever.When the capability is a competitive differentiator, when no product fits the process and the process is the advantage, or when the data cannot leave your boundary at all.

Source: Thinklytics AI readiness practice, 2026.

Usually asked as build versus buy. There are three options and the middle one is where most organisations belong.

Buy the product. A finished application for a common problem: contract review, invoice capture, service desk deflection. Fast, cheapest to start, and you accept someone else's definitions and workflow. Right when your process is standard, which it is more often than people like to admit.

Configure a platform. Your data, your definitions, your workflow, on infrastructure someone else maintains. This is where the majority of sensible enterprise AI work sits, and it is underrepresented in vendor material because it is less profitable to sell than either alternative.

Build. Justified when the capability is a competitive differentiator, when no product fits the process and the process is the advantage, or when the data cannot leave your boundary at all. Building because the platform did not quite fit is how organisations acquire software they have to staff forever.

The failure mode is not choosing wrong. It is choosing without writing down which of the three you chose and why, so that in eighteen months nobody can say whether the thing being maintained was supposed to be maintained.

The questions, condensed

Seven questions for the shortlist meeting

A firm that answers all seven without deflecting is worth shortlisting regardless of which model they recommend.

AskWhat it settles
Is our data used for training, contractually, and where is it processed?In healthcare, financial services, government or education, this determines whether the deployment is possible at all.
What is retained, for how long, and who can read it?Retention includes logs and abuse-monitoring systems. Ask whether it can be reduced to zero, and under what process a person at the vendor sees it.
Run this on twenty examples we choose, now.The response to that request tells you more than the result does.
What does year two cost, all in, including the people?Token price is the smallest line item in the estate.
How do we know if it degrades, and who holds the evaluation set?Models change under you whether or not you asked, and without a held evaluation set you cannot tell an upgrade from a regression.
What does leaving cost, and what transfers?An architecture where the model is a replaceable component costs slightly more to build and is worth it the first time pricing changes.
Which of build, buy or configure is this, and what would change the answer?The failure mode is choosing without writing it down, so in eighteen months nobody can say whether the thing being maintained was supposed to be maintained.

Source: Thinklytics AI readiness practice, 2026.

Take these to the shortlist meeting.

1. Is our data used for training, contractually, and where is it processed? 2. What is retained, for how long, and who can read it? 3. Run this on twenty examples we choose, now. 4. What does year two cost, all in, including the people? 5. How do we know if it degrades, and who holds the evaluation set? 6. What does leaving cost, and what transfers? 7. Which of build, buy or configure is this, and what would change the answer?

A firm that answers all seven without deflecting is worth shortlisting regardless of which model they recommend.

Frequently asked questions

Do model benchmarks predict whether a deployment will work?

Rarely. Benchmarks measure general capability on public tasks, and your task is narrow, specific and uses your vocabulary. A model ranked third on a public leaderboard routinely outperforms the first on a particular extraction job. The ranking is not about you.

What should we evaluate instead of benchmarks?

Four things. Where the data goes and under what contract. What it costs to leave. Whether you can tell when quality degrades. And fit to your actual task, measured by running the shortlist on the same fifty real examples, scored blind by the people who do the work today.

What data questions should we ask a model vendor?

Is our content used for training, by default or under any circumstance, and is that contractual or current policy. Where is it processed, in which jurisdiction, and can we pin it. How long is it retained, including in logs and abuse-monitoring systems, and can that go to zero. Who at the vendor can read it, under what process.

How do we avoid vendor lock-in with AI?

Assume in two years you want to move, then check three things. Are the prompts, evaluation sets and retrieval index portable or held in a proprietary format. Does the integration work transfer, or is it written against one vendor's SDK. Is the semantic layer and data model yours, in your systems. An architecture where the model is a replaceable component costs slightly more and pays for itself the first time pricing changes.

What is an evaluation set and why does it matter?

A fixed collection of real inputs with known-correct outputs, held by you, that runs on every change. Without one there is no way to tell an upgrade from a regression, and models change under you whether or not you asked. If the answer to how will we know it degraded involves users reporting it, your customers are doing quality assurance.

Is it build or buy for enterprise AI?

There are three options, not two. Buy a finished product when your process is standard. Configure a platform with your data, definitions and workflow on infrastructure someone else maintains, which is where most sensible enterprise work sits. Build when the capability is the differentiator, no product fits the process, or the data cannot leave your boundary.

Why is token price a poor basis for choosing?

It is a fraction of total cost. Retrieval infrastructure, evaluation, monitoring, integration work, and the person who keeps the content current will usually exceed the model bill. Optimising the smallest line item is a common and expensive mistake.

Topics covered

  • AI vendor selection
  • model selection
  • build vs buy AI
  • AI procurement
  • data residency
  • AI evaluation set

Frequently asked questions

Do model benchmarks predict whether a deployment will work?

Rarely. Benchmarks measure general capability on public tasks, and your task is narrow, specific and uses your vocabulary. A model ranked third on a public leaderboard routinely outperforms the first on a particular extraction job. The ranking is not about you.

What should we evaluate instead of benchmarks?

Four things. Where the data goes and under what contract. What it costs to leave. Whether you can tell when quality degrades. And fit to your actual task, measured by running the shortlist on the same fifty real examples, scored blind by the people who do the work today.

What data questions should we ask a model vendor?

Is our content used for training, by default or under any circumstance, and is that contractual or current policy. Where is it processed, in which jurisdiction, and can we pin it. How long is it retained, including in logs and abuse-monitoring systems, and can that go to zero. Who at the vendor can read it, under what process.

How do we avoid vendor lock-in with AI?

Assume in two years you want to move, then check three things. Are the prompts, evaluation sets and retrieval index portable or held in a proprietary format. Does the integration work transfer, or is it written against one vendor's SDK. Is the semantic layer and data model yours, in your systems. An architecture where the model is a replaceable component costs slightly more and pays for itself the first time pricing changes.

What is an evaluation set and why does it matter?

A fixed collection of real inputs with known-correct outputs, held by you, that runs on every change. Without one there is no way to tell an upgrade from a regression, and models change under you whether or not you asked. If the answer to how will we know it degraded involves users reporting it, your customers are doing quality assurance.

Is it build or buy for enterprise AI?

There are three options, not two. Buy a finished product when your process is standard. Configure a platform with your data, definitions and workflow on infrastructure someone else maintains, which is where most sensible enterprise work sits. Build when the capability is the differentiator, no product fits the process, or the data cannot leave your boundary.

Why is token price a poor basis for choosing?

It is a fraction of total cost. Retrieval infrastructure, evaluation, monitoring, integration work, and the person who keeps the content current will usually exceed the model bill. Optimising the smallest line item is a common and expensive mistake.

Related reading