AI Automation · 9 min read · May 2026
How to evaluate AI workflow automation vendors: 14 questions to ask before signing
By Thinklytics Partners, AI Automation Practice
A practical evaluation framework for AI workflow automation consultants. Fourteen questions that surface whether the consultant has actually shipped this kind of work in stacks like yours.
What are the 14 questions to ask an AI workflow automation vendor?
They cluster into four groups. Data integration (3 questions), agent observability (4), permission model (3), and pricing transparency (4). The full list is in the article body. Vendors that answer all 14 in writing are usually safe to pilot. Vendors that won't are not.
Most "AI workflow automation" pitches sound the same once you've sat through five of them. Same demo videos, same hand-wave on integration, same promise that the agent will learn your business. The selection problem isn't information. It's filtering. You need a question set that surfaces whether the consultant has actually shipped this kind of work, in stacks like yours, with the constraints you're going to hit.
This is that question set. We use it internally on prospect calls so we can disqualify ourselves quickly when we're a bad fit. The same questions work in the other direction for buyers who want to spot a vendor selling above their delivery capacity.
What this is
A 14-question evaluation framework for choosing an AI workflow automation consultant. It covers fit, scope, technical depth, change-management, post-launch operations, and commercial terms. Use it on every shortlisted vendor. Score the answers. Don't just listen to them.
What this is not
This is not a vendor scorecard for AI platforms (Zapier, n8n, Make, LangChain, custom Python). The platform choice is downstream of the consultant choice. A good consultant can build on whatever fits your constraints. Pick the partner first.
The 14 questions
Fit and scope
1. Show us a workflow you've shipped that's still running 12 months later.
Anyone can demo a clever prototype. The hard work is the workflow that survives org changes, vendor API rev-locks, the data source that gets re-permissioned, and the new compliance review. Ask for a live one. If they can't name a 12-month-old production workflow with a real owner, they're selling pilots.
2. Where do you draw the line between automation and AI agents?
If the answer is "everything is an agent now," they're using the marketing glossary. A useful answer: deterministic workflows for high-volume repeat work, agents for branching judgment calls, and a clear test for which side of the line a given task lands on.
3. What's the smallest engagement you've taken, and the largest?
Tells you whether they can actually scope down to your problem or whether they need an enterprise budget to be interested. A consultant who's done both is a better fit for a mid-market buyer than one whose smallest project was \$400k.
Technical depth
4. Walk us through how you'd connect this workflow to our data warehouse, specifically with our table-level access controls intact.
The honest answer involves service accounts, scoped credentials, and a permission audit trail. The dishonest answer is "we'll just give it admin access for now." If they pick the second answer, walk away.
5. What happens when the model output is wrong? What's the human-in-the-loop escalation path?
Every production AI workflow needs a fallback. The good answer involves confidence thresholds, human review queues, and a clear retry policy. The bad answer is "the model is very accurate."
6. How do you handle prompt versioning and rollback?
If they're shipping prompts straight to production with no version control, they'll silently break your workflow the next time they "improve" it. Ask to see the change log of a prompt that's been iterated on five times.
7. What's your stance on fine-tuning vs. retrieval-augmented generation?
Wrong answers: "we always fine-tune" or "we always use RAG." Right answer: RAG by default, fine-tune only when retrieval keeps missing high-value patterns and the data volume justifies the maintenance overhead.
Integration and operations
8. How long until the workflow is generating output we'd be willing to ship without manual review?
A vendor who quotes "two weeks" is either selling a toy or hasn't built one in production. The honest range for a non-trivial workflow is 4 to 12 weeks of supervised running before manual review goes away. Even then, most production workflows keep human-in-the-loop on the high-value tail.
9. Who owns the workflow once you leave?
If the answer is "we do, on a managed-services contract," your vendor is selling lock-in. The right answer is "your team, with documentation, runbooks, and 30 days of overlap support during handover." Managed services is fine as an explicit contract. It should be a choice, not the only option.
10. Show us your runbook for an outage.
A vendor with operational chops has a runbook. If they have to invent the answer on the call, they haven't run a workflow at scale.
Change management
11. Who on our side will use this workflow daily, and how will you train them?
The right answer involves shadowing, paired runs, and a written cheat sheet for the people on your side who will press the buttons. The wrong answer is "we'll send a recorded demo."
12. What's your usual post-launch usage pattern?
Get the truth on how often a delivered workflow gets adopted. Industry average for self-built AI workflows is brutally low, sub-30% three months post-launch. A vendor who quotes higher should be able to point at why their delivery method beats the average.
Commercial terms
13. How do you charge: fixed-scope, time and materials, or outcome-based?
Each model has a failure mode. Fixed-scope incentivizes shipping the cheapest thing that meets the spec. Time and materials incentivizes spending. Outcome-based incentivizes scoping problems where the outcome is easy to measure. What you want is a vendor who can articulate the failure mode of their own pricing model. That means they've thought about it.
14. What's a workflow you turned down, and why?
The point of this question is to verify they can say no. Anyone who has only ever said yes is selling capacity, not judgment.
How to use the framework
Score each answer 0 to 2:
- 0: vague, evasive, or contradicts evidence
- 1: credible but generic; could apply to any consultant
- 2: specific, demonstrates pattern memory from real deliveries
Anything below 18 of 28 across the 14 questions, walk. The bar isn't high. The median answer-quality on these questions is actually poor, which is why most "successful" pilots stall at the production-handover line.
Common questions
How long should the evaluation phase take?
Two to three calls per shortlisted vendor over two weeks. Anyone who pressures you to decide faster is managing their pipeline, not yours.
Should we run a paid pilot to evaluate?
Sometimes, yes. For a workflow you're going to build anyway, a paid 4-week pilot is a reasonable de-risking mechanism. Refuse to pay for a pilot whose output isn't part of the production scope. That's just consulting-flavored sales engineering.
What if the vendor we like answers question 9 with "managed services"?
Negotiate the option. Insist on documentation and runbooks as deliverables even if you stay on managed services. The deliverable is your insurance against vendor lock-in. The managed-services contract is a comfort layer on top.
How do we know when the workflow is ready to scale beyond the pilot?
Three signals together: usage rate above 70% of the eligible cases for two consecutive weeks, manual override rate below 10%, and at least one specific failure mode caught and fixed by the human-in-the-loop pathway during pilot. That last one matters more than the first two.
What's the fastest red flag in a sales call?
When the vendor's senior person stops talking and the demo person takes over without coming back. The senior person isn't there to close. They're there to sit with the technical risk. If they leave the room mid-call, the engagement is going to feel that absence later.
If you've got two or three vendors shortlisted and you want a second pair of eyes on the answers, that's exactly the kind of work covered by our AI Workflow Automation Consulting practice. We'll join your evaluation calls, score each vendor's answers against this framework, and give you a written recommendation with the failure modes to expect from each option.
You can also see how this evaluation logic plays out on real deliveries in two recent case studies: a regional health plan's intake workflow and a SaaS company's revenue-ops automation. Both of those started with a version of this question set on the buy side.
Frequently asked questions
What are the 14 questions to ask an AI workflow automation vendor?
They cluster into four groups. Data integration (3 questions), agent observability (4), permission model (3), and pricing transparency (4). The full list is in the article body. Vendors that answer all 14 in writing are usually safe to pilot. Vendors that won't are not.
Why does pricing transparency matter so much in AI vendor evaluation?
AI workflow vendors price by run, by token, by action, or by user. The pricing model decides the cost curve as you scale, and most vendors are vague on purpose. The four pricing questions force a per-run worked example with your expected volume so you can compare quotes on equal terms.
What is the most common AI vendor red flag?
Demos against vendor-prepared data. If the demo does not run on your actual sample data within the pilot, the demo is theater. Insist on a pilot with one of your own workflows in the first 30 days, with success metrics agreed upfront.
How do you evaluate AI agent observability before buying?
Ask for a sample agent action log from a real customer (redacted), the full prompt the agent received, and the tool calls it made. If the vendor cannot produce this, the agent is opaque and you will not be able to debug it when it goes wrong in your environment.
Should we pilot more than one AI workflow vendor at once?
Two, in parallel, on the same workflow, with the same success metric. One vendor in isolation gives no comparison baseline. Three or more dilutes the engineering attention. The pilot length should be 30 days, with a go/no-go decision the day after.
How does Thinklytics help with AI vendor evaluation?
We run the 14-question process with you, score each vendor on a common rubric, and recommend the one that fits your data layer and your buying constraints. The engagement is typically 6 to 8 weeks end to end. Read more at AI agent consulting.
Should we ask vendors for customer references?
Yes, and ask for references at companies in your size band running the workflow you're piloting. References at 50x your size tell you about enterprise concerns you don't have; references at 1/10th your size tell you about scale problems you'll hit. Match the reference profile to your own.
What if no vendor can answer all 14 questions in writing?
Pilot the one that comes closest, with a 30-day cap and explicit go/no-go criteria. Vendor maturity in AI workflow automation is uneven enough in 2026 that no single vendor is the obvious answer for every environment. The 14 questions filter; the pilot decides.
Frequently asked questions
What are the 14 questions to ask an AI workflow automation vendor?
They cluster into four groups. Data integration (3 questions), agent observability (4), permission model (3), and pricing transparency (4). The full list is in the article body. Vendors that answer all 14 in writing are usually safe to pilot. Vendors that won't are not.
Why does pricing transparency matter so much in AI vendor evaluation?
AI workflow vendors price by run, by token, by action, or by user. The pricing model decides the cost curve as you scale, and most vendors are vague on purpose. The four pricing questions force a per-run worked example with your expected volume so you can compare quotes on equal terms.
What is the most common AI vendor red flag?
Demos against vendor-prepared data. If the demo does not run on your actual sample data within the pilot, the demo is theater. Insist on a pilot with one of your own workflows in the first 30 days, with success metrics agreed upfront.
How do you evaluate AI agent observability before buying?
Ask for a sample agent action log from a real customer (redacted), the full prompt the agent received, and the tool calls it made. If the vendor cannot produce this, the agent is opaque and you will not be able to debug it when it goes wrong in your environment.
Should we pilot more than one AI workflow vendor at once?
Two, in parallel, on the same workflow, with the same success metric. One vendor in isolation gives no comparison baseline. Three or more dilutes the engineering attention. The pilot length should be 30 days, with a go/no-go decision the day after.
How does Thinklytics help with AI vendor evaluation?
We run the 14-question process with you, score each vendor on a common rubric, and recommend the one that fits your data layer and your buying constraints. The engagement is typically 6 to 8 weeks end to end. Read more at AI agent consulting.
Should we ask vendors for customer references?
Yes, and ask for references at companies in your size band running the workflow you're piloting. References at 50x your size tell you about enterprise concerns you don't have; references at 1/10th your size tell you about scale problems you'll hit. Match the reference profile to your own.
What if no vendor can answer all 14 questions in writing?
Pilot the one that comes closest, with a 30-day cap and explicit go/no-go criteria. Vendor maturity in AI workflow automation is uneven enough in 2026 that no single vendor is the obvious answer for every environment. The 14 questions filter; the pilot decides.