The demo is never the hard part. The hard part is the answer being right when it is grounded in your documents, the output being consistent enough to put in front of a customer, someone owning it when it is wrong, and a cost per use that still makes sense at volume. We pick the use cases where that is achievable, build one properly, and show you what it actually costs to run.
Generative AI consultants who pick the use cases that survive production, ground outputs in your own content, and show the real cost per use before you commit.
Generative AI consulting is the work of choosing which generative use cases are worth building, grounding them in your own content so the output is accurate, and putting the evaluation, review, and cost controls around them that production requires. The modeling is rarely the difficulty. Deciding what is worth doing, and proving it stays correct, is.
Generative AI consulting selects the use cases that survive contact with production, grounds them in the client's own content, and adds the evaluation and cost controls that keep output accurate at volume. Thinklytics scopes these against the data a client actually has, so a pilot that works in a demo also works in the month after launch.
Use case selection against your real data, so the shortlist is what is buildable rather than what is fashionable.
Grounding in your own documents and records, with citations, so an answer can be checked.
Evaluation before launch and after. A written quality bar, a test set, and a score you can watch move.
Cost per use measured at realistic volume, not at demo volume.
A model training project. Almost nobody needs to train one, and we will say so.
A chatbot on your website. If that is the goal, it is a smaller piece of work than this.
An innovation workshop. The deliverable is a running system, not a list of possibilities.
Use case assessment scored on value, data readiness, and the cost of being wrong.
One production use case built end to end, grounded in your content with citations.
An evaluation set and quality bar, so accuracy is measured rather than asserted.
Cost model at real volume, plus runbooks and handover to your team.
Annual claims recovery after three stalled machine learning pilots were diagnosed and rebuilt on reconciled data.
Rebalancing opportunities identified once advisor-facing analysis was grounded in unified portfolio data.
Blocked machine learning initiatives diagnosed and unblocked by fixing what sat underneath them.
It was never scoped against a real volume, a real owner, or a real cost per use, so there was nothing to put into production.
It is answering from the model rather than from your content, with no retrieval and no citation to check against.
Output quality drifts and nobody notices until a customer does.
There is no evaluation set, so there is no measurement between launch and complaint.
Cost was estimated at demo volume, with no caching, no routing to cheaper models, and no ceiling.
They work out which generative use cases are worth building for your business, then build one properly. In practice the job is mostly judgment and engineering around the model rather than the model itself: picking a use case your data can support, grounding output in your own content, writing the evaluation that proves it is accurate, and costing it at real volume.
Almost certainly not. Most problems that look like they need a custom model are retrieval problems, which means the answer exists in your documents and the system is not finding it. Fine tuning is worth it mainly for consistent output format or specialist vocabulary, and it is a late optimisation rather than a starting point.
Ground it in your own content and make it cite what it used, so an answer can be checked. Restrict it to answering from retrieved material rather than from memory. Require human review where the output goes to a customer or changes a record. And measure it against a fixed test set so you find out that quality dropped before a customer tells you.
It depends almost entirely on volume and on how much text goes into each call, and the demo number is misleading because demos are low volume with short context. We model cost at your expected volume, including retrieval, and look at caching and routing simpler requests to cheaper models before committing. Ongoing cost is usually predictable. Maintenance is the part people underestimate.
A use case assessment is typically two to three weeks. A first grounded, evaluated use case in production is usually eight to twelve weeks, and most of that is retrieval quality, evaluation, and the review workflow rather than the generative part.
The ones where the output is checkable, the volume is real, and being wrong is recoverable. Internal knowledge answering, drafting that a person reviews, summarising long documents, and extracting structured data from unstructured files all tend to hold up. Anything that goes unreviewed to a customer, or that depends on data nobody has reconciled, tends not to.
They overlap and they sit at different points. This engagement starts earlier, when the question is still which generative use case to build. RAG Consulting is the build when you already know you need grounded retrieval over your content. AI Agent Consulting is the build when the system needs to take actions rather than produce text.
Scope comes out of the use case assessment. These are the factors that move the effort.
Clean, structured documentation is quick. Scanned PDFs, duplicates, and no owner mean cleanup first.
Internal drafting tolerates error. Customer-facing or regulated output needs far deeper evaluation and review.
One document store is simple. Permission-aware retrieval across several systems is not.
Volume drives the cost model, and the caching and routing work needed to keep the bill sane.
Where a human approves output, that workflow has to be built and fitted into how the team already works.
The assessment prices the shortlist, so you can drop a use case before it is built rather than after.
Innovation engagements end with possibilities. This one ends with something running and measured.
You have budget for generative AI and no agreed shortlist of what to build.
A pilot worked and then stalled on accuracy, cost, or ownership.
Output needs to be grounded in your own documents and checkable.
You want the cost per use known before you commit to volume.
You already know you need grounded answers over your content: see RAG Consulting.
The system needs to take actions, not produce text: see AI Agent Consulting.
A security review is the thing blocking you: see AI Security Consulting.