- Judge an AI consulting engagement by what it leaves behind. A use case inventory scored on value and feasibility, a data readiness finding per candidate, and a working evaluation harness are executable by someone else. A capability maturity chart is not.
- Discovery that never queried the source systems has produced an opinion. The single most useful test of a proposal is whether the assessment includes hands-on inspection of the data, with access arranged before the engagement starts.
- Ask what the model does when it is wrong, and what the process does in response. A proposal with no position on error handling has scoped the model and not the workflow it sits in.
- Licence, consumption and labour have to be priced separately. Bundled into a single figure, the third-year run cost of the engagement cannot be estimated by anyone, including the supplier.
- The evaluation method and the metric belong to the client, defined before the build and computed internally. A supplier who defines the measure of success and also reports it has removed the only external check on the work.
Buying AI consulting is harder than buying most technical services, because the proposals converge. Everyone names the same model families, offers a workshop and a pilot, and shows a three-horizon roadmap at similar day rates. What separates them is easier to check: which artefacts the engagement produces, whether the data was inspected before the proposal, and who runs the result. This article sets out what to look for.
What an AI consulting engagement should leave behind
The most reliable test of an engagement is whether its outputs can be picked up and acted on by a party that did not produce them. Applied to a typical assessment phase, that means these artefacts:
- A use case inventory with a score per candidate, where the score has components that can be checked: the volume of the underlying activity, the cost of the current process, the decision the output would change, and who owns that decision.
- A data readiness finding per candidate. Which fields are needed, where they live, whether they are populated, how far back the history goes, and whether they are available at the moment the decision is made rather than only in a monthly extract.
- A recommendation per candidate on build or buy, with the reasoning recorded. Most enterprise use cases are better served by a product, and a firm whose answer is always to build has an incentive worth noting.
- An evaluation harness, meaning a dataset, a metric and a script that produce a comparable number. This is the artefact that lets you evaluate the next supplier, and it is the one most often absent.
- A target operating model. Who retrains, who is paged when output degrades, who owns the decision the model informs.
- A register of open questions with owners and dates.
Two common outputs carry little weight. A capability maturity assessment positions the organisation on a scale used nowhere outside the engagement. A technology landscape summarising available model providers is public information restated. Neither can be executed, and both are inexpensive to produce, which is why they are frequently the bulk of the document.
Two kinds of AI readiness assessment
The word covers two quite different pieces of work at similar prices.
The first is a set of interviews and workshops, producing a synthesis of what people said they wanted, arranged into themes and prioritised on a matrix. It is genuinely useful for alignment, and it establishes nothing about feasibility, because every feasibility statement in it is an inference from a conversation.
The second includes hands-on inspection: someone with database access querying the tables that would feed the model, counting populated rows, checking how far the history goes, looking at how a free-text field is used in practice, and establishing when each field is written relative to the event being predicted. That last check is the one that decides whether a use case is viable, and it cannot be done in a workshop.
The distinction matters because the failure it prevents is expensive. A pilot built on an extract can succeed on data that is unavailable at decision time, or on fields that are populated only after the outcome has occurred. Both produce excellent evaluation scores and nothing operable. Confirming availability at the point of decision is cheap during assessment and costly to discover in month five.
The practical implication for procurement: arrange data access before the engagement begins. An assessment that spends its first three weeks waiting for access has lost the part of its scope that had the most value, and it will deliver the interview-based version by default. In DNA Solutions engagements, arranging that access is a pre-sales task for this reason.
Model error handling in the surrounding workflow
Every model is wrong some of the time, so the question is what that triggers in the surrounding process. A proposal that treats accuracy as the deliverable has scoped a model rather than a workflow.
Specific things worth asking for a position on:
The error profile, separately from the rate. Which direction of error is tolerable, and at what ratio. In document classification a misfile that a human catches at the next step is cheap; in a maintenance decision a missed failure is not. The acceptable ratio is a business judgement, and the supplier should be asking for it rather than supplying it.
The confidence threshold and what sits below it. Most useful deployments route low-confidence cases to a person. That routing path is part of the build, and it needs capacity assumptions: if fifteen percent of cases go to review, someone has to review them.
The fallback. What the process does if the model is switched off. A production model without a fallback path cannot be switched off, which means a degradation becomes an incident rather than a decision.
Traceability. For decisions affecting a customer or an employee, what is recorded so the decision can be explained months later. Where the process falls under the EU AI Act, the classification work belongs at design time, because it determines documentation and oversight obligations and therefore the effort.
Reading the shape of an AI consulting proposal
A few signals sort proposals quickly.
Specificity about your estate. A proposal written after serious pre-sales conversation names your systems, your data constraints and your process. One that names none has been assembled from a template, and it will meet your estate for the first time after signature.
A stated position on the hard parts. Data access, integration into the operational path, evaluation, monitoring, and the run cost. Confidence distributed evenly across everything indicates that nothing has been examined closely.
Named people against named workstreams. Which individuals, with what background, on which part, for what share of their time, and what happens if one becomes unavailable. Team size and partner biographies answer a different question. It is reasonable to ask to speak with the person who would lead the technical work, and the answer to that request is itself informative.
Something the firm declines to do. A proposal that claims value across the entire scope is describing a sales position. A firm with judgement will tell you which of your candidate use cases it thinks are not worth attempting, and why.
Exit terms. Code in your repositories, models and artefacts in your accounts, training data and evaluation sets yours, throughout rather than transferred at the end. Where a supplier's tooling or platform is part of the delivery, establish what remains operable without them.
From AI pilot to production: how to scope it
Pilots are useful and the way they are scoped determines whether anything follows. A pilot designed only to demonstrate feasibility usually has to be rebuilt for production, and the rebuild is rarely budgeted.
Three conditions make a pilot extendable. It reads from the same data path production would use, rather than from a one-off extract. It writes its output somewhere a process can consume, even if nobody consumes it yet. And its evaluation criteria were agreed in writing before it ran, including the level of accuracy below which the process would be worse off.
That last condition is the one to insist on. A threshold set after the result is known becomes a conversation about the result. The gate between pilot and operation, and what has to be true to pass it, is covered in our article on the enterprise AI roadmap.
It is also worth asking what the pilot assumes about the state of the surrounding systems. Where the data has to be assembled across systems that do not currently reconcile, that work is the project, and a proposal that treats it as a precondition someone else will meet has quietly excluded most of the effort. The pattern is described in our article on technical debt and AI projects.
AI consulting pricing and the three cost components
For AI work specifically, three cost components behave differently and should be priced separately.
Labour is the engagement itself, and it ends. Licence covers any platform, tooling or model provider commitment, and it recurs. Consumption covers inference and compute, and it scales with usage, which means it grows exactly as the project succeeds.
Bundled into a single figure, the third-year cost of the arrangement cannot be estimated. Separated, a straightforward question becomes answerable: at ten times the current volume, what does this cost. A supplier who cannot answer has not modelled it, and the answer sometimes reveals that a use case is viable at pilot volume and uneconomic at production volume.
Two related items. Where a proprietary model or platform is central, establish the switching position: what would be involved in moving to an alternative provider, and whether prompts, fine-tuning artefacts and evaluation sets are portable. And where the supplier resells a platform, ask whether they receive margin on the consumption they specify.
Success metrics and who measures them
The metric and the measurement belong on the client side. A supplier who defines the measure of success, computes it, and reports it has removed the only external check on the work, and this arrangement is common enough to be worth naming as a condition rather than assuming.
Practically: the metric is agreed before the build, in business terms as well as technical ones, meaning both the model measure and the process outcome it is meant to move. The evaluation dataset is held by the client. And a named internal person can run the evaluation, which is one reason the harness matters as a deliverable.
What good AI consulting looks like
Assembled: an assessment that included hands-on data inspection with access arranged in advance; a use case inventory with a data readiness finding per candidate; a build or buy recommendation with reasoning; an evaluation harness the client keeps; a stated position on error handling and fallback; named senior people against named workstreams; labour, licence and consumption priced separately; a target operating model naming who runs the result; and a written gate the pilot has to pass.
None of it depends on assessing methodology. It is all checkable before signature, most of it during pre-sales.
Talk through your AI consulting engagement
DNA Solutions helps European enterprises with AI and machine learning engagements run the way this article describes: the assessment starts with hands-on data inspection, the use case inventory and the evaluation harness stay on your side, and we say early where we think a candidate use case is not worth attempting. Where the sequencing has to be reconciled with a modernisation programme already under way, that part sits under IT consulting. Talk to us.
Related services: AI & Machine Learning, IT Consulting



