What AI Implementation Consulting Should Deliver
What to expect from each stage of an AI engagement, how to judge the work, and what your team should retain when the project closes.
· 16 min read

Imagine you're reviewing an AI implementation proposal with your operations lead. The consultant has understood the problem, demonstrated a promising approach, and divided the engagement into milestones. Near the end of the proposal, a milestone promises a “working AI solution.” Your operations lead asks who will look after it once the project closes.
The consultant starts explaining the support options. As you listen, you realize you had pictured something slightly different from the team beside you. One person expected software the firm could run independently. Another assumed the consultant would keep operating it. Someone else thought the current fee covered an experiment, with the full build to follow.
All three arrangements could be reasonable. They involve different work, different prices, and different things you should expect to receive. The proposal needs to make those differences clear while there's still time to choose between them.
That's a useful place to begin buying AI implementation consulting. You need to understand what you're commissioning well enough to recognize when it has been delivered. This gets easier when each stage has a defined purpose, evidence you can inspect, and an explicit decision about what happens next.
Agree on what the current fee buys
A consultant may start with a workshop because the business problem still needs definition. If so, you should leave with an agreed description of the work, its limits, the people involved, and the question the next investment will answer. A short document can hold all of that. Its usefulness becomes apparent when you and the consultant can use it to explain the proposed scope without filling in missing details from memory.
Suppose the firm wants help preparing responses to prospective clients' security questionnaires. There is already a library of approved answers, but employees spend time finding the right material and checking whether it still applies. The immediate uncertainty might be whether AI can assemble a useful first draft from that library while making unsupported questions visible to a reviewer.
A paid prototype could answer that question using historical questionnaires and copies of approved material. It wouldn't necessarily include a connection to the firm's current systems, access for the whole team, or a service somebody monitors every morning. Those could be sensible later investments, once the experiment has shown enough promise.
The Australian Digital Transformation Agency's stage definitions provide a useful distinction. A proof of concept tests feasibility in an experimental setting. A pilot introduces real users and live or near-live data within a controlled scope, producing results and feedback from actual work. Production puts the system into normal operations, with live integration, monitoring, support, recovery arrangements, and much lower tolerance for failure. This is Australian government guidance; a smaller firm can borrow the distinction without adopting the government's entire assurance process.
In your proposal, those labels need to describe the work being purchased. Ask which data the system will use, who will use it, and which systems it will be able to change. The answers make the stage much more tangible than a progress slide marked “pilot complete.”

A prototype can therefore be completed successfully even when you decide to stop. Perhaps the system drafts convincing answers, but checking every statement against current policy takes more effort than writing from the existing library. The experiment has resolved a real uncertainty. You should receive the results, examples of the problems, and an explanation of the recommendation, including what could change it.
That outcome is much easier to accept if the agreement described the purchase as an experiment. It is harder when the proposal implied production software and explains its limitations only after the budget has been spent. Equally, a buyer shouldn't expect a consultant to absorb an unpriced production build because the prototype looked close to finished. Agreeing on the stage protects both sides from that misunderstanding.
Make the acceptance decision possible before the build
Once the scope is clear, the next question is how you'll recognize acceptable work. The consultant should help you define this in terms that an experienced member of your team can judge.
For the questionnaire example, a test might check whether a draft preserves an important qualification in an approved answer. Record the question, the answer it's allowed to draw on, and the qualification the reviewer expects to see. The reviewer retains responsibility for the response sent outside the firm.
The technical term for these checks is evaluation, often shortened to evals. In its guide to evaluating AI agents, Anthropic recommends building tests from real user tasks and known failures, with criteria clear enough for domain experts to agree on a result. It also distinguishes testing before launch from production monitoring and human review. The guide is practitioner engineering advice, not evidence that any particular consulting method guarantees success.
For a buyer, the practical implication is that the test cases are part of the delivery. Keep the inputs, expected behavior, results, and important failures together. You can inspect what the consultant tested, and the same cases can help whoever maintains the system check a future change.
Some examples should be unfamiliar to the people tuning the system until it's ready for review. If the team keeps correcting the same handful of demonstrations, you learn that it can improve those demonstrations; you have less evidence about the next piece of work. Have your experienced reviewer help select this separate set of cases and retain the results alongside those used during development.
This also changes how you read an overall score. A missed formatting preference and an unsupported assurance to a prospective client have different consequences. Ask to see consequential failures separately, alongside ordinary completion quality and the work needed to check the output. A serious failure may be grounds to withhold approval for the proposed use, even when the overall score looks encouraging.
The acceptance decision should record any remaining limitations and the scope in which the system can be used. If it works for one service line but struggles with another, you could accept the narrower use, commission further work, or stop. Writing down that choice prevents a qualified acceptance from becoming an unrestricted rollout through assumption. Payment timing, and any link between payment and acceptance of deliverables, belongs in the agreed commercial terms.
Put the client's contribution on the schedule
These decisions require time from your team. An implementation partner can build the tests and explain the technical tradeoffs, but someone inside the business must decide what a professionally acceptable answer looks like. If two experienced colleagues disagree, that disagreement needs resolution before a developer can turn their judgment into a reliable test.
The proposal should identify who supplies representative material, who reviews results, who grants access, and who can accept a stage. It should also estimate the time those people need and when they'll need to be available. A firm without an internal technology team may ask the consultant or its existing IT provider to handle administration, while retaining a business owner with authority to make decisions.
That division of work belongs in the price and schedule discussion. If a consultant prices a build assuming that approved source material is ready, discovering that half the library is obsolete creates work someone must do. You might clean it internally, pay the consultant to help, or narrow the initial scope. Each choice changes what can be delivered and when.
Ask for the fee to distinguish the experiment, connections to business systems, testing, training, and the agreed period of support where those are included. Also establish who pays software subscriptions and usage charges. The consultant needn't quote every possible future change, but the agreement should explain how a new request is assessed and approved before it adds cost.
This gives you a fairer comparison between proposals. One provider may include administrator training and a period of post-launch assistance while another expects your IT team to handle both. A lower total price tells you little until you understand those obligations. Compare the defined work, the assumptions your firm must satisfy, and the point at which the engagement ends.
Rehearse the handover while help is still available
As a pilot approaches normal use, bring the discussion back to the operations lead's original concern. Who will look after this, and can they do the work the arrangement assigns them?
We would make part of the handover a practical rehearsal. The designated operator uses their own account and the supplied instructions while the consultant observes. This gives both sides a chance to find missing access or unclear procedures before the project closes. It can be a modest session for a modest system; it should still exercise the responsibilities the firm is accepting.
Begin with ordinary work. The operator opens a case, checks the output, completes any required approval, and finds the record of what happened. If they need the consultant to explain an undocumented step, improve the instructions and let them try again. The goal is that the person responsible can perform the task with the materials they'll keep.
Next, show what the questionnaire reviewer sees when the approved answer library cannot be reached. Where a safe test environment is available, interrupt that connection; for a small configured tool, a simulated case or walkthrough of a known failure may be proportionate. The reviewer should recognize that the draft is incomplete and know how to recover or escalate. They shouldn't have to guess whether retrying will create a duplicate or whether work has disappeared.
The DTA's AI evaluation guidance includes testing unusual inputs, real-world scenarios, system performance, and recovery. Our recommendation is to make the relevant failure behavior visible to the buyer. A test report is more useful when your operator can also recognize the situation it describes and knows what to do about it.
Finish by rehearsing a small change, such as replacing an approved questionnaire answer after a policy update. The operator needs to know who can make that change, how it will be checked, and how to return to the previous working version if necessary. Locate the instructions for stopping or retiring the workflow as well. The person running it needn't become an AI engineer; their responsibilities must match their training and access.

Attendance alone is thin acceptance evidence. Being able to complete the delivered workflow and handle a foreseeable problem is more informative. Whether people continue using it consistently is a question for subsequent operational review; the handover establishes that they know how to begin.
Know what you can carry forward
The rehearsal may reveal another dependency: essential access belongs to the consultant. That arrangement can be intentional in a managed service, but you should understand its consequences before you choose it.
Ownership, access, and portability are related questions. You may own custom code while the system runs in an account only the supplier can administer. You may control your data without being able to move the configured workflow to a different platform. A third-party product may remain the vendor's intellectual property even when your firm pays for its implementation. Ask what the agreement gives you and what you can practically do with it.
The DTA's AI procurement guidance recommends addressing rights to data, models, and code up front, alongside cost transparency, knowledge transfer, and exit arrangements. These are useful purchasing questions for a smaller firm, too. They don't mean every engagement should transfer every piece of a provider's technology to the buyer; the appropriate terms depend on what's being bought.
For your own handover, identify the accounts, configuration, instructions, test cases, and documentation you will receive, plus the licenses or services that must remain active. If another competent provider takes over, establish what it can access and where it would need to rebuild. An export button helps only if the exported material contains what you need and can be used elsewhere. Migration may still require work, so ask for the known limits rather than a promise of effortless portability.
Continuing with the original consultant may be the best choice. A defined support arrangement can buy someone to review failures, maintain connections, test changes, and help users resolve problems. Specify who watches for issues, the hours and channel through which help is available, the response commitment, and which changes carry an additional fee. “Ongoing support” needs enough detail for both parties to understand the job.
Also agree what happens when support ends. Someone must retain administrative access, know which charges continue, and be able to stop the workflow while preserving the business records the firm needs. That knowledge makes the arrangement easier to manage even if you never change providers.
Bring five questions to the proposal review
You can use the following questions to turn the next conversation into decisions about the engagement:
- What are we buying at this stage? Name the deliverables, the data and systems involved, and whether the stage may reasonably end with a recommendation to stop.
- What evidence will let us accept it? Agree on representative tasks, consequential failures, retained review work, and the person who makes the decision.
- What must our team contribute? Put access, source preparation, subject-matter review, and decision time beside the consultant's schedule and fee assumptions.
- What will we rehearse before handover? Have the people inheriting the responsibilities demonstrate ordinary use, failure handling, and the route for a change.
- What remains after the engagement? Confirm the assets and access you retain, the ongoing costs and support obligations, and how you can change provider or retire the system.
Before signing a proposal, have the prospective partner answer these questions against the scope and fee on the page. If you're considering The Foregrounds for implementation, use the same review with us: agree on the deliverables, your team's contribution, and the responsibilities each side will retain when the project closes.
Related articles

AI Readiness Assessment: Is Your Business Ready to Put AI to Work?
A practical way to decide what your firm can implement now, what needs preparation, and how much responsibility to give AI.
Sep 10, 2026 · 18 min read

Build, Buy, or Fix the Process First?
How to combine the software you own, the products you buy, and the parts you build into a system your team can run.
Sep 10, 2026 · 19 min read

