Skip to content

Guide · 9 min read · Sep 18, 2026

How to choose an AI implementation partner: 10 questions to ask

Choosing an AI implementation partner comes down to ten questions about scoping, testing, human control, ownership, lock-in, security, integration feasibility, who does the work, how success is measured and what happens after launch. Ask every shortlisted firm the same questions and compare the answers side by side. Then verify with reference calls and a small, paid first piece of work before committing to a large build.

Tuaha JawaidFounder, business lead

Key takeaways

  • Ask every shortlisted firm the same ten questions and write down the answers, so you compare evidence rather than impressions.
  • The strongest signal is whether a firm insists on understanding the workflow before proposing a technology.
  • Ownership of code, data, accounts and documentation should be settled in the contract, not discussed at handover.
  • Reference calls with companies of your size tell you more than a polished portfolio does.
  • Start with a paid discovery or audit and a small first scope, so both sides learn before the money gets large.

Most AI proposals read well. They describe a problem you recognize, show an interface you like and end with a number. What they rarely show is how the firm will decide what to build, how it will prove the result is correct, and what you are left holding when the engagement ends. Those are the things that separate a working system from an expensive pilot, and they are all answerable in a first conversation if you ask directly.

What is an AI implementation partner, and when do you need one?

An AI implementation partner builds and integrates AI into your existing operations: connecting systems, automating steps in a workflow, extracting information from documents, or making company knowledge searchable. That is different from a strategy consultancy, which advises but does not build, and different from a software vendor, which sells you a product to configure.

You need one when the value comes from your own processes and systems rather than from a product anyone can buy, and when no one internally has the combination of AI engineering, integration and cloud skills available. If the choice between hiring, contracting and partnering is still open, in-house, agency or partner works through that decision first.

The stakes are worth stating plainly. In BCG's 2024 survey of 1,000 senior executives across 59 countries, only 26 percent of companies had developed the capabilities to move beyond proofs of concept and generate tangible value, leaving 74 percent with nothing to show. The same research attributes around 70 percent of the difficulties to people and process issues, 20 percent to technology and only 10 percent to the algorithms themselves. Choosing a partner who treats the work as an operations problem rather than a modeling problem is therefore not a stylistic preference.

The ten questions, and what good answers sound like

Send the same list to every firm on your shortlist. Write the answers down. The comparison is far more revealing than any individual conversation.

#QuestionA good answer sounds likeRed flag
1How will you scope this before building anything?A paid discovery or audit with named workflows, interviews with the people doing the work, and a ranked list of options including "do not use AI here"A fixed proposal produced from one call, or a demo presented as a scope
2How will you test accuracy, and against what?A test set built from our real past cases, acceptance criteria agreed with our business owner, and scores reported by failure type"You will see when it works", or an accuracy percentage quoted before seeing our data
3Where does a human stay in control?Named approval points for consequential actions, a defined route for low-confidence cases, and an owner for the exception queueEmphasis on how little human involvement will be needed
4Who owns the code, the data and the accounts?We own the repository, the cloud accounts, the model provider accounts and the documentation, stated in the contractCode held in the supplier's environment, or ownership described as something to sort out later
5What if we want to change model or provider later?Model calls behind one interface, provider choice in configuration, and evaluation sets that can be rerun against a replacementA proprietary platform that "handles all the AI", or an answer that treats a model change as a rebuild
6How will our data be handled and secured?Specific answers on where data is processed and stored, retention, who can read prompts and logs, and the terms of each provider usedGeneric assurance that everything is encrypted, with no reference to the actual provider terms
7How will you confirm the integrations are feasible?Checking APIs, permissions, license tiers, rate limits and data quality before committing to a date, with each system's administrator identifiedA connector list or logo wall offered as proof that integration is straightforward
8Who exactly will do the work?Named people, their roles, the share of their time, and how specialists are brought inA team assigned at kickoff, or a refusal to name anyone before signature
9How will we know whether it worked?A baseline measured before the build, a target agreed with the sponsor, and the measurement method written into the proposalBenefits expressed as percentages from a market report rather than from our operation
10What happens after launch?Monitoring, model and dependency updates, incident response and support scoped explicitly, with hours and owners, under a separate agreementThe assumption that hosting includes support, or silence until something breaks

A firm does not need a flawless answer to all ten. What you are looking for is specificity, willingness to say what is not included, and consistency between what they say in the meeting and what appears in the proposal.

The four questions that matter most

Scoping before building (question 1)

The most reliable predictor of a good engagement is whether the firm insists on understanding the work before proposing technology. A paid discovery or audit is a reasonable thing to buy: it is how a partner earns the right to quote responsibly, and it should leave you with a current-state assessment, prioritized opportunities, a recommended approach and an implementation roadmap that another team could act on. What happens in an AI readiness assessment sets out what to expect and what to demand as deliverables.

Be wary of the opposite pattern, where every answer arrives at the same platform. A recommendation that could have been written before meeting you was.

Testing and accuracy (question 2)

Ask how the partner will prove the system does what you agreed, and listen for real cases rather than adjectives. Managed evaluation tooling exists on both major clouds, including Amazon Bedrock evaluation jobs that can score outputs automatically, with a judge model or with human reviewers using your own prompt dataset, so "we will build a test set from your data" is an ordinary request rather than an exotic one. The part that cannot be bought is your business owner agreeing in advance what correct means for each task. How to test AI before it touches your operations covers the method in detail, and is a fair thing to send a prospective partner before a meeting.

Ownership (question 4)

Settle four things in writing before work starts:

  1. Code and configuration. In your repository from the beginning, not delivered at the end.
  2. Cloud and model provider accounts. Ideally yours, with your own billing relationship. If a partner hosts, agree in writing who has administrator access, where logs live, what the exit process is and how usage is billed.
  3. Data. What the partner may access, where it is processed, how long anything is retained, and what is deleted at the end.
  4. Documentation and prompts. Detailed enough that a different team could take over. This is the test that reveals whether ownership is real.

Security and data handling (question 6)

The answers should reference the actual terms of the providers involved, not general reassurance. Those terms are public and specific. OpenAI states that data sent to its API is not used to train or improve its models unless you explicitly opt in, that abuse monitoring logs may be retained for up to 30 days by default, and that Zero Data Retention is an optional control requiring prior approval. Anthropic states that it will not use your chats or coding sessions to train its models unless you choose to participate in its Development Partner Program. AWS describes Amazon Bedrock running models in deployment accounts operated by the Bedrock service team that model providers cannot access, meaning providers do not see customer prompts and completions, while the shared responsibility model leaves your data, identities and configuration as your accountability.

If you need a framework your risk team already recognizes, the NIST AI Risk Management Framework, released in January 2023 with a generative AI profile added in July 2024, organizes the conversation around four functions: govern, map, measure and manage.

One more point belongs under question 5. Model retirement is published policy, not a hypothetical. Microsoft Foundry sets a retirement date 18 months from launch for generally available models, 12 months for models from Anthropic, DeepSeek, Fireworks and Mistral AI, with at least 60 days notice and no extensions. Amazon Bedrock moves models through Active, Legacy and end-of-life states, with a Legacy notice period of 6 months or 45 days, and states that migration will not happen automatically. A partner who has no answer for that is proposing work you will have to redo.

How do you run a fair evaluation?

Illustrative example: Consider an operations director at a 250-person logistics business shortlisting three firms for document processing work. She sends all three the same ten questions and the same one-page description of the workflow, with a two-week deadline. One replies with a fixed price and a platform demo. One asks for sample documents and volumes, then proposes a paid audit with named deliverables. One asks who administers the transport management system and whether its API tier is licensed, which is the question nobody internally had checked. She scores the responses on a shared sheet with her IT manager, then takes two reference calls each for the two firms that engaged with the workflow.

Four practices keep an evaluation honest:

  • Ask everyone the same questions, in writing, with the same deadline. Score the answers as a group rather than deciding alone after the most persuasive meeting.
  • Take reference calls, and prepare them. Ask what went wrong and how it was handled, whether the timeline held, who actually did the work, and what the company would scope differently. References from companies close to your size and sector are worth more than logos.
  • Buy a small piece of work first. A paid discovery, audit or single workflow lets both sides learn how the other operates while the commitment is still small. What you learn about responsiveness and honesty in four weeks is not available from a proposal.
  • Keep the first scope narrow and measurable. One workflow, one owner, one baseline. Broader programs are easier to justify once something is live and measured.

Then read the contract for the things that are awkward to raise later: ownership of code and data, what happens to accounts at the end, notice periods, what is excluded from support, and whether any part of the arrangement requires exclusive use of a particular provider.

Warning signs across the whole process

  • Urgency created by the supplier rather than by your business.
  • Results quoted as percentages with no named client, baseline or measurement period behind them.
  • No willingness to say "buy this product instead" or "fix the process first."
  • Undisclosed reseller or referral relationships that could shape the recommendation.
  • Scope expanding during the sales process, before anyone has understood the brief.
  • Hosting presented as though it includes monitoring, backups and incident response.
  • Discomfort when you ask to speak to the engineers who would do the work.

How Kastling answers these questions

For fairness, here is how Kastling would answer its own list. Kastling is an AI solutions and implementation partner whose lead service is AI Integration & Automation, alongside AI & Cloud Infrastructure and Custom Solutions, and the same ten questions should be put to us with everyone else.

Engagements start with a free discovery call. For AI and operations work a separately scoped, paid audit usually follows, producing a current-state assessment, prioritized opportunities, a recommended approach and a roadmap; it is not mandatory for every project, and a well-defined software brief can move straight to a scoped proposal. Development is tested against real business scenarios with feedback from the people who will use the solution, then integration and training follow, with optional maintenance under a separate agreement.

The principles behind those answers are consistent: understand the work before the technology, connect existing systems rather than replacing them, keep a named person in control of consequential decisions, agree how success will be measured, involve the people who will use the solution, and agree data access and ownership up front. Deployment can run in your AWS account or Azure subscription or in Kastling's accounts, with account ownership, access, operating costs and responsibilities set out in the proposal, and Kastling-hosted usage billed on a pay-as-you-go basis. The team is founder-led, a business lead and an AI engineering lead, with specialists including AWS and Azure expertise brought in as the work requires. Pricing is scoped per engagement on a call rather than published, and we do not present intended outcomes as achieved results.

Questions

Should we shortlist a specialist AI firm or our existing software supplier?

Ask both the same questions and see who is specific about testing, data handling and human control. An existing supplier already knows your systems and your people, which is worth a great deal, but general development skill does not automatically include evaluating model accuracy or designing approval steps. A reasonable answer is often to use the incumbent for integration work and bring specialist help in for the AI-specific parts.

Is it a problem if a partner cannot name existing clients?

Not on its own, since many engagements are covered by confidentiality and newer firms will have fewer. What matters is whether they can describe, in detail, how they would approach your workflow, and whether they can offer any reference at all, including anonymized ones. Treat the absence of references as a reason to start smaller, not necessarily as a reason to walk away.

How much should we expect to pay for an AI implementation?

It depends on the number of systems involved, the state of the data, the level of security review and how much human process design is needed, which is why credible ranges only appear after scoping. Ask what drives cost up or down in your specific case, and ask for the ongoing costs (model usage, hosting, support) to be quoted separately from the build. Any figure quoted before anyone has seen your workflow is a placeholder.

What if the audit concludes we should not build anything?

That is a useful outcome and worth paying for. It usually means the process needs fixing first, the data is not available, or a product you can buy already does the job. A partner willing to reach that conclusion is demonstrating the judgment you are hiring for, and the assessment still leaves you with a ranked list of what to do instead.

Can we run the evaluation ourselves without a consultant?

Yes. The ten questions here, a shared scoring sheet and three reference calls are within reach of any operations or IT leader. Bring in outside help only if the workflow touches regulated data or a system nobody internally administers, where a second technical opinion on feasibility is worth the cost.

Sources

  1. BCG: AI Adoption in 2024, 74% of Companies Struggle to Achieve and Scale Value
  2. NIST: AI Risk Management Framework
  3. OpenAI: How your data is used
  4. Anthropic: How do you use personal data in model training?
  5. AWS: Amazon Bedrock data protection
  6. AWS: Evaluate the performance of Amazon Bedrock resources
  7. Microsoft Learn: Foundry Models lifecycle and support policy
  8. AWS: Amazon Bedrock model lifecycle

Guide · 14 min read

AI workflow automation: a practical guide for operations teamsAI workflow automation puts AI models inside a defined business process to read, classify and prepare work, while fixed rules, integrations and named people handle the steps that must be predictable. It suits repetitive, document-heavy or request-heavy work such as intake, routing, document checks and approvals. Start with one measurable workflow, keep consequential approvals with an accountable person, and expand only once the numbers show it works.Read

Article · 9 min read

AI document processing for invoices, contracts and formsAI document processing combines three techniques: OCR to turn images into text, layout models to understand where values sit on a page, and language models to extract and interpret fields that vary between documents. The technique matters less than what surrounds it: validation rules that check extracted values against your own records, confidence thresholds that decide what a person sees, and a review queue somebody owns. Measure field level accuracy and the share of documents that pass without a human touch, not a single headline accuracy number.Read

Article · 9 min read

How to test AI before it touches your operationsTesting AI before deployment means scoring it against a set of real past cases, with the business owner agreeing in advance what counts as correct and what target has to be met. Run it in shadow mode or a limited pilot before it can change anything, keep human review on consequential decisions, and rerun the same test set every time a prompt or model changes.Read

One useful email a month.

New guides, tools and practical notes on putting AI to work. No sales sequences.

Bring us a workflow.

Tell us where the work slows down. We will help you see where to start.

Book a discovery call