AI

How to choose an AI development company

Editorial · Reveneau · July 28, 2026

How to choose an AI development company

Almost anyone can set up an API call to a model and get an impressive result on the first try. We see it constantly: a founder shows us a demo built in a weekend, and it genuinely works. The question that matters is not whether someone can produce that demo. It is whether the company in front of you can turn it into software that keeps working when real users arrive with inputs nobody planned for. That is a different skill, and most companies claiming to build "AI features" or "AI agents" have never actually tested it.

We build these systems for a living, and AI vendors try to sell to us as often as to anyone. The pattern for spotting the real ones is consistent: ask about the failure case, not the success case. Here is the checklist we use.

1. Ask them to show you a demo failing, not succeeding

Any vendor can show you their system getting the right answer. Ask them to show you what it does when the input is ambiguous, when there is little retrieved data, or when the user asks something outside the system's scope. A company that has actually released AI features will have this ready, because they have seen it happen in production. A company that has only ever built demos will look surprised by the question, or will change the subject to how rare that failure is.

A demo is a single run, in a clean setting, with an input someone chose because it works. Production traffic is thousands of runs a day, full of inputs nobody predicted, and the system has to fail safely when it gets confused. If a vendor cannot show you the failing case, they have not built the part of the product where the real risk is.

2. Ask to see their evaluation set

This is the single most revealing question you can ask, and it is the one most vendors are least prepared for. An evaluation set is a list of real inputs paired with known good answers, that a team runs the system against every time they change a prompt, a model, or a retrieval setting. It turns "it seems to work" into a number you can track.

Ask specifically: how many cases are in it, where did the cases come from, and what happened the last time a change made the score go down. A company doing real evaluation work will answer in specifics, often naming real user queries collected after an early launch and a concrete story about catching a regression before it was released. A company that is not doing this will answer in generalities, "we test thoroughly" or "quality is a priority," language that describes an intention rather than a process. If a vendor cannot name a number of test cases, they are not measuring anything. They are hoping.

3. Ask what happens when the model is wrong

This is the question that separates a company that only builds demos from a company that builds production AI. Every model gets things wrong sometimes, and the systems that stay reliable are built around that fact rather than around the hope that it will not happen.

A real answer names mechanisms: checking answers against source data before showing them to a user, routing low-confidence outputs to a human, logging what happened so the team can see the failure pattern, and having a plan for a bad output that already reached someone. A weak answer talks about the model being "very accurate" or the team being "confident in the results." Only a specific check that runs on every output protects you, and confidence is no substitute for one. If the person you are talking to cannot describe what stops a wrong answer from reaching your customer, nothing is stopping it.

4. Check whether they have tied you to one model or one vendor

Ask how the AI calls are connected to the product's code. If your application talks directly to one provider's API everywhere a model is used, you have a hidden dependency. Model prices change, providers deprecate versions, and a better model for your exact task appears more often than most teams expect. A company that understands this builds a separate layer: the rest of your product calls an internal interface, and that interface calls whichever model is the right fit today.

Ask them directly: if your provider raised prices tomorrow, or deprecated the exact model version you are running, how much of the product would need to be rewritten? If the honest answer is "most of it," that is a business risk you are inheriting, not a technical detail. The stronger answer is that swapping models means changing a configuration and re-running the evaluation set.

5. Ask how they choose a model in the first place

There is no single best model, only the right model for a given task and budget. A good AI development company chooses based on what the feature actually needs: how accurate the answer must be, how fast it must respond, what it can cost per call, whether the data is allowed to leave your environment, and whether a smaller model tuned for the task beats a larger general one.

Watch for a vendor whose answer to "which model would you use for this" is the same regardless of what "this" is. That usually means they have a favorite tool, not a method. Ask them to test candidate models against a small version of your own evaluation set rather than pointing you to a public leaderboard. A benchmark score tells you how a model performs on someone else's tasks, not how it performs on yours.

6. Watch for the "it worked on my test case" mistake

This is the mistake almost every team makes once, including good ones. Someone runs the AI feature against a handful of inputs, the answers look good, and the team calls it done. The inputs were picked, consciously or not, because they were likely to work. That is not testing a feature. That is confirming a hope.

Ask what inputs they tested before calling something ready, and whether any were adversarial: confusing phrasing, missing information, a request slightly outside the system's intended scope. A team that has made this mistake before will have a real answer, often a specific story about a case that looked fine in testing and broke in production. A team that has not made it yet will describe a small number of clean examples and call that validation.

7. Ask who watches the system after launch, and what they look at

Launch is not the end of the work for an AI feature the way it often is for ordinary software, because model behavior changes over time and a prompt that worked well last month can quietly get worse. Ask who owns monitoring after the feature is released, and what specifically they watch: what users are actually asking, where answers go wrong, and how cost and response time move over time.

If nobody on the team can answer this in one sentence, nobody is watching, and you will find out the feature has degraded only when a user complains loudly enough.

8. Ask what the feature costs at your expected volume

A demo that runs a few dozen times costs nothing worth mentioning. A feature running against real traffic costs money on every single call, and that cost can rise fast if nobody is watching it. Ask directly: what would this cost per 1,000 requests at the volume we expect, and what do you do to control it?

A team that has actually run AI features at high volume will talk about caching repeated queries, routing simple requests to smaller and cheaper models, and reducing how much context gets sent with each call. A team with no answer here has not operated anything at real volume, and you will be the one who discovers the cost problem, on your invoice, after launch.

Warning signs, put plainly

Three signals are worth weighing on their own, and are a reason to reject a vendor when they appear together: an all-demo portfolio with nothing described about what happens after launch, an inability to explain what happens when the model is wrong, and no plan for cost at high volume. Any one of these is a caution worth a follow-up question. All three together mean you are looking at a company that only builds demos, not a development partner.

What we do differently

We should declare our own position before giving advice about choosing between suppliers: we write one hundred percent of our code with AI, and a named engineer reviews every line of it before it reaches your branch. Ask us the questions in this post and you will get numbers rather than adjectives, and you should expect the same from anyone else you are considering.

We build AI features and agents the way we build any product: starting with the problem and the failure case, not the model. Before we write a line of integration code, we ask what a good answer looks like and how we will know if the system gets it wrong. We keep an evaluation set for every feature we release, we build a separate layer so you are never tied to one model or provider, and we watch real traffic after launch. We cover the full detail of that approach on our AI development page.

Related reading

The checklist above is the short version. These cover specific parts of this decision in more depth:

None of the questions in this checklist require special technical expertise to ask. They require paying attention to whether the answers name real mechanisms or just describe good intentions. A company that has actually released AI features will answer plainly, because they have experienced the failure case you are asking about. A company that has only built demos will avoid answering it directly. That difference is the whole decision.

Common questions

How do I tell if an AI development company can build a real product and not just a demo?

Ask what happens when the model gives a wrong answer. A company that only demos will describe normal use in detail and say little about failure. A company that releases production systems will describe monitoring, a fallback, and a way for a wrong answer to get caught before a user sees it.

What questions expose whether a company actually does evaluation work?

Ask to see their evaluation set: real inputs paired with known good answers that they score against every time they change a prompt, a model, or a retrieval setting. If they cannot describe a specific set of test cases, they are not measuring quality, they are guessing at it.

What does it mean for an AI vendor to tie you to one model?

It means your product code calls a specific provider's API directly, with no separate layer of code between the product and the model, so switching providers means rewriting the feature. Ask how hard it would be to swap the model if the price doubled or the provider shut down. If the honest answer is "very hard," that is a real business risk, not a technical detail.

What is the "it worked on my test case" trap in AI development?

It is the mistake of judging an AI feature by a handful of inputs someone picked because they already knew the answer would look good. Production traffic is full of inputs nobody predicted. A team that only ever shows you hand-picked examples has not tested the feature, they have tested their own confidence in it.

What are the biggest warning signs when evaluating an AI development company?

Three: an all-demo portfolio with no monitoring story, an inability to explain what happens when the model is wrong, and no answer for what a feature costs to run once real traffic hits it. Any one of these on its own is a caution. All three together mean the company has not released an AI feature that kept working with real users.

How should an AI development company handle model selection?

By matching the model to the task and the budget, not by defaulting to whichever model they know best. Ask what they would use for your specific case and why, and ask what they would use for a cheaper or faster version of the same feature. A vendor with one answer for every problem has a favorite tool, not a method.

Why does cost at high volume matter when choosing an AI vendor?

Because model calls cost money on every single request, and that cost is invisible in a demo that runs a few dozen times. Ask what a feature costs per 1,000 requests at your expected volume, and ask what they do to control it: caching, routing simple tasks to smaller models, or reducing context. A vendor with no answer has not run anything at real volume.

What should I ask about safety checks before hiring an AI development company?

Ask what stops the system from doing or saying something it should not. A real answer names specific checks: validating output before it reaches your database, blocking responses that are not grounded in your data, and limiting which tools an agent is allowed to call. A vague answer about "safety" with no mechanism behind it is not a safety check.

Is a working prototype enough to judge whether a company can build my AI product?

A prototype tells you the idea is possible. It does not tell you whether the company can make it reliable under real traffic. Ask what the same team's plan is for turning that prototype into a monitored, evaluated, cost-controlled feature, and treat that answer as the real test.

How is hiring an AI development company different from hiring a general software agency?

The core hiring checks are the same: who writes the code, how they communicate, who owns the output. AI adds more questions, because the system's behavior is not fixed the way normal code is. The extra questions are about evaluation, safety checks, model flexibility, and monitoring, since those are the parts that decide whether an AI feature stays correct after it is released.

Should I ask an AI development company for references from AI-specific projects?

Yes, and ask the reference the same question you would ask about any vendor: did the system ever produce a wrong or unsafe output in production, and how did the team find out and respond? A reference who says nothing ever went wrong either has not used the feature enough to know, or is not being candid.

What is a reasonable first engagement to test an AI development company before committing to a full build?

A small, scoped piece of work: one use case, a real evaluation set, and a working version with safety checks, measured against real cases before the scope grows. How they handle that first piece, what they measure, and how honestly they report the results, tells you more than any pitch deck.