A scorecard for choosing an AI development partner
Shortlists get decided on impressions unless something forces them to be decided on evidence. This is a weighted scorecard for comparing development suppliers on the things that predict how an engagement goes, rather than on how well the firm presents. Score every supplier on the same criteria, including this one, and be suspicious of any score you cannot point at a piece of evidence for.
- Weight verification highest. It predicts more than anything else on the list.
- Score only what you have evidence for. An unevidenced score is an impression.
- A supplier who scores badly on one thing you do not need is not disqualified.
- Run it on your incumbent too, if you have one.
Score each supplier 0 to 5 on each criterion, multiply by the weight, and total it. The number matters less than the argument it forces.
The criteria
Verification, weight 5. What runs automatically on every change, does it block a merge, and can they show a check failing? Score 5 only if you watched a check fail. This is weighted highest because it is the strongest available predictor of what reaches your customers.
Specification practice, weight 4. Do they treat the specification as the expensive artefact, and will they show you one? A supplier working from loose tickets is going to pass unclear requirements directly to a model.
Accountability, weight 4. Who owns it when something breaks after release, and for how long? Score on the arrangement, not on how friendly the promise sounds.
Handover, weight 4. What do you get at the end besides code? Specification, checks, decision record, runbook. Score 1 if the answer is "the repository".
Security, weight 3. Automated checks on the paths where untrusted input reaches a page, a log, or a query. "Our developers are experienced" scores 1.
Domain fit, weight 3. Have they worked on something structurally similar? Not the same industry necessarily, the same type of problem.
Track record, weight 2 to 5. Set this weight yourself, honestly. If a checkable history is what your decision needs, weight it 5, and accept that new firms score 0 and should. If method matters more to you, weight it 2.
Communication, weight 3. Did they tell you something you did not want to hear during the sales process? That is the single best predictor of whether they will during delivery.
Commercial clarity, weight 2. Do you understand what happens to the price when scope changes? Vagueness here becomes an argument later.
How to use it
Fill it in from evidence, not impression. If you cannot name what a score is based on, it is a feeling and should be left blank rather than guessed.
Do it independently if several people are deciding, then compare. Where two people scored the same supplier differently, that gap is the useful conversation.
Scoring us
Reveneau scores 0 on track record and cannot argue otherwise. If you weight that at 5, we lose to an established firm on arithmetic, which is a legitimate way to make the decision. We would rather you used a scorecard and reached that conclusion than picked on a feeling and got it wrong in either direction.
Where we fit
Reveneau suits this when
- You have two or more credible suppliers and no clear way to separate them
- Several people have to agree on the decision
- You need a written record of why a supplier was chosen