The scorecard

The scorecard stays with you. So take it now.

The scorecard is the instrument of the head to head in phase 1, and of the measured assessment in phase 2. It turns impressions into scores anchored to cases. Here are the families of criteria I use, the scoring principles, and a builder that produces your own version, weighted your way.

No sign-up, no email requested. The file is produced in your browser and never leaves it.

The principles

Three rules, without which the scorecard is worthless.

  1. Every score is anchored to a case.

    The rationale column is not decoration. It holds the number of the case from the evaluation set that produced the score. A score with no case is an opinion, and an opinion does not survive a committee.

  2. Weights are set before results are seen.

    Weighting after the fact means picking the winner and then writing the rules. If a weight must change, the change is dated and justified.

  3. The same scorecard applies to every solution, without exception.

    Including the solution you already have, and including the option of doing nothing. Those two are often the most instructive candidates.

How to score

A short scale beats a fine one. Zero to five, forcing yourself to describe what a 0 and a 5 look like before you start.

The weighted score orders, it does not decide. Two solutions three points apart out of a hundred are tied: what settles it is the knockout criteria, the gaps on critical cases, and the trajectory across campaigns.

Mark as knockout any criterion whose failure disqualifies, whatever the overall score. Compliance is often one of them.

The six families

Five families covering use, a sixth covering what comes after.

Family The question it answers What makes it measurable
Answer quality Does the solution answer correctly on our cases, including the ambiguous ones? Ground truth established by your experts, before any test.
Functional coverage Does it handle the whole scope, or only the easy part? The share of cases in the set that the solution accepts and handles.
Robustness on edge cases What happens at the boundaries, on traps, and under load? The edge and trap cases in the set, and the spread between two runs.
Cost in use What does the real volume cost, once in production? Unit cost multiplied by your observed annual volume, not by the pilot volume.
Compliance Where does the data go, who can justify a decision, and to whom? Contractual and technical documentation, checked with the papers in hand.
Exit, lock-in and sovereignty What is left if you change your mind in three years, and who do you depend on in the meantime? Actual portability of your data and configuration, the existence of alternatives, and where processing and decisions are located.

The first five families are those of the evaluation protocol. The sixth applies as soon as the commitment goes beyond a pilot: it does not measure the solution, it measures what leaving would cost you. It covers three things that often get conflated: vendor lock-in, meaning the exit cost designed into the product; risk concentration on a single supplier with no equivalent on the market; and sovereignty, meaning the law your data falls under and your organisation's ability to continue if the relationship ends.

The builder

Compose your scorecard, take the file with you.

Uncheck what does not apply, adjust the weights, add your own criteria, name the solutions you are comparing. The resulting file opens in a spreadsheet and is filled in by hand, case by case.

Appears at the top of the file. For example: drafting assistance for written replies, call transcription, extraction of supporting documents.

A short scale forces a judgement. Five levels are almost always enough.

Solutions compared

Each solution adds two columns to the file: the score and the reference case behind it. Eight at most.

Criteria and weighting

Weights are relative: only their proportion matters. They are normalised to one hundred in the file.

  • Answer quality
  • Share of answers matching ground truth on the most frequent cases.
    9
  • Quality of the answer when a case admits several defensible readings.
    8
  • Can the solution say it does not know, rather than produce a confident wrong answer.
    7
  • Functional coverage
  • Share of the cases in the set that the solution accepts and can handle.
    7
  • Real integration effort with existing applications, reference data and working habits.
    6
  • Robustness on edge cases
  • Behaviour on cases that look like an everyday case without being one.
    8
  • Spread in the answers when the same input is submitted several times.
    7
  • Behaviour and response time at the volume observed in production, not at pilot volume.
    5
  • Cost in use
  • Cost relative to real annual volume, not to the trial period pricing.
    7
  • Internal time needed to monitor, correct and evolve the solution once in service.
    5
  • Compliance
  • Where data is processed, by whom, under what contract, and what happens to submitted data.
    8
  • Ability to reconstruct and justify an answer, months later, in front of a third party.
    6
  • Exit, lock-in and sovereignty
  • What you can retrieve, in what format, and how long it takes, on the day you leave.
    5
  • Exit cost designed into the product, closed formats, risk concentrated on a single component with no market equivalent.
    6
  • Law applicable to your data and processing, actual location, and your organisation's ability to continue if the relationship ends.
    5

          

The file is CSV, semicolon separated, UTF-8 with a byte order mark. It opens directly in a spreadsheet configured for European settings.

One caveat

A scorecard without an evaluation set is useless.

This builder gives you the measuring instrument. It does not give you what gets measured. The real work of the method is at step 2: assembling, with your experts, the cases the solutions will be scored on, and fixing the right answer before seeing what the candidates produce.

A scorecard filled in from vendor demos is still a scorecard filled in from vendor demos. It merely looks more serious, which is a drawback.

If you want the part that cannot be downloaded, that is where I come in.

Free to use

This scorecard is yours the moment you download it. Use it, change it, circulate it internally, with no conditions and no attribution required.

If it proved useful, tell me. That is the only feedback I ask for.