Two phases, five steps each

Evaluate an AI solution before you buy it, then before you deploy it.

Vendor demos all look alike and they always go well. They say nothing about what matters: how the solution behaves on your data, your edge cases and your volumes. Methodia is the method I apply to decide on observed facts rather than on a favourable impression, first on the market, then inside your organisation.

Transferable to any sector and any type of solution. When the engagement ends, the scorecard and the evaluation set stay with you.

The problem

Buying on the strength of a demo is buying blind.

A demo is a prepared object. It runs on examples chosen by the seller, in an order chosen by the seller, on a version the seller controls. It proves the solution can do what it was taught to show. It says nothing about what comes next.

The cost of the mistake only surfaces once deployment is under way: when real cases arrive, when volumes rise, when an ambiguous case lands on a caseworker who has to make a decision. By then the contract is signed and the programme has started.

Then comes the second trap. The solution has been chosen, a pilot starts, early feedback is encouraging, and nobody can any longer say under what condition it would be stopped. The pilot settles in, becomes a de facto service with no budget and no owner, and the decision ends up making itself without ever having been made.

Neither difficulty is technical. Both are methodological. Nobody is short of tools for testing. What is missing is a protocol that makes results comparable and defensible, and decision criteria set before anyone has seen the results.

The three most common traps

  • The vendor's own test set. It is calibrated on what the solution does well. Whoever builds the set decides who wins.
  • The single run. A generative system does not answer the same way twice. A good answer observed once is not a good answer.
  • Criteria rewritten after the fact. At the end, the real value of the project turns out to lie exactly where the results happen to be strong.
The method

Two phases. The first says what to choose, the second says whether to go ahead.

Every step has a deliverable and a gate. You do not move to the next step because the calendar says so, you move because the gate is met.

Phase 1 Upstream, on the market What is out there actually worth, and which of these solutions deserves to go further?
1

Qualify maturity before comparing

The TRL scale, to place a technology and avoid treating a research topic as an available product. Many projects fail because they were premature, not because they were badly run.

2

Build an evaluation set anchored in the business

The set is built with your domain experts, from real cases and above all from the hard and ambiguous ones, never from the examples supplied by the vendor.

3

Run every solution on an identical protocol

Same inputs, same criteria, several runs. Criteria cover answer quality, functional coverage, robustness on edge cases, cost in use and compliance.

4

Iterate across successive campaigns

A single campaign measures a moment, on one version. Several campaigns measure a trajectory, and it is the trajectory that supports a multi-year commitment.

5

Produce a decision document

A comparative scorecard, a reasoned recommendation and the conditions for success. Not a test report nobody reads.

Gate: one solution is selected, or none is.

If none is, the work stops here and it has done its job, for the price of an evaluation.

Phase 2 In house, in real conditions Does the selected solution hold up in your organisation, and should it be industrialised?
1

Framing

A short note setting out the stake, the scope and the decision indicators, fixed from the start. You move on when the stake is stated in a verifiable way, not merely an interesting one.

2

Diagnosis

A shared assessment of the organisation, the data, the technology and the regulatory constraints. You move on when the options are down to those that deserve a real test.

3

Targeted pilot

A precise hypothesis, a limited scope, a fixed deadline, an explicit protocol. You move on when observable results exist, favourable or not.

4

Measured assessment

A scorecard covering feasibility, real business impact, maturity and risk, held against the framing criteria, without rewriting them after the fact.

5

Decision: industrialise, adjust or stop

A reasoned note, worth just as much when it concludes that the work should stop. A documented stop is a management decision, not a failure to be hidden.

What you are left with

  • The evaluation set
    Real cases, with ground truth established by your own experts.
  • The completed scorecard
    Criteria, weights, scores, and the case behind every score.
  • The protocol
    Everything needed to rerun the campaign next year, with or without me.
  • Two decision notes
    Which solution to select, then whether to industrialise it and on what conditions.
The method transfers to any sector and any type of solution. I leave you the scorecard so you can run your next evaluations without me.
Track record

This method was not written for a website. It was built by practising it.

Applied to very different subjects inside an IT department of 1,500 people, with the business teams concerned.

5

successive campaigns evaluating large language models for claims handling support, on a set of 129 questions

9

speech recognition solutions compared on an identical protocol

4

low code platforms compared before commitment

TRL

a sign language technology placed on the maturity scale before any comparison

Above all, this work served to rule out solutions that were compelling in a demo, and to secure those that did make it into production. An evaluation concluding that the work should stop is a management decision, not a failure.

The scorecard

A tool, not a brochure.

The promise of the method is that you can rerun it without me. So let us start now: the next page holds the scorecard builder. You pick the families of criteria, adjust the weights, name the candidate solutions, and leave with a file ready to fill in.

Everything happens in your browser. No data is transmitted, no sign-up is required, no account is created.

The six families of criteria

  • Answer quality
  • Functional coverage
  • Robustness on edge cases
  • Cost in use
  • Compliance
  • Exit, lock-in and sovereignty
The position

An evaluator who sells the solution being evaluated is not evaluating anything.

That is the whole point of this page. The only interest I have in the outcome of your evaluation is that it be correct.

What I do not do

  • No development
  • No solution reselling
  • No vendor partnership, no referral fee
  • No automatic renewal
  • No publication of your results, in any form

A decision to make, and a demo that went well.

A first thirty minute call is enough to say whether the method applies to your case, on which phase, and within what timeframe. If there is nothing there, I say so on that call.