The method

Two phases, five steps each, one deliverable per step.

The first phase plays out on the market: it says what the available solutions are worth and which one deserves to go further. The second plays out inside your organisation: it says whether the selected solution holds up in real conditions, and whether you should industrialise it. The order matters, and each step has a gate before the next one.

Phase 1, upstream

What is out there actually worth?

  1. Qualify maturity
  2. Build the evaluation set
  3. Run a level playing field
  4. Iterate across campaigns
  5. Decide which solution

Each new campaign sends step 4 back to step 3, on the same set.

Gate: one solution is selected, or none is.

Phase 2, in house

Should you go ahead, and on what conditions?

  1. Frame the stake and the indicators
  2. Diagnose the real context
  3. Pilot on a bounded scope
  4. Assess against the framing criteria
  5. Decide and document

Three outcomes, of equal weight

  • Industrialise
  • Adjust the scope
  • Stop, and document it
The two phases of the method Phase 1, upstream, five steps: qualify maturity, build the evaluation set, run a level playing field, iterate across campaigns, decide which solution. Step four sends back to step three for the next campaign. A gate separates the two phases: one solution is selected, or none is. Phase 2, in house, five steps: frame, diagnose, pilot, assess, decide. The last step leads to three outcomes of equal weight: industrialise, adjust the scope, or stop and document it. Phase 1, upstream · what is out there actually worth? next campaign 1 Qualify maturity 2 Build the evaluation set 3 Run a level playing field 4 Iterate across campaigns 5 Decide which solution Gate: one solution is selected, or none is Phase 2, in house · should you go ahead, and on what conditions? 1 Frame the stake 2 Diagnose the real context 3 Pilot on a bounded scope 4 Assess against the framing 5 Decide and document Industrialise Adjust the scope Stop, and document it

The three outcomes carry equal weight. A process concluding that the work should stop has rendered the same service as one concluding in favour of deployment, at a fraction of the cost of a failed industrialisation.

Phase 1 Upstream, on the market What is out there worth, and which of these solutions deserves to go further? This phase ends in a selection decision, which may be to select none of them.
1

Qualify maturity before comparing

I use the TRL scale, the nine technology readiness levels, to place a technology and avoid treating a research topic as an available product. Many projects fail because they were premature, not because they were badly run.

Three kinds of maturity are constantly conflated in procurement files, and they need separating: that of the underlying technology, that of the product wrapped around it, and that of your own organisation faced with this use. A mature technology inside a young product sold to an organisation with neither the data nor the roles to exploit it does not add up to a mature project.

This step is short. Its purpose is to rule out early whatever has no business being in the comparison, and to align everyone on the nature of what is being examined.

What it produces

A qualification note: the maturity level retained, what justifies it, and the list of candidates ruled out at this stage.

Gate

The remaining solutions sit in a comparable maturity band. Otherwise you are not comparing, you are illustrating.

Order of duration

A few days, building on what you have already collected.

The mistake to avoid. Taking the maturity of the demo for the maturity of the product. A highly polished demo can rest on unstable research foundations, and the reverse is just as true: a solid product sometimes demos poorly.

2

Build an evaluation set anchored in the business

The set is built with your domain experts, from real cases and above all from the hard and ambiguous ones, never from the examples supplied by the vendor. This is the step that determines the value of everything else: whoever builds the set decides who wins.

A useful set mixes four kinds of case. Everyday cases, in the proportion in which they actually arrive, because a solution that excels on the exotic and is mediocre on the ordinary is a poor solution. Edge cases, sitting on the boundary between two rules. Trap cases, which look like an everyday case and are not. And out of scope cases, where the right answer is to say that the system does not know: a solution that never abstains is a risk, not a help.

Ground truth is established by your experts before they have seen a single answer from a single solution. It is a tedious constraint and it is not negotiable: a reference written after the fact always drifts, unconsciously, towards whichever answer sounded most convincing.

What it produces

The evaluation set, its ground truth, and the note explaining how it was assembled. It belongs to you.

Gate

The set holds enough hard cases to separate the solutions. A set everyone passes decides nothing.

What is needed from you

Domain expert time, the scarcest and most decisive resource in the whole exercise.

The mistake to avoid. Accepting the vendor's test set to save two weeks. Those two weeks are then paid for in years of commitment to a solution chosen for the wrong reasons.

3

Run every solution on an identical protocol

Same inputs, same criteria, several runs. The criteria cover answer quality, functional coverage, robustness on edge cases, cost in use and compliance, to which I add exit cost whenever the commitment is long.

Several runs, because a generative system is not deterministic: the same question asked twice does not produce exactly the same answer. A single run measures luck as much as quality. The spread between two runs of the same solution is itself valuable information: if that spread is of the same order as the gap between two solutions, then the comparison says nothing and the set needs widening.

Scoring is done blind wherever possible, meaning the assessor does not know which solution produced the answer being scored. It is a simple precaution, and it changes the results more often than people expect.

Every score points back to a specific case in the set. A score resting on no case is not a score, it is an opinion.

What it produces

The raw campaign results and the completed scorecard, with the reference case behind each score.

Gate

Gaps between solutions exceed the variability observed between two runs of the same solution.

Watch point

Configuration must be frozen and documented for every candidate, including the front runner.

The mistake to avoid. Letting each vendor tune its configuration between runs. You are then no longer comparing solutions, you are comparing the responsiveness of their pre-sales teams.

4

Iterate across successive campaigns

A single campaign measures a moment, on one version. Several campaigns measure a trajectory, and it is the trajectory that supports a multi-year commitment.

In practice, you rerun the same set at a set interval, on the successive versions of the shortlisted solutions. Three things then appear that a single campaign never shows. The direction and speed of progress: a vendor steadily gaining ground on your cases is often worth more than a vendor who led on the first pass and has stalled since. Regressions: a new version that improves the average may have broken precisely the edge case that matters to you. And the stability of the vendor itself, readable in how it handles the gaps you report.

On fast moving technologies, the gap between two campaigns a few months apart is frequently larger than the gap between two competitors at a given moment. That alone is reason enough never to decide on a snapshot.

What it produces

Results campaign by campaign, the regressions identified, and a protocol that can be rerun.

Gate

The trajectory is established enough to carry a multi-year commitment, or flat enough to justify walking away.

Cadence

Two campaigns are enough to see a direction, more to measure a speed. The interval follows the vendors' release cycle.

The mistake to avoid. Changing the evaluation set between campaigns. The set must stay stable for campaigns to be comparable. New cases go into a clearly identified extension, scored separately.

5

Produce a decision document

A comparative scorecard, a reasoned recommendation and the conditions for success. Not a test report nobody reads.

The test of the document is simple: a committee must be able to decide on that document alone, with no preparatory meeting to decode it. So it states the recommendation first, then the conditions it rests on, and puts the detailed results in an annex for those who want to check.

It also states what was not measured. Every evaluation has blind spots. Naming them is better than letting the reader assume the scope was complete.

Three outcomes are possible at this point, and only one of them leads to the next phase: select a solution and put it to the test in house, widen the search because none is convincing, or walk away because the need does not justify the cost. Walking away here costs the price of an evaluation. Walking away after deployment costs something else entirely.

What it produces

The decision note, the comparative scorecard, the conditions for success, the watch points and the blind spots owned up to.

Gate

The decision can be taken in committee on this document alone, and defended to someone who was not involved.

What stays with you

The set, the scorecard, the protocol. Everything needed to run the next campaign without me.

The mistake to avoid. Rewriting the decision criteria after seeing the results. Criteria are set at the start. If they must change, the change is dated, justified and signed off.

Gate: one solution is selected, or none is.

If none is, the work stops here and it has done its job. If a solution is selected, everything measured so far was measured on cases, in laboratory conditions. The second phase answers the only question left: does this hold up in your organisation, with your processes and your people, and should it be industrialised.

Phase 2 In house, in real conditions Does the selected solution hold up in your organisation, and should it be industrialised? This phase ends in a management decision, with three outcomes of equal weight.
1

Framing

A short note setting out the stake, the scope, and the decision indicators fixed from the start. You move on when the stake is stated in a verifiable way, not merely an interesting one.

The difference fits in one sentence. "Improve case handling with AI" is interesting and cannot be verified. "Reduce first response time on category B cases, measured over the past month, without degrading the rework rate" can be verified, and makes it possible to say at the end whether this was won or lost.

Indicators are fixed now, not at the end. That is the only protection against the bias of discovering, after the fact, that the real value of the project lay elsewhere, precisely where the results happen to be good.

The framing also names what is out of scope. A pilot whose edges are not drawn grows until it can no longer demonstrate anything.

What it produces

A short framing note: the stake, the scope and its edges, the decision indicators, the stakeholders.

Gate

The stake is stated verifiably, and the indicators are accepted by those who will decide at the end.

Watch point

A framing note the decision maker has not read is not a framing note. It needs their signature, literal or otherwise.

The mistake to avoid. Skipping this step because "everyone knows what this is about". That is almost always false, and it surfaces at the moment of concluding, when each party defends whichever indicator suits them.

2

Diagnosis

A shared assessment of the organisation, the data, the technology and the regulatory constraints, then the options and their trade-offs. You move on when the options are down to those that deserve a real test.

Four dimensions, and technology is rarely the blocker. The organisation: who does what today, who will have to work differently, and who has an interest in nothing changing. The data: does it exist, is it accessible, in what state, and at what cost to make it usable. The technology: what integration with the existing estate actually requires, and what it costs. Regulatory constraints: what the nature of the processing imposes, and what your sector adds on top.

The important word is shared. A diagnosis the business disputes is useless, even when it is correct. It is built with the people whose work it describes.

By the end, the options are few and each carries an explicit trade-off. An option with no downside is an option that has not been properly examined.

What it produces

The assessment across the four dimensions, the shortlisted options and the trade-off each one carries.

Gate

Options are down to those worth a real test, and the business recognises its own situation in the diagnosis.

What is needed from you

Access to the teams concerned, including those who did not ask for this project.

The mistake to avoid. Leaving the data question until last. It is the constraint that shifts the most timelines, and it is always discovered too late when it has not been faced squarely at the start.

3

Targeted pilot

A precise hypothesis, a limited scope, a fixed deadline, an explicit test protocol. You move on when observable results exist, favourable or not.

Targeted is the word carrying the weight. A pilot set up to demonstrate that "it works" demonstrates nothing, because it has no failure condition. A pilot testing a hypothesis stated in advance, on a scope where measurement is possible, produces a result either way.

The deadline is fixed at the outset, and it is short. A pilot that drags on becomes a de facto service: users get used to it, stopping it becomes politically expensive, and the decision then makes itself without ever having been taken.

The protocol states who uses what, for how long, what is measured and how. It also states what happens in the event of an incident, because these are real users on real cases.

What it produces

The pilot protocol, then the observed results and the qualitative feedback from actual users.

Gate

Observable results exist, favourable or not. An unfavourable result is a result.

Watch point

Plan the exit from the pilot at the start, including the return to the previous way of working.

The mistake to avoid. Widening the scope mid-course because early feedback is good. You then lose comparability, and the final measurement no longer bears on what was framed.

4

Measured assessment

The scorecard is completed on feasibility, real business impact, maturity, regulatory and lock-in risk, then held against the framing criteria, without rewriting them after the fact.

Real business impact is what can be observed, not what can be inferred. Time saved per case only becomes a gain on two conditions: that the freed time goes to something else useful, and that the workload has not simply moved elsewhere. Both are checked with the people who do the work.

This is also where the internal cost of the pilot is measured, in team time and in management attention. That cost is real and it weighs on the industrialisation decision, since it will recur at greater scale.

Holding the results against the framing criteria is the moment of truth for the phase. You take the framing note, read it exactly as it was written, and answer point by point.

What it produces

The completed scorecard and a line by line comparison against the indicators set at framing.

Gate

Every framing indicator has an answer: met, not met, or not measurable, with what prevented measurement.

Rule

Framing criteria are not rewritten. If they must evolve, the change is dated and visible.

The mistake to avoid. Swapping the framing indicators for the ones on which results happen to be good. It is the most common bias, the most understandable, and the most destructive to the credibility of the whole exercise.

5

Decision: industrialise, adjust or stop

A reasoned note, worth just as much when it concludes that the work should stop. A documented stop is a management decision, not a failure to be hidden: that is what separates a consulting engagement from a showcase of successes.

The three outcomes carry equal weight and are stated as such from the framing. Industrialise: the note then says on what conditions, with what monitoring, and what will need rechecking in six months. Adjust the scope: the value is there, but not where it was being looked for, or not yet at this scale. Stop: the note says why, what the exercise taught, and on what conditions the subject would deserve reopening.

That last part is the one people skip, and it is the most useful. A technology ruled out today for lack of maturity may become relevant in eighteen months. Writing down the reopening conditions saves having to reinvestigate everything, and equally stops the subject coming back too soon.

Whatever the outcome, what has been built stays with you: the evaluation set, the scorecard, the pilot protocol and the indicators. That is the real asset of the exercise, the one that serves the next decision.

What it produces

The decision note, the conditions for success or for reopening, and the monitoring arrangements if you industrialise.

Test of success

The decision is taken, it is written down, and it is defensible to someone who was not involved.

What stays with you

The whole apparatus, reusable on the next subject, without me.

The mistake to avoid. Not deciding. A pilot that drags on without a conclusion becomes an unacknowledged service, with no operating budget, no owner and no service level. It is the most expensive of the three outcomes, and the only one nobody chose.

What the method assumes

  • That your domain experts are available, a few days spread across the period.
  • That you can extract real cases, anonymised or reconstructed if need be.
  • That vendors accept a protocol they do not control. Their reaction to that request is already information.
  • That a decision maker accepts in advance that stopping is an acceptable outcome.

Next

The scorecard is the instrument of step 3 in the first phase, and of step 4 in the second. You can build it right now, without providing anything.

Build the scorecard

The question is where you stand in that sequence.

Some arrive before the first phase, others halfway through the second with a pilot that has been running far too long. Thirty minutes are enough to place your starting point and to say what is missing before a decision can be made.