I

AI in practice

How to run an AI pilot a firm can actually judge

One domain, two people, three to five weeks, and a pass mark agreed before anything is built.

The short answer

A pilot a firm can judge has five properties: it covers one work domain, involves two people, runs three to five weeks, targets a deliverable the firm already produces by hand, and is measured against a pass mark agreed before anything is built. Everything else is a demonstration, and demonstrations do not change how a firm works.

Why most pilots cannot be judged

Most AI pilots in professional firms fail quietly rather than loudly. A partner sees a tool, a few people try it, and three months later nobody can say whether it worked. The problem is rarely the model. It is that nobody wrote down what "worked" would mean.

Broad pilots make this worse. A pilot that touches bids, archive search and project tracking at once produces three half-results and no clear decision. The organisation loses patience before anything ships.

Five design constraints

One domain, not a cross-cutting process. Broad workflows are hard to implement and harder still to assess afterwards.

Two people, weeks not months. Enough to be real, small enough to change course cheaply.

The output already exists. Pick a deliverable the firm already produces manually, so "better" is measured against something real, not against a vendor benchmark.

Start where use is zero. Where nobody uses AI yet, every improvement is visible immediately, and visibility buys permission for the harder work.

Learn from artefacts, not interviews. Busy experts struggle to explain their tacit knowledge. Past inputs and the outputs they produced teach a system far more, and the experts only have to approve the structure it proposes.

Five measures, agreed in advance

Cycle time: hours recorded before the pilot, then measured again on two real cases. A reasonable target is a 40% reduction.

Output accuracy: tested on material the firm has already processed, so the right answer is known. Above 90% of material items identified is a fair bar.

Actual use: how many participants open the tool in a given week without being prompted. Above 70% means it has become part of the work.

Load displacement: how much repetitive work the team can now absorb without adding headcount.

New capability: how many actions that were not possible before were actually performed, counted against a list fixed during discovery. This is the only measure that is not about efficiency, and it is often the one that matters most commercially.

The stop rule

If the pilot misses its targets, the answer is to change the tool or the approach, not to expand. Writing that rule down before the pilot starts is what makes the measures real rather than decorative. It is also the whole reason to start small: a mistake costs weeks, not quarters.

What it costs the firm

Separate two kinds of time. One internal lead needs four to six hours a week, treated as a project in its own right. Each pilot participant needs two to three hours a week, mostly reviewing outputs, and only for the pilot period.

Clients usually quote a single number for "what this will cost us". Splitting it makes the ask smaller and far easier to approve, because only one person is asked for a standing commitment.

In practice

We used exactly this design with a 25-person architecture practice. The first pilot was administration and bids: repetitive, hard to recruit for, and with zero AI use at the start. The full method, from discovery to the 24-month economics, is in the case study.

FROM NOTE TO VENTURE

What physical system are you seeing?

Share the opportunity