Two weeks is a forcing function
A fortnight is too short to build everything, which is the point: it forces the scope down to one workflow and one question. Longer pilots expand to fill the time and drift into the pilot that never ships.
It also keeps the decision cheap. A two-week commitment can be approved by one person and abandoned without embarrassment, which is precisely what makes an honest negative result possible.
Before day 1: pick the right candidate
The workflow has to be high-volume enough to matter, verifiable in seconds, and low-stakes enough that a wrong answer is recoverable. Choosing badly is the only failure this format cannot absorb — run the pre-project questions first if there is any doubt.
Day 1: write down what would make this a yes
Agree the specific result that means proceed and the one that means stop, before anything is built. Without both, the pilot cannot conclude, only continue — and this is the same artefact as an evaluation written first.
Write the stop condition as carefully as the go condition. “If more than a handful of these twenty need substantial correction, we stop” is a decision made in advance by people who are not yet invested in the answer, which is the only time it can be made cleanly.
Days 2–3: collect twenty real inputs
Real tickets, real documents, real records — including the messy ones. Curated examples answer a question you don’t have. This set is both your test data and, later, your regression suite.
Sample deliberately rather than taking the most recent twenty: a few typical cases, a few edge cases your team argues about, one or two that went wrong historically. If a category of input is missing here, it is missing from every conclusion you draw.
A pilot that can’t fail hasn’t been designed. It’s been budgeted.
Days 4–8: build the thinnest version
One workflow, no interface work beyond what is needed to see the output, no integrations that aren’t essential. You are testing whether the judgment step works, not building the product.
Concretely: a script, a spreadsheet of outputs, and whatever retrieval is genuinely required. Anything spent on interface at this stage is spent on a thing you may throw away, and it makes the eventual decision harder because somebody now has something to defend.
Days 9–11: have a human score it
Someone who does the job reviews every output against the criteria from day one. Their disagreements are the finding — they tell you whether the gap is capability, knowledge, or an unclear rule.
That distinction determines what happens next, and the three have completely different costs. A capability gap may mean a different model. A knowledge gap means writing documentation. An unclear rule means the business has never decided something — the discovery that makes the exercise valuable even when the automation is abandoned.
Days 12–14: decide, and write down why
Proceed to a narrow production release, iterate once more on a specific known gap, or stop. Recording the reasoning matters even for a stop, because the same idea returns in six months and the analysis is reusable.
If the answer is proceed, the next step is production with one team on real traffic — not a bigger pilot. Scope grows after corrections have flattened, not before.
What the two weeks leaves behind
Even a stopped pilot produces four durable things: twenty labelled real inputs, a scored evaluation set, a written record of which rules were undocumented, and a clear statement of the conditions under which this becomes viable.
That is why the format is worth running even where you suspect the answer is no. The alternative — an open-ended investigation with no stopping rule — costs more and produces an opinion instead of an artefact.