Rank by verifiability, not by ambition

The right first job is one where a human can check the output in seconds and a mistake costs nothing. That single criterion produces a better ordering than any assessment of what the technology can theoretically do.

It also produces an order most people find anticlimactic, because the jobs that satisfy it are the ones nobody wanted to talk about in the meeting. That is the point: the first deployment is buying trust and documentation, not impressiveness.

1. Summarising what already happened

Call notes, long email threads, handover summaries, weekly digests. The source material is right there, so correctness is obvious at a glance, and the time saved is immediate. Nobody’s job depends on it being perfect, which makes it the safest place to build trust.

It also produces the first real lesson cheaply: summaries reveal immediately whether the system understands your vocabulary, your product names and your shorthand. Those gaps are trivial to fix here and expensive to discover in front of a customer.

2. Sorting and routing

Classifying inbound work and sending it to the right place is high-volume, invisible to customers, and self-correcting — every mis-route a human fixes is a labelled example. This is the same reasoning behind putting triage ahead of deflection in a support workflow.

Measurement is unusually clean here too, because your team is already routing this work and you can compare directly. Few AI deployments offer a like-for-like baseline that easily.

Give it the job where being wrong is noticed immediately and costs nothing. Trust is built there, not claimed.

3. Drafting the reply a human sends

Draft-and-edit captures most of the speed benefit while a person stays accountable for what goes out. The edit distance between draft and sent message is also a free, continuous quality metric — the closest thing to an honest measure of whether it earns its seat.

Watch for the failure mode specific to this one: if agents routinely delete the draft and start again, the feature is costing time rather than saving it. That shows up in the edit distance long before anyone complains.

Why this order and not another

Each job produces the input the next one needs. Summarising establishes vocabulary. Routing produces labelled decisions. Drafting produces a corpus of approved replies. By the time you consider anything customer-facing, you have months of corrections and a documented set of rules rather than a vendor demo.

What comes later, not first

Autonomous customer conversations, anything touching money, and anything irreversible belong after the first three are stable — by which point you have documented the tribal knowledge the system was missing and know where it needs a supported way to escalate.

Teams that invert this order don’t move faster. They spend the same months, just with an audience — and a public failure sets the programme back further than the sequencing would have cost.

How long each takes to settle

Expect corrections to be frequent in the first weeks and to flatten as the missing knowledge gets written down. The signal to move to the next job is that flattening, not a date — which is the same trigger as expanding scope generally.