An AI automation trial can draft every reply correctly and still leave you with more work. If nobody owns the review queue before the tool goes live, the subscription price is only part of what you’re committing to.
Writer Elena Brooks built a five-row readiness card for founders and operators to work through before they even look at tools. The argument is simple: settle these five conditions first. Passing the card means you’re ready to run a bounded pilot. It does not mean AI will save you money.
Gate 1: A Process Stable Enough to Observe
Choose a slice you can describe without naming software. Brooks uses a hypothetical online shop routing refund requests. The starting point is an incoming request. The endpoint is a proposed route and reason for a reviewer. No money moved. No reply sent.
Her suggested baseline: record every request over a normal two-week window using one version of the policy. Keep timestamps, the route chosen, missing information, and policy disputes, including requests outside the proposed pilot scope.
Two weeks is an illustrative starting point, not a guarantee. A quiet fortnight could miss the problems that arrive during a sale. The practical gate is narrower: can the owner apply the same written eligibility and completion rules to the sample? If those rules change daily, fix the boundary or collect another window before comparing results.

Gate 2: A Named Owner With Review Capacity
An owner needs time and authority to intervene, not just a title. For the hypothetical shop, that means writing down who checks proposed routes, when they review the queue, and who covers an absence. A name without review capacity won’t clear a backlog.
Brooks references the NIST AI Risk Management Framework Core, specifically GOVERN 2.1 on documented roles and MAP 3.5 on defined human oversight. She frames the practical interpretation plainly: don’t send live work into a queue nobody has agreed to watch.
Gate 3: A Mapped Exception Path
Before connecting live inputs, every exception needs a destination. In the refund routing example, a missing order number goes to the reviewer. So does a request that conflicts with the written policy. The exception log records the reason, recipient, and resolution for each case.
Brooks recommends testing the handoff with a sample case before going live. Confirm the reviewer receives it and can complete the work manually. A notification alone doesn’t show the handoff works.
Unknown cases belong in that same review path. You won’t enumerate every exception beforehand, but every unrecognized case needs a safe destination. Pause if the reviewer can’t keep up.
Gate 4: A Captured Baseline
Measure through review and correction, not just processing speed. For each sampled request, record active handling minutes through completion. Also record elapsed time from arrival to completion separately, so waiting in a queue doesn’t disappear inside a speed claim.
Brooks pairs two metrics for refund routing:
- Median active minutes: time spent actually handling a request
- Correction rate: requests needing correction divided by all completed requests
Apply the same definitions before and during the pilot, including human checking time. She references the Institute for Healthcare Improvement’s guidance on establishing measures, which distinguishes outcome, process, and balancing measures. Balancing measures check whether a change improves one area while creating problems elsewhere.
Count eligible requests against all incoming requests too. If a trial accepts only straightforward cases, compare it with that same slice of the baseline. Faster handling of easy requests doesn’t establish savings across the whole inbox. Report the count beside any percentage, especially when the sample is small.
I wouldn’t approve recurring software spend from a pilot that leaves its cleanup work uncounted.
Gate 5: Written Success and Stop Conditions
Write the pass and stop rules before seeing results. Brooks gives hypothetical test settings for the refund routing example:
- Boundary: 30 eligible requests or one week, whichever comes first
- A reviewer checks every proposed route
- Illustrative success target: a 20% reduction in median active handling time with no increase in the correction rate
- Immediate stop: the trial sends an unauthorized message or changes a payment
Set a spending cap too. If the trial ends before the agreed minimum sample arrives, label the comparison inconclusive. NIST’s MANAGE 2.4 addresses responsibilities for disengaging systems whose performance or outcomes conflict with intended use.

The Honest Limit of This Framework
Brooks acknowledges the fair objection: trying automation can expose exceptions you couldn’t predict. The IHI guidance on testing changes supports small initial tests, recording unexpected observations, and using what happens to plan another test. You don’t need a finished process map to explore.
Her recommendation for unclear boundaries: use synthetic or appropriately de-identified cases. Verify data permissions before anything goes to a provider.
And even a pilot that hits its target has limits. Passing the success condition justifies another bounded test. It doesn’t establish causation, annual savings, or permission to remove the reviewer.
When This Works
This framework fits best when you have a repeatable process you can observe over a normal operating window, a person who can own the review queue before a single live input goes through, and clear rules for what counts as success and what forces a stop.
When It Does Not
If your process rules change frequently, if no one has agreed to watch the queue, or if your baseline data doesn’t separate active handling time from queue wait time, the card will flag a no-go. Brooks is direct on this: any missing row is a no-go for live inputs. Tool suitability and data-security approval remain separate checks on top of all five gates.


