Service Robot Pilot Exit Criteria: Scoring Go or No-Go Before You Commit
At a glance: A pilot that ends with a positive feeling and no written decision has produced a memory, not evidence. The fix is structural: agree the thresholds before the first unit arrives. This guide sets out six categories of exit criteria, how to capture the baseline they are compared against, and the clauses that turn a passing pilot into a contract.
Most Pilots End Without Ever Being Scored
A service robot pilot that finishes with a positive feeling and no written decision has not produced evidence; it has produced a memory. The failure is structural rather than cultural. Teams write pilot plans that describe what will be done, and rarely write the thresholds at which the answer becomes yes or no. Without pre-agreed thresholds, every pilot conclusion becomes an argument about interpretation, and the argument is usually won by whoever is more senior rather than by the data.
Exit criteria fix this by moving the decision from the end of the pilot to the start. They are written before the first unit is delivered, signed by both sides, and they name the exact numbers that constitute pass and fail.
What Separates an Experiment From a Demonstration

There are two activities that are both called pilots and they produce opposite results.
A demonstration shows that a robot can clean a floor when conditions are favourable. Vendors run demonstrations because they reliably succeed. A real pilot tests whether the fleet performs in your building, on your hardest surface, during your worst shift, with your staff operating it after two days of training. Demonstrations are marketing; only the second produces evidence that survives contact with a finance committee.
The practical difference is in the conditions. A pilot is genuine if it runs unassisted by vendor engineers for at least three consecutive operating days, includes the highest-task-hour surface type, and runs during a peak constraint window at least twice. If all three are missing, the exercise is a demonstration with a longer invoice.
Six Categories of Exit Criteria and the Thresholds That Matter

Exit criteria should cover six categories. Covering fewer leaves a gap that appears later as an unresolvable dispute.
| Category | Metric | Typical pass threshold | Why it is included |
|---|---|---|---|
| Productive performance | Coverage rate on named surface | Within 10% of contracted rate | Drives fleet size and payback |
| Quality | First-pass clean rate, audited | ≥ 90% of inspected zones | Prevents "clean but you clean it again" |
| Reliability | Availability across the pilot window | ≥ 88% of scheduled hours | Determines real capacity, not datasheet capacity |
| Autonomy | Interventions per shift | ≤ 2 per unit per shift | Hidden labour cost of supervision |
| Safety and compliance | Reportable incidents, near misses | Zero reportable | Non-negotiable in occupied buildings |
| Operational integration | Task completion without manual scheduling | ≥ 80% of tasks | Tests the software, not just the hardware |
Two of these deserve scrutiny because they are the ones most often written loosely.
Interventions per shift is the real autonomy metric
A robot that completes its route but needs a staff member to clear it three times a shift has transferred labour rather than reduced it. The intervention count should be logged with a reason code, and the exit criterion should apply to the total, not just to faults. Interventions caused by the environment, such as chairs left in a path or a door propped open, are as much a finding as interventions caused by the machine, because in production they will recur daily.
First-pass clean rate must be independently audited
Quality assessed by the party that supplied the equipment is not an audit. The pilot budget should include a small independent inspection allocation, even if it is only a supervisor with a defined checklist sampling a fixed percentage of zones. The threshold should be set against the existing manual cleaning standard, not against an idealised one, so the comparison is like for like.
Writing the Baseline Before Day One
Exit criteria comparing robot performance to nothing are meaningless. Every metric above needs a manual baseline captured before the robots arrive, using the same measurement method and ideally the same auditor. The baseline is normally collected over one to two weeks on the same surfaces.
Three baseline numbers carry disproportionate weight. The current labour hours per surface per week establishes the denominator for any savings claim. The current first-pass clean rate, as judged by the same checklist, establishes the quality bar. The current incidence of reactive, unplanned work establishes how much of the day is already spent firefighting, which is the most common reason a robot productivity figure looks disappointing in week one.
Capturing a baseline is unglamorous and frequently skipped. When it is skipped, the savings case is later built on the vendor's assumptions instead of the facility's history, and the facility has no defence when finance asks how the number was derived.
The Three Failure Modes of Pilot Governance
Pilots fail in three recurring ways, all of them governance rather than technology problems.
- The criteria are written late. Thresholds agreed in week six are shaped by what the pilot turned out to show. Thresholds agreed before delivery are honest. The date of agreement should be recorded.
- No single decision owner. A pilot steered by a committee with operations, procurement, IT and finance all represented produces a conclusion of "interesting, let's run another pilot". One named owner with authority to recommend go or no-go, and a decision date, produces a decision.
- The pass threshold has no cost attached. Criteria should state what happens on pass and on fail: units deployed, order placed, or exit with data shared. Criteria without consequences are surveys.
Structuring the Commercial Terms Around the Criteria

The strongest pilot agreements convert exit criteria directly into commercial terms, so that the pilot is not a separate exercise but the first phase of the contract. Three clauses do most of the work.
First, a price hold that fixes unit and service pricing for a defined period after a passing pilot, so the vendor cannot reprice once it knows the facility is committed. Second, a measured-performance clause that carries the coverage rate and availability figure achieved in the pilot into the operating contract, with a defined remedy if they are not maintained. Third, a data ownership clause confirming the facility receives the pilot's raw logs, including intervention records and fault codes, whether the decision is go or no-go.
That third clause matters more than it appears. Raw logs are the input to any later comparison of vendors, and a facility that has to start its measurement from zero with the next supplier has effectively subsidised the first one's learning.
What Good Looks Like at the End of the Pilot

A well-run pilot concludes with a one-page scorecard listing each criterion, the threshold, the measured value and a pass or fail. It includes the intervention log with reason codes, the availability calculation with its scheduled-hours denominator, and the audit sample size. It names the decision owner and the decision date.
The document is short because the criteria were agreed in advance, which is the point. The value of exit criteria is not that they make a pilot rigorous in the abstract; it is that they make the final conversation a comparison against numbers rather than a negotiation about impressions. Facilities that write them once tend to write them for every subsequent technology evaluation, because the difference in decision quality is immediate and obvious.
