Service Robot Pilot Exit Criteria: Scoring Go or No-Go Before You Commit

At a glance: A pilot that ends with a positive feeling and no written decision has produced a memory, not evidence. The fix is structural: agree the thresholds before the first unit arrives. This guide sets out six categories of exit criteria, how to capture the baseline they are compared against, and the clauses that turn a passing pilot into a contract.

Photorealistic view of a white autonomous cleaning robot alone on a polished office floor during a quiet evaluation run, instrumented floor markings, no people and no text

Most Pilots End Without Ever Being Scored

A service robot pilot that finishes with a positive feeling and no written decision has not produced evidence; it has produced a memory. The failure is structural rather than cultural. Teams write pilot plans that describe what will be done, and rarely write the thresholds at which the answer becomes yes or no. Without pre-agreed thresholds, every pilot conclusion becomes an argument about interpretation, and the argument is usually won by whoever is more senior rather than by the data.

Exit criteria fix this by moving the decision from the end of the pilot to the start. They are written before the first unit is delivered, signed by both sides, and they name the exact numbers that constitute pass and fail.

What Separates an Experiment From a Demonstration

Photorealistic photograph of a single white autonomous cleaning robot operating alone during an overnight trial on a polished office floor with visible route markings, no people and no text

There are two activities that are both called pilots and they produce opposite results.

A demonstration shows that a robot can clean a floor when conditions are favourable. Vendors run demonstrations because they reliably succeed. A real pilot tests whether the fleet performs in your building, on your hardest surface, during your worst shift, with your staff operating it after two days of training. Demonstrations are marketing; only the second produces evidence that survives contact with a finance committee.

The practical difference is in the conditions. A pilot is genuine if it runs unassisted by vendor engineers for at least three consecutive operating days, includes the highest-task-hour surface type, and runs during a peak constraint window at least twice. If all three are missing, the exercise is a demonstration with a longer invoice.

Six Categories of Exit Criteria and the Thresholds That Matter

Photorealistic overhead photograph of a commercial hard floor after cleaning, uniformly reflective surface with clean edge lines along the wall, no people and no text

Exit criteria should cover six categories. Covering fewer leaves a gap that appears later as an unresolvable dispute.

CategoryMetricTypical pass thresholdWhy it is included
Productive performanceCoverage rate on named surfaceWithin 10% of contracted rateDrives fleet size and payback
QualityFirst-pass clean rate, audited≥ 90% of inspected zonesPrevents "clean but you clean it again"
ReliabilityAvailability across the pilot window≥ 88% of scheduled hoursDetermines real capacity, not datasheet capacity
AutonomyInterventions per shift≤ 2 per unit per shiftHidden labour cost of supervision
Safety and complianceReportable incidents, near missesZero reportableNon-negotiable in occupied buildings
Operational integrationTask completion without manual scheduling≥ 80% of tasksTests the software, not just the hardware

Two of these deserve scrutiny because they are the ones most often written loosely.

Interventions per shift is the real autonomy metric

A robot that completes its route but needs a staff member to clear it three times a shift has transferred labour rather than reduced it. The intervention count should be logged with a reason code, and the exit criterion should apply to the total, not just to faults. Interventions caused by the environment, such as chairs left in a path or a door propped open, are as much a finding as interventions caused by the machine, because in production they will recur daily.

First-pass clean rate must be independently audited

Quality assessed by the party that supplied the equipment is not an audit. The pilot budget should include a small independent inspection allocation, even if it is only a supervisor with a defined checklist sampling a fixed percentage of zones. The threshold should be set against the existing manual cleaning standard, not against an idealised one, so the comparison is like for like.

Writing the Baseline Before Day One

Exit criteria comparing robot performance to nothing are meaningless. Every metric above needs a manual baseline captured before the robots arrive, using the same measurement method and ideally the same auditor. The baseline is normally collected over one to two weeks on the same surfaces.

Three baseline numbers carry disproportionate weight. The current labour hours per surface per week establishes the denominator for any savings claim. The current first-pass clean rate, as judged by the same checklist, establishes the quality bar. The current incidence of reactive, unplanned work establishes how much of the day is already spent firefighting, which is the most common reason a robot productivity figure looks disappointing in week one.

Capturing a baseline is unglamorous and frequently skipped. When it is skipped, the savings case is later built on the vendor's assumptions instead of the facility's history, and the facility has no defence when finance asks how the number was derived.

The Three Failure Modes of Pilot Governance

Pilots fail in three recurring ways, all of them governance rather than technology problems.

Structuring the Commercial Terms Around the Criteria

Photorealistic photograph of a wall-mounted charging dock for service robots in a service corridor with status indicator lights, no people and no text

The strongest pilot agreements convert exit criteria directly into commercial terms, so that the pilot is not a separate exercise but the first phase of the contract. Three clauses do most of the work.

First, a price hold that fixes unit and service pricing for a defined period after a passing pilot, so the vendor cannot reprice once it knows the facility is committed. Second, a measured-performance clause that carries the coverage rate and availability figure achieved in the pilot into the operating contract, with a defined remedy if they are not maintained. Third, a data ownership clause confirming the facility receives the pilot's raw logs, including intervention records and fault codes, whether the decision is go or no-go.

That third clause matters more than it appears. Raw logs are the input to any later comparison of vendors, and a facility that has to start its measurement from zero with the next supplier has effectively subsidised the first one's learning.

What Good Looks Like at the End of the Pilot

Photorealistic photograph of a clean empty modern facility management office desk with a closed notebook and pen under soft daylight, no people, no screens, no text

A well-run pilot concludes with a one-page scorecard listing each criterion, the threshold, the measured value and a pass or fail. It includes the intervention log with reason codes, the availability calculation with its scheduled-hours denominator, and the audit sample size. It names the decision owner and the decision date.

The document is short because the criteria were agreed in advance, which is the point. The value of exit criteria is not that they make a pilot rigorous in the abstract; it is that they make the final conversation a comparison against numbers rather than a negotiation about impressions. Facilities that write them once tend to write them for every subsequent technology evaluation, because the difference in decision quality is immediate and obvious.

Products