Weighted Procurement Scoring for Service Robots

At a glance: By the time proposals are on the table, every shortlisted vendor has a plausible story. The difference between a good fleet and an expensive mistake is rarely the hardware, it is the decision method. A weighted scoring matrix converts opinion into arithmetic: each criterion is scored on a fixed scale, multiplied by a weight agreed in advance, and totalled. The result is a ranked shortlist you can defend to a board, to an auditor and to the losing bidder.

Photorealistic cover image for the article, no people faces and no text

Why Unweighted Evaluation Fails in Robot Procurement

Most buying committees evaluate robot vendors on a single axis: price, or price plus a gut feeling about the demo. That works when the product is a commodity. A service robot fleet is not a commodity. It is a multi-year operating commitment whose true cost is dominated by uptime, service response, integration effort and the replacement cycle, none of which are visible in the purchase price.

The failure mode is predictable. The vendor with the slickest demonstration and the lowest headline price wins, and eighteen months later the buyer is absorbing downtime, paying for integration the proposal implied was free, and discovering the battery pack is a proprietary part with a six-week lead time. An unweighted decision has no place to record those risks before they become costs.

A weighted matrix does not remove judgement. It forces judgement into the open, where each assumption has a number and each number can be challenged. The output is not a proof, but it is auditable, and auditability is what protects the decision when the fleet underperforms and someone asks how the vendor was chosen.

The Seven Criteria That Carry the Decision

Boil the evaluation down to seven criteria. More than that and the committee loses the thread; fewer and you omit a deal-breaker. The table below gives the standard set, the evidence to request for each, and a starting weight. Adjust the weights to your building type, but fix them before proposals are opened.

CriterionWhat it measuresEvidence to requestWeight
Total cost of ownership5-year cost including service, parts, powerItemised TCO model, not a price sheet25%
Technical fitCoverage, navigation, payload against your scenesMeasured coverage rate on comparable floors20%
Service and uptimeResponse time, spares availability, first-time fixSLA with credits, named service region20%
Integration capabilityAPI, BMS/CMMS hooks, data exportDocumented API, sample export file10%
Vendor stabilityFinancial health, install base, referencesTwo reference sites you can visit10%
Safety and complianceCertification against the standards you must meetTest certificates, not marketing claims10%
Contract termsData ownership, exit, escalators, remediesDraft contract, not a proposal annex5%

Notice where the weight sits. TCO, technical fit and service together carry 65%. Price alone is a component of TCO, not a criterion in its own right, which is precisely the re-framing that stops the cheapest bid winning by default. The commercial depth of that cost model is set out in the five-year cost of ownership method, and the arithmetic behind the technical-fit scores is covered in the throughput planning guide.

Photorealistic photograph of a bright empty meeting room with a large printed comparison table pinned to the wall and a stack of vendor proposal folders on the table, no people faces and no text

Scoring Each Criterion Without Fooling Yourself

Use a fixed 1-to-5 scale for every criterion and define what each point means in writing. The scale is the guard against grade inflation, where every criterion drifts to 4 or 5 and the matrix stops discriminating.

ScoreMeaningUse when
5Exceeds requirement, documentedEvidence proves it beats your spec
4Meets requirement fullyEvidence proves the spec is met
3Meets with caveatsYes, but with a stated condition
2Partially meetsGap you must absorb or pay to close
1Does not meetUnresolved or unproven

Score on evidence, not on the response document. A vendor who claims a 98% uptime figure but supplies no measurement method scores a 3, not a 5. A vendor who supplies the last twelve months of fleet telemetry scores a 5 even if the number is 96%, because it is verifiable. The rule is simple: no evidence, no top score. Apply the procurement-adjacent due diligence in the vendor evaluation framework as the evidence trail that feeds these scores.

Normalising and Applying the Weights

Once every criterion is scored 1-5, multiply by the weight and total. The arithmetic is deliberately simple so a committee member can reproduce it by hand.

CriterionWeightVendor A scoreA weightedVendor B scoreB weighted
Total cost of ownership25%51.2530.75
Technical fit20%40.8051.00
Service and uptime20%40.8030.60
Integration capability10%30.3050.50
Vendor stability10%50.5030.30
Safety and compliance10%50.5040.40
Contract terms5%40.2020.10
Total100%4.353.65

Vendor B wins the demo on technical fit and integration, yet loses the decision on cost and service. That inversion is exactly what the matrix exists to surface. Had the committee scored on demo impressions alone, B would have won, and the 5-year cost gap would have arrived as a surprise.

Handling Disagreement and Grade Inflation

Two safeguards keep the scores honest. First, have each evaluator score independently before the committee meets, then reconcile. The reconciliation discussion is where hidden assumptions surface: one reviewer scores service a 3 because a 4-hour response is a caveat, another scores it a 4 because it is contractual. Neither is wrong until the group defines the standard.

Second, run a sensitivity check. Raise and lower each weight by five points and see whether the ranking changes. If a two-point shift in the TCO weight flips the winner, the decision is genuinely close and the committee should look for another differentiator, such as a pilot. If the ranking holds across the range, the decision is robust and you can stop debating weights. The go/no-go gate that a pilot adds to this arithmetic is set out in the site survey and readiness method.

Tie-Breakers and the Award Decision

A matrix can produce a tie or a near-tie. Resolve it in a fixed order, written down before scoring, so the tie-break is not a backdoor for the preferred bidder.

Record the final matrix, the evidence behind each score and the tie-break used. That record is what makes the award defensible in a procurement audit and what lets you hold the winning vendor to the claims that earned them the score. For a live fleet, the same scoring discipline reappears at renewal, where the incumbent is scored against challengers using identical criteria, a process that keeps pricing honest without the disruption of a full re-tender.

Frequently Asked Questions

How many vendors should be scored? Three to five. Two gives no spread and becomes a binary argument; more than five and the evidence-gathering cost exceeds the value of the extra comparison.

Should price be its own criterion? No. Price is an input to TCO. Scored separately it double-counts and drifts the decision back toward the cheapest bid, which is the outcome the matrix exists to prevent.

Who should score? At least three people with different exposure: operations, who live with uptime; finance, who own the TCO; and IT or facilities, who own integration. Single-evaluator scoring defeats the purpose. The scoring discipline here also underpins the lease versus buy decision, where the same weighted logic decides the financing route rather than the vendor.

Products