Report 172 · Defense Tech
What a bomb robot demo doesn't prove
Every ground robot sold to a bomb squad comes with a video: it climbs the stairs, opens the door, reaches into the car. One clean run shows the robot can do the task once. The standard test methods that NIST built with DHS and bomb technicians ask a different question, whether it will do it the next time, driven from out of sight, with nothing breaking. That takes 10 to 30 repetitions per task, and the reason is plain statistics.
In counter-IED work, the robot exists for one reason: to keep a person away from the device. NIST's own C-IED guidance puts the responder's goal in four words, "start remote and stay remote." A robot that quits halfway down a stairwell doesn't just fail a task. It sends a human downrange to finish the job, which is the exact outcome the robot was bought to prevent. That is why the way these robots get tested matters more than how good the brochure looks.
Who wrote the test, and why
Since the mid-2000s, NIST's Intelligent Systems Division has led a program, sponsored mainly by the DHS Science and Technology Directorate with support from DOJ, the Army Research Laboratory, JIEDDO and DARPA, to build standard test methods for "response robots": remotely operated ground, aerial and aquatic systems used by bomb squads, hazmat teams, urban search-and-rescue teams and soldiers. The methods are balloted through an ASTM International committee on response robots (numbered E54.08.01 in the 2014 and 2015 NIST documents, E54.09 on NIST's current pages). NIST's ground-robot page says the suite now has "more than 50 test methods" and has been used to specify more than $60M of robot procurements for organizations doing C-IED missions, with more than 250 civilian and military bomb technicians using it.
The problem it was built to solve is stated plainly in NIST's 2015 C-IED training document. New robots promise advanced capabilities and friendlier interfaces, "but it is hard for responders to sift through the marketing." NIST's web page goes further, describing its validation exercises as a way to "cut through the inevitable marketing, even within development programs." That last clause is worth noticing. Even government development programs need help separating a demo from a capability.
A test method is not a spec sheet
NIST is careful to say it is not defining a "standard robot." It is defining standard ways to measure one. Each test method has four parts, per NIST: a cheap, reproducible apparatus (stairs at a set angle, a pipe pattern for manipulator tasks, a car-sized prop for vehicle-borne IED tasks), a scripted procedure, a quantitative metric, and a defined fault condition, meaning a failure that stops the robot completing its required run of repetitions. Every fault gets logged, along with every field repair.
Three rules in the 2014 guide do most of the work of separating a test from a demo:
- The operator can't see the robot. Testing is "always remotely operated, out of sight and sound of the test apparatuses," so the operator relies on the robot's cameras and interface the way they would downrange. Practice with eyes on is allowed. Scoring is not.
- Forward and reverse. The training rules have robots drive the maneuvering tests forward to the far end and back through the obstacle in reverse. In real work, getting the robot back out is half the mission.
- The whole configuration, every time. Robots are tested across typically 20 to 30 methods per configuration, and any change, "even changes as simple as tracks vs. wheels," means retesting across the suite. The guide gives the example that "addition of the manipulator inhibited stair climbing." A spec sheet will often list stair climbing and the arm as two features. The test asks whether you get both at once.
The number that matters: 10, 20 or 30
Here is the core of it, from NIST's 2015 C-IED document. Each test method involves 10 to 30 repetitions, chosen by the standards committee to give "roughly 80% reliability with 80% confidence that the robot can perform the next repetition." The pass rule:
This means that within the first 10 repetitions in any test method, no faults are allowed; 1 fault is allowed in 20 repetitions; and 3 faults are allowed in 30 repetitions.
The arithmetic behind those thresholds is standard binomial statistics, and it is worth doing once, because it shows exactly how little one clean run proves. Suppose a robot really succeeds only 80% of the time on a task. The chance it would still go 10 for 10 by luck is 0.810, about 11%. So passing 10 clean repetitions gives you about 89% confidence the robot is at least 80% reliable. The 20-with-one-fault rule works out to about 93%, and 30-with-three to about 88%. All three clear the 80% bar.
Now run the same arithmetic on a demo video. If a robot that is only 80% reliable gets one attempt, it succeeds 80% of the time. One clean run gives you 20% confidence it meets that bar. Three clean runs in a row: about 49%. That is a coin flip, on a task where failure means a person walks downrange. The NIST training rules say the same thing in plain English: "A few successful repetitions are meaningless indicators of performance."
Who drives during the test
One more design choice is easy to miss and quietly clever. Capability testing uses "expert" operators supplied by the manufacturer. That sounds like it favors the vendor, and it does, on purpose. The guide's logic is that incentives are aligned: the vendor's best driver produces the best performance the robot can deliver, which becomes the baseline. In the guide's words, "No robot developer should promise any more than the capability demonstrated."
That baseline then becomes the yardstick for everyone else. A squad's own operators are scored as a percentage of the expert, and NIST offers example tiers (novice 0 to 39%, proficient 40 to 79%, expert 80 to 100%) that user communities can adopt or adjust. The robot gets measured, and so does the human skill, which NIST calls "very perishable."
A missing result is a result
Vendors are allowed to abstain from a test method or to withhold the data from a trial that went badly. The page is then stamped "ABSTAINED." This is a deliberate choice to let developers try tests without consequence. But the guide is clear about how a buyer should read the gap:
If the test method is considered critical to the operational needs of the sponsor or user, the test should be considered failed until the robot can demonstrate satisfactory performance at a later date.
That is the single most useful sentence in the whole document for anyone writing a purchase requirement. A brochure only shows you the tasks the robot did well. The test record shows you which tasks it skipped.
What to ask before you believe the video
- How many repetitions, and how many faults? One run, or 10 clean in a row? If there were faults, did it reach 20 with one, or 30 with three?
- Was the operator out of sight? A robot driven by someone watching it from five feet away is being tested on the operator's eyes, not the robot's cameras.
- Which configuration? Stairs with the manipulator fitted, with the battery or tether you will actually use, at the angle your buildings actually have. The standard stair test, per NIST, runs up to 45 degrees with rounded steel treads and wet surfaces; the training variant is gentler.
- Reverse too? Getting in is half of it.
- What did it abstain from? Treat a skipped critical test as failed until shown otherwise.
What I could not confirm
I read NIST's guidance, not the ASTM standards themselves. The individual ASTM test methods are sold separately and I did not open them. Everything here about procedures and pass rules comes from NIST's 2014 guide, its 2015 C-IED training document and its current web pages, which NIST wrote to explain the standards.
The documents are 10 years old. Counts have grown since (the 2015 document says fifteen methods were standardized at that point; NIST's current page says more than 50 test methods exist). I can't tell you which thresholds, if any, the committee has revised since 2015. NIST's page also notes that test sponsors can adjust the 80/80 thresholds for their mission.
The confidence figures are my arithmetic. They are the standard binomial calculation, assuming each repetition is independent with the same success chance. NIST states the 80/80 target and the 0/10, 1/20, 3/30 rule; it does not print the percentages I derived from them.
Nothing here evaluates or recommends any robot or manufacturer.
The signal
A demo answers "can it?" Bomb technicians need "will it, next time, from out of sight?" The NIST-DHS-ASTM test methods turn that into a number: 10 straight successes, or 20 with one fault, or 30 with three, driven remotely, forward and back, in the exact configuration you will field. Below that, the statistics say you are guessing. When a vendor shows you one perfect run, ask for the test form. And if a critical test says "ABSTAINED," read it as a fail.
Sources
- Adam Jacoff, Ann Virts and Kam Saidi, National Institute of Standards and Technology, "Counter-Improvised Explosive Device Training Using Standard Test Methods for Response Robots" (v2015, PDF dated November 2015, 11 pp). (PRIMARY. Read in full. Source of, verbatim: "start remote and stay remote"; "but it is hard for responders to sift through the marketing"; "roughly 80% reliability with 80% confidence that the robot can perform the next repetition"; the 10/20/30 fault rule quoted in full; "even changes as simple as tracks vs. wheels"; "A few successful repetitions are meaningless indicators of performance"; and "very perishable." Also source of the forward-and-reverse rule, the 20 to 30 methods per configuration, the novice/proficient/expert tiers, the 2015 count of fifteen standardized methods, and the standard stair test up to 45 degrees with rounded steel treads and wet surfaces.)
- National Institute of Standards and Technology, sponsored by DHS Science and Technology Directorate, Office of Standards, "Guide for Evaluating, Purchasing, and Training with Response Robots Using DHS-NIST-ASTM International Standard Test Methods" (test director Adam Jacoff; PDF dated June 2014, 40 pp). (PRIMARY. Read the sections on testing approach, configuration identification, statistical significance, resets, abstaining and documentation. Source of, verbatim: "always remotely operated, out of sight and sound of the test apparatuses"; "addition of the manipulator inhibited stair climbing"; "No robot developer should promise any more than the capability demonstrated"; "ABSTAINED"; and the abstain passage quoted in full. Also source of the expert-operator logic.)
- National Institute of Standards and Technology, Engineering Laboratory, "Ground Robot Tests" (updated 3 October 2024) and "Standard Test Methods for Response Robots" (updated 16 August 2024). (PRIMARY. Read in full. Source of the sponsors list, ASTM committee E54.09, "more than 50 test methods," the more than $60M of C-IED robot procurements and more than 250 bomb technicians, "cut through the inevitable marketing, even within development programs," the four elements of a test method (apparatus, procedure, metric, fault condition), the not-a-standard-robot distinction, and that sponsors can adjust the 80/80 thresholds.)
- Onur Oncer, "What a drone trial actually proves," The Signal Report 067, and "A mine detector finds metal, not mines," The Signal Report 097. (Earlier reports in this beat on reading test claims.)
Scope note: this report explains publicly released NIST guidance on how response robots are evaluated and what that implies for reading capability claims. It evaluates no specific robot and contains no operational procedures for handling explosive devices.
Onur Oncer
U.S. Army combat veteran (Counter-IED / Electronic Warfare), peer-reviewed researcher in microwave spectroscopy, and founder & CEO of Shroombiosis. Consults on laboratory operations, AI, and supplement formulation.