← The Signal Report Work with me

Report 121 · Defense Tech

Fifty-six bomb dog teams, none passed

In three regions of the United States, 56 explosive-detection dog teams ran the searches that the new national certification standard prescribes. Not one of them cleared the bar. The failure rate is not the interesting part. The interesting part is that the same study showed the standard is doing its job, and that what separated the teams was not talent or training quality. It was whether they could get their hands on the explosives at all.

I spent my Army career on the counter-IED side of the house, which means I spent a lot of it around detection. Different sensors, same recurring lesson: a detector is only as good as the conditions it was calibrated under, and the number on the certificate is a claim about those conditions, not about the world. I have written that argument here about metal detectors, about radar, and most recently about what heat does to a working dog.

The dog is the oldest sensor in that inventory and still, for a lot of tasks, the best one. It is also the one with the weakest paper trail. Until recently there was no national consensus standard describing what an explosive-detection dog team has to demonstrate before you put it to work. Now there is, and somebody finally took it into the field and measured what happens.

The answer, published in Frontiers in Veterinary Science in October 2025 by a team led by Michelle Karpinsky and Lauryn DeGreeff and funded by NIST, is that out of 56 teams, zero would have certified.

That sentence is going to get quoted without the rest of the paper attached. So here is the rest of the paper.

What the standard actually asks for

The document in question is ANSI/ASB Standard 092, first edition, 2021, published by the AAFS Standards Board. It was approved by the ASB in July 2021, by ANSI in December 2021, and added to the OSAC Registry on 6 December 2022. It covers training and certification for explosives-detection canines, person-screening canines, and the combination of the two.

Three clauses matter for what follows. The threshold, in section 6.7:

For successful certification, the canine team shall achieve at least an overall 90% positive alert rate and an overall false alert rate not to exceed 10%.

The material, in section 6.5.2: "Certification shall be conducted with actual explosive, or explosive precursor chemicals." Not pseudo-scents, not simulants. The real thing.

And the blinding, in section 5.8.1.1.9.2: "The canine handler shall not know the total number or placement of target odors for the totality of the exercise(s)." That clause is the standard quietly acknowledging the Clever Hans problem, which is that a handler who knows where the explosive is will tell the dog without meaning to.

Ninety percent with under ten percent false alerts is a demanding bar, and deliberately so. For comparison, the paper notes that the North American Police Working Dog Association certification requires "a minimum accuracy rate of 91.6% and permitting only a single miss over the course of the certification evaluation." This is not an outlier standard. It is roughly where the field sets the bar.

What they did

The study ran three trials in the southwestern, southeastern, and western United States, with 20, 17, and 19 canine and handler teams respectively, over two days each. Every team ran two kinds of search: the assessments prescribed by Standard 092 (rooms, parcels, vehicle exteriors, luggage, odor recognition) and a set of operational scenarios written to look more like real work, including people carrying bags, piles of luggage, and vehicle interiors.

It was a black box study, which in this context means the researchers measured outcomes without controlling or coaching the teams' methods. The blinding was single-blind: "The searches conducted were single-blind, meaning the evaluators knew the placement and number of the targets, but the handlers did not." Six explosive types were used, referred to in the paper only as Explosives 1 through 6 for security reasons.

The numbers

Correct alert and false alert rates, by trial:

Trial 1: 38% correct and 6% false on the Standard 092 assessments, 45% correct and 7% false on the scenarios. Trial 2: 72% correct and 13% false on Standard 092, 65% correct and 15% false on the scenarios. Trial 3: 42% correct and 9% false on Standard 092, 30% correct and 13% false on the scenarios.

Against a 90 percent bar, none of those trial averages are close. At the individual level: "Team 9 from Trial 2 was the highest scoring team with an 88% correct alert rate, with a 6% false alert rate on Standard 092 assessments." That team came within two percentage points on detection with a comfortable false alert rate, and still would not have certified. Five teams achieved 70 percent or better with under 10 percent false alerts; three more cleared 70 percent but with false alert rates above 10 percent. Twenty-one teams landed between 36 and 50 percent, and 15 came in below 35 percent.

The paper's conclusion is blunt: "Analyzing the individual canine/handler team responses revealed that no team would have passed the OSAC certification."

The finding underneath the headline

Here is the part that gets lost, and it is the reason the study was run in the first place.

A brand new certification standard has an obvious open question: does passing it mean anything operationally, or is it a bureaucratic hoop? You cannot answer that by writing more standard. You have to test the same teams on the certification searches and on realistic searches and see whether the two track each other.

They track each other. The eight teams that scored above 70 percent on Standard 092 averaged "79% ± 6% on Standard 092 and 86% ± 16% for the scenarios with 4 out of the 8 teams achieving a 100% positive alert rate on the scenarios," which the authors read as high performance on the standard predicting high performance on realistic tasks. Their summary sentence: "This suggests that meeting the Standard 092 criteria may predict success in operational contexts too."

Lauryn DeGreeff put it this way in the journal's announcement of the work:

Here we show that the performance of canine teams in the new official assessment is indeed informative, in that it may predict their performance in real-world scenarios.

So the correct one-line summary of this paper is not "bomb dogs fail." It is "the new certification appears to measure the right thing, and almost nobody can currently meet it." Those are very different findings, and only one of them is about the dogs.

Why nobody passed

The authors offer two explanations and lean hard on the second.

The first is that the bar is simply high: "One possibility is that the OSAC standard represents a particularly stringent and demanding certification process." Fair, and worth keeping in view.

The second is logistics, and this is where the paper stops being a story about animals and starts being a story about supply. "Second, four of the six explosives used within the trial, and required by Standard 092, can be difficult or expensive to obtain." Teams that struggled with those specific materials "likely had limited exposure to them during routine training." And then the sentence that reframes everything:

In many cases, these teams may only encounter these explosives once or twice a year when completing their certification hosting by other agencies, such as the North American Police Working Dog Association (NAPWDA) and the National Police Canine Association (NPCA).

Read that again with the threshold in mind. The standard requires certification on actual explosive material. Working teams may touch four of those six materials once or twice a year, at somebody else's certification event. Then they are assessed at 90 percent on all six.

Paola Prada-Tiedemann, one of the co-authors, said the same thing directly: "We found that limited access to explosive training materials and training opportunities were the primary challenges for the teams and largely explained the geographic variation seen in the data."

The data point that settles it

If the training-access explanation is right, then exposure should produce improvement, and fast. The trials ran over two days, which is not a training program by any definition. It is barely an introduction.

It was enough. "The most notable improvement of detection was on Explosive 4 and 5 in Trial 1 with an increase of 28 to 74% alert rate and 8 to 50% alert rate, respectively, and on Explosive 5 in Trial 2 with an increase of 25 to 74% alert rate, implying that exposure to the target improved later detection." One material in Trial 1 went from being found 5 percent of the time on day one to 67 percent on day two.

A dog that goes from 8 percent to 50 percent on a substance overnight was not a bad detector the day before. It was an unfamiliar one. That is a generalization gap, and generalization gaps close with exposure to varied material, which is exactly what the authors recommend.

The trend was not universal. Trial 3 went the other way, with performance flat or declining across day two, most sharply on Explosive 2, which fell from 92 percent on day one to 35 percent on day two. The authors do not resolve that, and neither will I. Fatigue, heat, handler stress and search-order effects are all live candidates, and the paper is explicit that it "did not directly measure cognitive or biological state variables."

Why this reads like a supply problem to me

In counter-IED work, the difference between a capability that exists on paper and one that exists in the motor pool is almost always consumables, access, and time on the equipment. You can write a superb standard and still field teams that cannot meet it, because the binding constraint is upstream of the standard entirely.

Real explosive training aids are regulated, expensive, hazardous to store, and legally awkward to move between jurisdictions. A department that wants to train its dog on a peroxide-based explosive cannot simply order one. The result is a system where the certification is written around six materials and the training pipeline reliably supplies two.

That is not a dog problem or a handler problem. It is an inventory problem that shows up on a score sheet as a detection failure, and it will not be fixed by more testing. It gets fixed by making varied training material routinely available to the teams expected to perform on it, which is precisely the recommendation the paper lands on and precisely what the journal's own headline emphasized: wider access to explosives.

What this study does not say

It does not say working dogs are ineffective. It measured performance against a specific, deliberately stringent written standard, with six specified materials, on two days, in a test environment. Deployed teams typically work a narrower and more familiar threat set.

It does not establish a national error rate. Fifty-six teams across three regions is a proof-of-concept, and the authors call it that. The teams volunteered, which introduces selection in an unknown direction.

It was not double-blind, and the authors flag the consequence themselves: "since the study was not double-blinded, it was possible that assessors provided unintentional clues to handler about the location of targets." Note the direction of that bias. Cueing would tend to inflate scores, not depress them, which makes the failure rate harder rather than easier to explain away.

And it does not tell you which chemistries the dogs missed. The six materials are anonymized, so nobody reading this can conclude anything about a specific explosive. That is a reasonable security decision and a real limit on what the public can learn from the paper.

What I could not confirm

I read the full text through Europe PMC and the complete Standard 092 PDF, both of which are open. I did not review the study's supplementary material, which the authors reference and which holds the per-team data. Figure-level values that appear only in plots, rather than in the text, are not quoted here: an earlier pass at this paper produced two different sets of per-trial correct-alert percentages before I pulled the full text and pinned them to sentences, which is a good argument for not citing numbers read off a chart.

I have not independently verified the NAPWDA and NPCA certification requirements described in the paper. Those are quoted as the study characterizes them.

Finally, this is one study of one standard at one moment. Standard 092 is a first edition. If a second edition changes the threshold or the material list, the arithmetic in this report changes with it, and I will note it here if that happens.

Sources

  1. Karpinsky M, Browning H, Quigley-McBride A, Bunker P, Chapman W, Prada-Tiedemann PA, DeGreeff LE, "Explosive detection canines in the field: a multi-site black box validation study," Frontiers in Veterinary Science 12:1668317 (2025), DOI 10.3389/fvets.2025.1668317, published 16 October 2025. Funded by the National Institute of Standards and Technology, grant 2021-NIST-MSE-01. (Primary source. Full text retrieved and read as JATS XML from Europe PMC, PMC12571827. Source of: the 56 teams across three geographic trials at 20 / 17 / 19; the two-day structure; the single-blind design and its verbatim description; the six anonymized explosive types; all per-trial correct and false alert rates; the "no team would have passed" finding; Team 9's 88% and 6%; the five teams at 70% or better with under 10% false alerts and the additional three above 10%; the 21 teams between 36 and 50% and 15 below 35%; the 79% ± 6% and 86% ± 16% figures for the high-performing group and the 4-of-8 at 100% on scenarios; the "may predict success in operational contexts" conclusion; the "particularly stringent and demanding" possibility; the training-aid difficulty and expense passage and the "once or twice a year" quotation; the day-one to day-two improvement figures for Explosives 4 and 5; the Trial 3 decline from 92% to 35% on Explosive 2; the statement that cognitive and biological state variables were not measured; the double-blinding limitation quotation; and the NAPWDA 91.6% comparison as characterized by the authors.)
  2. AAFS Standards Board, ANSI/ASB Standard 092, "Standard for Training and Certification of Canine Detection of Explosives," First Edition, 2021. ASB approved July 2021, ANSI approved December 2021. (Primary source. Full 48-page PDF downloaded and read directly. Source of the section 6.7 certification threshold quoted verbatim, the section 6.5.2 requirement that certification use actual explosive or explosive precursor chemicals, the section 5.8.1.1.9.2 handler-blinding requirement, and the note that organizations may set a stricter false alert threshold than 10%.)
  3. NIST, "ANSI/ASB Standard 092-21 Standard for Training and Certification of Canine Detection of Explosives, 2021, 1st Ed.," OSAC Standards Library entry. (Opened and read. Used only to confirm the standard's status on the OSAC Registry and the 6 December 2022 registry date.)
  4. Frontiers, "Sniffer dogs tested in real-world scenarios reveal need for wider access to explosives, study finds," 16 October 2025. (Publisher announcement, opened and read. Used only for the two author quotations from Lauryn DeGreeff and Paola Prada-Tiedemann, both reproduced verbatim, and to confirm how the publisher framed the finding. No claim in this report rests on the announcement alone.)
Onur Oncer
Onur Oncer

U.S. Army combat veteran (Counter-IED / Electronic Warfare), peer-reviewed researcher in microwave spectroscopy, and founder & CEO of Shroombiosis. Consults on laboratory operations, AI, and supplement formulation.

← All reports