What It Took to Make a Dog's Nose Into Medicine
Sensitivity, specificity, and an AUC of 0.962. Here's what actually went into that number, and what it does and doesn't prove yet.


A Superpower That Never Quite Made It Into Medicine
In 1989, a woman went to her doctor because her dog would not leave a mole on her leg alone, sniffing at it insistently for months until she finally had it checked. It was a malignant melanoma, caught early enough to treat. The case made it into The Lancet, one of the world's most respected medical journals. Medicine noticed, cited it occasionally, and largely moved on.
Fifteen years later, a UK team ran the first properly controlled test of the idea, and found dogs could pick out cancer samples at a rate well above chance. Over the two decades since, similar studies have piled up across cancer types and sample types, breath, urine, blood, tissue, many reporting sensitivities and specificities above 90%. The pattern kept repeating: disease, it turns out, leaves a chemical trace, and a dog can find it.
So why did this never become something medicine actually used? The honest answer is that a nose, however extraordinary, cannot scale on its own. What was missing was a way to read what the dog was detecting, standardised and reproducible enough to screen populations rather than a few dozen people at a time, and rigorous enough for medicine to trust. That is what we spent the last few years building, with the dogs' own noses at the centre of it. The results were published in the Journal of Clinical Oncology in April 2026.
The Study, and the Numbers
We enrolled 3,275 participants across six hospitals in Karnataka, split between Hubballi and Bengaluru. Some had biopsy-confirmed cancer across seven cancer types, recruited before treatment began. Others were healthy or living with non-cancerous conditions, recruited from the same clinical settings, so that any signal the dogs picked up reflected cancer specifically, not just the scent of a hospital.
Each participant did one simple thing: breathed normally into a surgical mask for ten minutes. The mask was sealed and sent to our lab, where seven dogs on our team took over. Sessions run on SniffSpace, our sample-presentation platform, with barrier gates, QR-logged samples, and infrared sensors timestamping every sniff. Each dog gives a simple read, alert or pass, and at least three dogs independently assessed every sample. Those individual reads were combined through a Bayesian model that weighted each dog's own historical accuracy and factored in each participant's clinical risk profile, producing a single calibrated score.
On a held-out set of 1,502 participants the dogs had never encountered before, sensitivity came out to 90.8%, specificity to 91.3%, and the AUC to 0.962.

Where the Accuracy Actually Comes From
Performance held steady across all seven cancer types, and across disease stages, sensitivity for Stage 1 and 2 cancers was 90.6%, nearly identical to late-stage performance. That matters, since how early a cancer is caught shapes almost everything about what happens after.
It's worth being specific about what is driving the number. A model built on clinical risk factors alone, age, tobacco history, family history, reached an AUC of 0.832. The dogs alone reached 0.932. Combined through the Bayesian fusion model, the system reached 0.962, a result neither the clinical data nor the dogs could reach on their own.

We also stress-tested the system the way you would need to for something meant to screen millions of people rather than hundreds: removing a dog from the ensemble entirely to simulate an off day, and checking whether results held regardless of a participant's sex, smoking history, or common conditions like diabetes. They did.
Teaching a Machine to Listen More Closely
The Bayesian model decodes what the dogs find. The next layer is learning to read the dogs themselves more precisely than any human handler could. In this study, computer vision tracked each dog's sniffing behaviour at every sample port, how long they lingered, the path of their approach, the rhythm of each sniff, feeding those signals into the model alongside each dog's binary read.

In a smaller pilot subset of 425 participants, this vision-augmented approach reached an AUC of 0.938, the first time automated behavioural sensing had been folded into the fusion framework at all. It's an early signal that some of what currently depends on a human observer watching a dog work could eventually be sensed automatically.
How the Field Responded
The Journal of Clinical Oncology is where oncologists go to update how they treat patients. Dr. Jonathan Friedberg, the journal's Editor-in-Chief, wrote that the results "pave the way for larger validation studies of canine olfaction as a cancer screening modality, with particular relevance for populations with limited health care resources." The study was also selected for a commissioned editorial, where independent reviewers weigh in on the work separately. Dr. Shiran Shapira and Prof. Nadir Arber called it an effort to move canine detection "beyond anecdotal observations toward a more structured analytical framework," concluding that the reproducibility of the pattern across labs suggests a real biological signal is at work.
What This Does, and Doesn't, Prove Yet
This study establishes what researchers call analytical validity: proof that the system can accurately tell cancer from non-cancer under controlled conditions, in a population that already included both confirmed cases and confirmed controls. What it does not yet establish is clinical utility, whether running this in a real, unselected screening population, made up of people with no symptoms and no prior cancer history, produces the outcomes we believe it can.
That is what our FirstAlert study is built to test next: higher-risk but otherwise healthy people, and people who have completed cancer treatment, to see whether the system can flag a returning cancer before it shows up clinically. A larger community screening study follows after that. Together, they're the path from a validated study to something that actually works at the scale of the population that needs it.

