Resources
Are veterinary AI scribes accurate?
By Owl ·
The short answer
Mostly, yes — modern veterinary AI scribes transcribe and summarize visits well most of the time. But accuracy isn't the right question. The risk isn't the 99% they get right; it's what they do with the fraction they're unsure about. Most AI scribes fill the gap with a plausible-sounding invention. That's called a hallucination, and it's the thing to actually evaluate.
What "accuracy" hides
Most veterinary AI scribes market an accuracy percentage. It sounds reassuring, but an average tells you how the tool behaves on the easy majority of a note — not what it does on the hard parts: the mumbled dosage, the finding you mentioned once, the moment two people talked at once. An accuracy number can't distinguish between a tool that leaves a gap honestly and one that papers over it with a confident guess. For a medical record, the second behavior is the dangerous one.
What is a hallucination?
In AI, a "hallucination" is when a model generates content that wasn't in its input — text that sounds right but was never said or seen. In a clinical note, that means a symptom, dose, or diagnosis that the visit didn't actually contain. A peer-reviewed 2025 review in npj Digital Medicine describes the failure plainly: AI systems can generate "entirely fictitious content, such as documenting examinations that never occurred or creating nonexistent diagnoses." (npj Digital Medicine, 2025)
Does this actually happen? Yes — and it's been measured
The most-cited example is OpenAI's Whisper, a speech-to-text model widely adopted in healthcare. In 2024, researchers reported to the Associated Press that Whisper invented content that was never spoken — including imagined medical treatments. A University of Michigan researcher found hallucinations in 8 of every 10 audio transcriptions of public meetings; a machine-learning engineer found them in more than half of 100+ hours of audio; one developer found them in nearly all of 26,000 transcriptions. OpenAI responded that its policies prohibit using Whisper "in certain high-stakes decision-making contexts." (TechCrunch / AP, 2024)
Even under clean, controlled conditions the problem doesn't vanish. In a study presented at the 2024 ACM Conference on Fairness, Accountability, and Transparency, researchers ran 13,140 short audio segments through Whisper and found roughly 1% contained hallucinations — including invented medications — on research-grade audio. The honest reading of both findings together: on clean audio the rate is low; in the noisy reality of a working clinic, it climbs.
Why exam rooms make it worse
Clinic audio is the hard case: overlapping speech, background noise, a patient who won't stay still, terminology and drug names a general model hasn't heard much of. Those are exactly the conditions where a model is most likely to guess. And a veterinary record has no human stenographer double-checking it — whatever the scribe writes is what goes in the chart, and often into the client's takeaway summary, unless the vet catches it.
This isn't a reason to avoid AI scribes — it's a reason to choose carefully
AI scribes genuinely help. Documentation burden is one of the most consistently cited stressors in the profession; the Merck Animal Health Veterinary Wellbeing Study, run with the AVMA, has found roughly half of veterinarians report burnout, with after-hours charting and record-keeping near the top of the list. (AVMA) Giving vets their evenings back is a real and worthy goal. The point isn't to distrust the category — it's to ask the right question of any tool you let near a medical record.
The question that actually matters: does it guess, or does it flag?
When an AI scribe isn't sure, it can do one of two things. It can write its best guess — fluent, plausible, and sometimes wrong. Or it can leave the field blank and tell you it wasn't sure, so you fill it in. The first protects the illusion of a complete note. The second protects the patient. When you evaluate a veterinary AI scribe, this is the behavior to test for — not the accuracy percentage on the box.
How Owl approaches it (and what we won't claim)
Owl is an AI scribe, and like every AI scribe it's built on models that can hallucinate — we won't pretend otherwise. What's different is the engineering around that fact. Owl runs on an anti-fabrication contract: when it isn't confident about a detail, it leaves the field blank and flags it for you rather than inventing something to fill the space. Its transcription layer also isn't Whisper — it uses a speech engine that the same 2024 research found did not exhibit those transcription fabrications — but that only addresses one layer; the blank-and-flag contract is what handles the rest. The goal isn't a tool that's magically never wrong. It's a tool that tells you where it might be, instead of hiding it. That's the difference between a note you have to second-guess and one you can trust.
Key takeaways
- Veterinary AI scribes are accurate most of the time; "accuracy %" is the wrong thing to judge them on.
- The real risk is hallucination — invented detail that reads as real. It's documented (Whisper: from ~1% on clean audio to a majority of real-world transcriptions) and peer-reviewed (npj Digital Medicine, 2025).
- Exam-room audio makes guessing more likely, and there's no stenographer to catch it.
- Ask one question of any scribe: when it's unsure, does it guess or does it flag?
FAQ
Questions, answered.
Are veterinary AI scribes accurate enough to trust?
For most of a routine note, yes. The caveat is the minority of cases where the model is unsure — there, some scribes invent a plausible detail rather than flag the gap. The safe choice is a scribe that surfaces uncertainty instead of hiding it, and a vet who reviews the note before it's signed.
Do AI scribes really make things up?
Yes, it's a documented behavior called hallucination. Researchers found OpenAI's Whisper invented content — including medical treatments — in transcripts, and a 2025 npj Digital Medicine review describes AI scribes documenting "examinations that never occurred." Rates range from ~1% on clean audio to a majority of real-world recordings.
How do I stop an AI scribe from hallucinating in my notes?
You can't eliminate the underlying possibility, but you can pick a tool designed to leave uncertain fields blank and flag them rather than guess, keep the human review step, and prefer scribes that don't rely on transcription models with known fabrication issues. Always confirm the note before it enters the record.
Is Owl immune to hallucination?
No — and any tool that claims to be is overselling. Owl is built on AI models that can hallucinate; what it does is flag uncertainty (leave the field blank) instead of inventing detail, so you see the gap rather than a confident guess.