Pulse — news for clinicians

← All posts

explainer   July 19, 2026

How an ambient AI scribe actually works — and where it fails

Mic to transcript to note: the four-stage pipeline inside every AI scribe, explained without marketing, including the failure modes vendors don't lead with.

Every ambient AI scribe on the market — ours included — is the same four-stage pipeline. Vendors differ in the quality of each stage and in how honestly they describe it. Here is the whole thing.

Stage 1: The microphone

A phone, watch, or laptop mic captures the room. This is the stage most failures actually start in, because an emergency department or a busy clinic is an acoustically hostile place: overhead pages, monitors, a crying child two bays over, two people talking at once.

What to ask a vendor: what happens when a phone call interrupts the recording? What happens when the app is backgrounded? Does audio survive a crash? Recording loss mid-visit is one of the most common real-world complaints about deployed scribes, and it is a capture-engineering problem, not an AI problem.

Stage 2: Speech-to-text

The audio becomes a transcript, either on a server or on the device itself. Modern medical STT is good but not magic. Documented weak spots: accented speech, disordered speech, drug names that sound alike, and negations ("no chest pain" vs. "chest pain"). A University of Cincinnati review this month listed accent and noise performance among the top risks of the whole category (coverage).

The on-device vs. cloud choice is a privacy trade, not a quality trade: on-device transcription means raw audio never leaves the hardware; cloud transcription is typically more accurate for hard audio but means the recording travels.

Stage 3: The language model

A large language model takes the transcript and rewrites it as a clinical note in the requested format. This is where the two failure modes you should actually worry about live:

Omission. The model summarizes. Summaries drop things. A detail you said once, quietly, in the middle of a tangent, may not survive into the note.

Fabrication. The model produces fluent text, and fluent text can include findings never said aloud — a normal exam element you didn't perform, a plausible-sounding medication dose. Published analyses of scribe output put combined error rates (omissions plus hallucinated content) in the mid-single-digit percent range, which is low per sentence and high per shift.

There is no vendor whose model does not do both of these sometimes. The differences are in rate, in whether the system shows you the verbatim transcript so you can check, and in whether the note's claims are traceable back to what was actually said.

Stage 4: Your review

The final stage of the pipeline is you, and it is not optional. The note is a draft. You are signing it, your name is on it, and every guideline, malpractice carrier, and honest vendor says the same thing: read the whole note, not the first three lines. The Cincinnati paper's own conclusion was that full human review removes a large share of the risk it catalogs.

What this means for choosing one

Skip the demo-day accuracy claims — no vendor publishes an independently reproduced word-error rate. Instead ask four questions that map to the four stages:

  1. What happens to the recording when the phone rings, the app backgrounds, or the battery dies?
  2. Where does transcription run, and can I see the verbatim transcript alongside the note?
  3. What is the measured edit burden — how much do real users change before signing?
  4. Does the workflow force review, or make it easy to skip?

A scribe that answers all four plainly is worth trialing. One that answers with adjectives is not.

tebIQ shows you the verbatim transcript before the AI note — that ordering is deliberate. See it at tebiq.com.

Charts done before you leave the building.

tebIQ is physician-built AI documentation. Try it free — no demo call, no sales rep.

Start free