Skip to content
← Help

Live and real-time transcription

Daisy transcribes while you record, on-device, and the text it produces during the meeting is your transcript. When you stop, Daisy promotes it. Nothing is sent anywhere at any point.

Whether you see that text while the meeting runs is a separate, purely cosmetic choice.

What you see while recording

The live transcript separates speech by source: your system audio is labeled Them, and your own microphone lines appear alongside it unlabeled. Each settled line carries a timestamp so you can find that moment again later. Words appear as they're spoken and settle into clean text a moment later; interim words look slightly faded and italic until they settle.

Daisy primes the transcriber with the meeting's own vocabulary — the title, the attendee names from a calendar event, and the names and term lists on any tags you attached. Those words are recognised more reliably, and they keep their capital letters.

Before any speech arrives you'll see waiting for live transcript….

Captions also appear in the floating mini-window, which shows a tail of the last few far-side lines as bubbles. Your own voice is left out there. Until the other side speaks it reads Waiting for others to speak….

"Possible missed speech" markers

Occasionally you'll see a line like:

[possible missed speech — meeting audio was faint for 4s]

That means Daisy heard energy on that track, a voice-activity check agreed it looked like speech, but the transcriber produced no words for it — usually very faint or badly clipped audio. It's an annotation: the decode already ran and the audio is on disk, and the marker is there so a quiet stretch reads as "something was here".

You can turn these off under Settings → Providers → Advanced → Transcription tuning → Faint-audio markers. They aren't available on macOS.

Showing or hiding captions

Settings → Recordings → Live captions → Show captions while recording is a single on/off switch, saved per machine. The group hint says it plainly: Text as you talk. Display only — the transcript is always produced.

Turn it off and the recording view shows a card titled Live captions are hidden, noting that the transcript is still being written. Transcription runs either way; the switch only changes whether you watch it happen.

When there's no model installed

If no transcription model resolves on this machine, Daisy tells you twice:

  • A banner across the top: No transcription model installed. Recordings are audio-only until one is installed in Settings → Providers, with an Open Settings button.
  • In the recording view, a card titled Live transcription is unavailableNo speech model was found on this machine — audio is still being recorded and can be transcribed once a model is installed.

Your audio is still captured normally, and transcribing it later once a model is in place produces the same result. See Transcription language for the model list.

If an engine stumbles mid-meeting

Daisy supervises the running transcriber. If it fails to start it falls through to the other installed engine; if it dies after producing text it gets one restart before Daisy moves on; and if it falls sustainedly behind real time — a machine under heavy load — Daisy demotes the session to record-only. Audio is kept in every one of these cases. The Recording status toast shows the engine currently in use (Transcribing live: …) and follows the switch.

You may see the same line twice — that's normal

During a call, the other person's words can appear twice — once on your own microphone line and once on the labeled Them line. This is expected.

Daisy captures two independent audio streams:

  • Your microphone — your voice. In the live view these lines carry no label.
  • Them — the system audio coming out of your speakers (the other side of the call), labeled Them.

When you use speakers, your microphone picks up the other person's voice, so their words get transcribed on both sides. Daisy detects these mic-bleed duplicates by comparing the wording and drops one copy. In the live view it always drops the microphone copy and keeps the Them line. At finalize it compares the two recordings acoustically to work out which copy is the echo, and keeps the true source.

Using headphones eliminates this at the source: the other side's audio never leaks back into your mic.

What changes when you stop

When you press ✦ Finish & summarize, Daisy promotes the live transcript, fills in any stretch the live pass didn't cover, runs echo cancellation, sorts out who said what on-device, removes the mic-bleed duplicates, and writes the summary. See the FAQ for rough timings.