Skip to content
← Help

Transcription models and language

Daisy never asks you to pick a language before a meeting — the model detects the spoken language itself. What you can change is which transcription engine and model are active, because only some of them handle languages other than English.

All of them run on your machine. None of them send audio anywhere.

Where the models live

Settings → Providers → Local transcription. One list, grouped by engine, with the active row highlighted and marked ● Active. A row that ships inside the app reads Built in. Below each name you get English or multilingual, the install size, and the licence.

One model is active at a time; switching applies to the next recording.

The three engines

Apple Speech — built into macOS

On macOS 26 and later, Daisy uses Apple's own on-device speech engine. Nothing to download from us — the operating system fetches its speech assets once, and Daisy shows a progress state while it does. It covers the languages your Mac's system settings do. This is the default on a Mac that supports it; older macOS falls through to Whisper.

Parakeet — fast on-device English

TDT 0.6B v2 (int8) — bundled with the Windows and Linux builds, so the first recording works with nothing downloaded and nothing configured. English only. It runs on your CPU, so it stays fast on machines with no usable GPU and leaves the graphics card to your other apps during a call. Mac users can download it too.

Whisper — multilingual, larger models available

The fallback and the multilingual path. Ten models from Tiny (78 MB) up to Large v3 (3.1 GB), each either English-only or multilingual. You download the one you want. Whisper uses your GPU where one is available (Metal on Mac, Vulkan on Windows and Linux) and your CPU otherwise.

Recording in another language

  1. Open Settings → Providers → Local transcription.
  2. Find a row tagged multilingual — any Whisper model without (English) in its name, or Apple Speech on a Mac.
  3. Click Download if you don't have it yet. Bigger models are more accurate but slower and use more memory.
  4. Click Use to make it active.

The change takes effect on your next recording, and the model auto-detects the spoken language from the audio. To go back to English, set an English row active again.

You can also pin Whisper to a specific language, under Settings → Providers → Advanced → Transcription tuning → Whisper → Language — a two-letter code such as en, or blank for automatic. This is worth doing only when auto-detection keeps guessing wrong on short or noisy audio.

Downloading, switching, and deleting

  • Download fetches the model and verifies it. A file download can be cancelled at any point; Apple's OS-managed install can't, and shows no percentage. A download that fails leaves a Retry download button and tells you to check your connection and free disk space.
  • Use switches the active model.
  • Delete removes your downloaded copy. For a Built in row this removes only the copy in your profile — the bundled one stays and the row reverts to it. Apple Speech can't be deleted; it belongs to the OS.
  • Delete the last copy of the active model and Daisy falls back to the bundled model for your platform, then to any other model you have installed.

If nothing is installed

Uninstall everything on a Mac too old for Apple Speech and you'll get a banner: No transcription model installed. Recordings are audio-only until one is installed in Settings → Providers. Recording still works — the audio is captured and kept, and you can transcribe it once a model is back. See Live and real-time transcription.