Overview

Requirements

Operating system macOS 14 or later
Hardware Apple silicon
Language English
Default model Large v3 Turbo, with Base, Small and Medium available
Network Needed once to download the speech model, then optional
Permissions Microphone, Accessibility, Input Monitoring
Location on disk Applications folder, so macOS remembers the permissions

On Apple silicon specifically#

Transcription runs through whisper.cpp with Metal acceleration. That is what makes local transcription fast enough for a hold-and-release loop rather than something you wait on. Intel Macs are not supported.

On English#

Phantom Voice transcribes English. If you need another language, this is not the tool for you today.

On the network#

The speech model downloads once, on first run. After that dictation works offline. Purchase, model download and update checks still use the internet when needed, and none of them carry your dictation audio. See Privacy on Device.

On disk#

The speech model is the main cost, and its size depends on which one you pick. Optional cleanup adds a second local model of about 1.1 GB, only if you turn it on.

Next#