比較

Superwhisper review: the best voice-to-text app for Mac in 2026
A detailed look at the AI dictation tool that turns your voice into polished, formatted text across every app on your Mac
Log in
Last updated October 2026
Superwhisper sits in your menu bar and waits. You press a hotkey, speak for as long as you like, release, and clean text appears wherever your cursor happens to be. That is the pitch, and for the most part it delivers. Built by SuperUltra, Inc., a bootstrapped company in Toronto, Superwhisper launched on Hacker News in August 2023 and has been shipping updates at a startling pace ever since, with more than forty releases in 2026 alone.
What makes it different from Apple's built-in dictation is that it does not just transcribe. It can rewrite, reformat, translate, and adapt your words to whatever context you are working in. Dictate a stream-of-consciousness ramble and get back a tidy email. Mumble a commit message and get clean lowercase text without punctuation. Speak meeting notes and receive bullet points. The app calls these "modes," and they are the reason most people end up paying for it.
This review covers everything you need to know before buying: how it works, what the free tier gives you, whether the lifetime deal is worth it, and where the app still falls short. If you have been thinking about switching from typing to talking, this is the place to start.
What Superwhisper is
The basics
Superwhisper is a system-wide dictation app that runs on macOS, Windows, and iOS. It records your voice through your Mac's microphone, transcribes it using AI models (either locally on your device or via cloud providers), and pastes the result into whatever app you are using. There is no dedicated writing window. The text goes straight into Slack, Notion, your email client, your terminal, a Google Doc, or any other app that accepts text input.
The default interaction is simple. Press Option-Space (or any hotkey you configure), hold it down while you speak, release it, and wait a beat. The transcription appears at your cursor. On Apple Silicon Macs with recent chips, the wait is barely noticeable. An M4 can process Whisper Large V3 at 17 to 27 times real-time speed, which means a ten-second recording is transcribed in well under a second.
You can also toggle recording on and off rather than holding a key, which is more comfortable for longer dictation sessions. The app supports push-to-talk, toggle recording, and per-mode shortcuts, so you can set up different keys for different tasks.
Who it's for
Writers, developers, and anyone who spends hours typing will get the most out of Superwhisper. The people who rave about it tend to be the ones who were already fast typists but wanted to go faster, or who deal with repetitive formatting tasks that a well-configured mode can handle automatically.
It is also a strong fit for multilingual professionals. The app supports over 100 languages through Whisper and 25 languages through NVIDIA's Parakeet V3 model, with automatic language detection. You can speak in one language and have the output translated to English, or set up separate modes for each language you work in.
People with RSI or other conditions that make sustained typing painful are another core audience. Speaking is simply easier on the body than typing, and the accuracy of modern transcription models means you lose very little in the translation from voice to text.
Features
Modes are the core of the app
Every Superwhisper workflow starts with a mode. A mode is a bundle of settings: which transcription model to use, whether to run the transcript through an AI for cleanup, what formatting rules to apply, which language to expect, and which apps to activate the mode for. The free tier gives you up to three custom modes. Pro unlocks unlimited modes.
The pre-built modes cover the obvious cases. Voice mode gives you a raw transcript. Message mode produces casual, short text suited to Slack or iMessage. Email mode generates more formal, structured output. Notes mode creates bullet points. Meeting mode records longer sessions and produces formatted summaries.
Custom modes are where things get interesting. You can create a mode that strips punctuation and lowercases everything for terminal commands. Another that detects your language and transcribes without translating. A third that reformats your words into a specific writing style. Each mode gets its own keyboard shortcut (added in version 2.12.0, April 2026), and modes can switch automatically based on which app you are using. Open Slack and your casual mode activates. Switch to Mail and it shifts to formal.
Super mode and AI rewriting
Super mode is the headline feature for Pro subscribers. It takes your raw transcript and sends it through an AI language model that rewrites it according to rules you define. This is not just grammar correction. It is a full rewrite that can change tone, restructure sentences, add formatting, and remove filler.
The practical effect is that you can speak messily and get polished output. You do not need to self-edit while talking. Say "um" and "you know" and trail off mid-sentence. The AI cleans it up. Filler word removal was made automatic in version 2.11.0, and version 2.18.0 introduced S1-mini, an on-device language model with tone control, so you can get some of this rewriting without sending anything to the cloud.
You choose which AI model does the rewriting. Superwhisper supports a wide range of LLMs through a bring-your-own-key (BYOK) system. You can use OpenAI's GPT-5 family, Anthropic's Claude (Sonnet, Haiku, Opus), Google's Gemini 3, xAI's Grok, Meta's Llama, and others. You pay the API provider directly for usage, and Superwhisper proxies the request so the provider does not see your account details.
On-device transcription and privacy
The privacy story is one of Superwhisper's strongest selling points. Every Whisper model can run entirely on your Mac. Audio never leaves your device. There is no internet connection required. For anyone working with sensitive material, legal documents, medical notes, or confidential business information, this matters.
The local model selection is generous. You get the full Whisper family (Fast, Nano, Standard, Pro, Ultra V3, and Ultra V3 Turbo), plus NVIDIA's Parakeet V2 (English-only) and Parakeet V3 (25 languages). As of version 2.18.2, all local Whisper models are free on macOS, which is a meaningful change from the earlier restriction that limited free users to smaller models.
On-device LLMs are available too. If you want AI rewriting without cloud access, you can run GPT OSS 20B, DeepSeek R1 Distill, Llama 3.x, Mistral 7B, or Phi-2 locally. The S1-mini model introduced in August 2026 is particularly notable because it was built specifically for Superwhisper's use case and includes tone control for messages, emails, and notes.
There are privacy caveats worth knowing about. Audio recordings are saved to disk by default, though you can configure retention settings (added in version 2.12.1). API keys for cloud services are stored in plaintext JSON in your Application Support folder rather than in macOS Keychain. Version 1.45.11 added encryption for custom model API keys on save, but the broader key storage concern remains. The app is closed-source, so you are relying on the vendor's claims about data handling.
Cloud transcription models
If you want the highest possible accuracy, the cloud models are worth considering. Superwhisper offers its own proprietary models, Ultra and S1-Voice, alongside third-party options like ElevenLabs Scribe V2 (added in version 2.10.0, made generally available in 2.11.0), Deepgram Nova 3, Nova 2, and Nova Medical, and the newer Cohere Transcribe (version 2.17.0, July 2026).
The S1 model family, introduced in August 2026, is Superwhisper's own cloud offering. S1-Voice handles transcription and S1-mini handles on-device language processing. These models were built specifically for dictation rather than adapted from general-purpose speech recognition, and the company has been iterating on them rapidly.
Realtime streaming arrived in version 2.10.0 (February 2026) and has been refined since. With streaming-capable models like ElevenLabs Scribe V2 realtime and Parakeet Realtime, you can see text appear as you speak rather than waiting until you stop recording. This makes longer dictation sessions feel more natural and helps you catch errors early.
Vocabulary and custom dictionary
Technical jargon, proper nouns, and unusual terminology are the bane of any dictation tool. Superwhisper addresses this with a custom vocabulary system that lets you add words, names, abbreviations, and phrases the transcription engine would otherwise get wrong.
The vocabulary system has been substantially improved in 2026. Version 2.16.0 added support for very large vocabulary lists. Version 2.16.5 introduced CSV import. Version 2.6.2 brought vocabulary support to Parakeet models. Forced alignment for offline Whisper vocabulary was added in 2.16.2. If you work in a field with specialised terminology, this is worth setting up properly.
Replacements work alongside vocabulary. You can define text substitutions that run after transcription, so abbreviations expand, formatting is corrected, and recurring phrases are standardised.
Developer and coding integration
Version 2.13.0 (April 2026) added integration with Claude Code and Open Code, allowing you to pipe voice transcripts directly into a terminal-based coding agent. This is macOS only and aimed at developers who want to describe code changes or ask questions verbally rather than typing commands.
The coding agent ecosystem has expanded quickly since then. Version 2.14.0 added direct coding agent plugin installation and Codex hook support. Version 2.17.0 brought Grok as a supported coding agent. Version 2.13.2 added Pi coding agent support. DeepSeek Harness was added as an agent plugin in version 2.18.2.
Shell command triggers let modes run scripts after transcription, which opens up automation possibilities beyond simple text insertion. You could, for example, have a mode that transcribes a command and then executes it.
History and file transcription
The rebuilt History feature (version 2.8.0, January 2026) is a quiet strength. It includes full-text search across all your past transcriptions, segmented playback so you can listen back to specific portions, and multi-select delete for housekeeping. History updates in real time and you can copy items or reveal them in Finder.
File transcription is also available. You can drag audio or video files into Superwhisper and get a transcript. Supported formats include standard audio files plus .mov video files (added in version 2.16.5) and .ogg and .opus formats (version 2.3.0). This is useful for transcribing recorded interviews, lectures, or meeting recordings that were not captured live through the app.
Pricing
Free tier
The free tier is more generous than it used to be. As of version 2.18.2, all local Whisper models are free on macOS, which means you get access to every model size from tiny to Ultra V3 Turbo without paying. You also get support for 100+ languages, up to three custom modes, meeting recording and transcription, and the iOS keyboard extension.
New users get a 3,000-word Pro trial that lets you test Super mode, cloud models, and advanced formatting before committing to a subscription. This replaced the earlier 15-minute time-limited trial, which was too short for most people to form a proper opinion.
The free tier is a legitimate daily driver for simple dictation. If all you need is accurate voice-to-text without AI rewriting, you can use Superwhisper indefinitely at no cost.
Pro monthly: $8.49 per month
Pro monthly unlocks everything. All transcription models (local and cloud), unlimited custom modes, cloud LLM post-processing through BYOK, priority support, and all platform apps. The monthly plan is the right choice if you want to try Pro features for a month or two before committing to a longer term.
Keep in mind that cloud AI rewriting has separate costs. When you use GPT-5, Claude, or Gemini for post-processing, you pay the API provider directly. For moderate use, this might add $5 to $15 per month depending on how heavily you lean on AI rewriting and which model you choose.
Pro annual: $84.99 per year
The annual plan works out to about $7.08 per month, a 17% saving over the monthly price. For anyone who has used the free tier or the monthly plan and knows they will stick with it, annual is the sensible choice. It includes the same features as monthly.
Lifetime: $249.99
The lifetime plan is the equivalent of roughly three years of annual billing. If you plan to use Superwhisper for longer than that, it pays for itself. The plan covers Mac, Windows, and iOS under a single purchase and includes all future updates.
Whether the lifetime deal is "worth it" depends on your confidence in the app's longevity. SuperUltra is a small, bootstrapped company. The development pace has been excellent, with dozens of releases in 2026, but bootstrapped companies carry risk that larger operations do not. A 30-day refund policy applies to all paid plans, so you can try it risk-free.
Enterprise
Enterprise pricing is custom and includes SOC 2 Type II compliance, team licensing, organisation-level configuration, and recording retention controls. If you need compliance guarantees beyond what the standard plans offer, this is the route.
Student discount
A 40% student discount on Pro plans is available, though it is not prominently advertised on the main pricing page. Check the Superwhisper website for eligibility requirements.
Pros and cons
What works well
Accuracy on Apple Silicon is excellent. On an M4 Mac, the Whisper Large V3 Turbo model delivers near-perfect accuracy at speeds that feel instant. The combination of on-device processing and modern Apple Silicon hardware means you rarely need to correct transcription errors.
The mode system is flexible and well-designed. Being able to set up different configurations for different apps and switch between them with keyboard shortcuts turns dictation from a novelty into a genuine workflow tool. The per-mode shortcut feature (April 2026) was a welcome addition.
Development velocity is impressive. More than 40 releases in 2026, covering new models, new integrations, UI improvements, and feature requests. For a bootstrapped company, this pace is remarkable and bodes well for the product's future.
The free tier is competitive. With all local Whisper models now free on macOS, the barrier to entry is essentially zero. You can test the core dictation experience thoroughly before spending anything.
Privacy options are strong. Full on-device transcription with no internet requirement is a meaningful differentiator. Most competing products require cloud processing.
The S1-mini on-device language model is a step change for people who want AI rewriting without cloud dependence. It is not as capable as GPT-5 or Claude, but for basic tone adjustment and cleanup, it works well enough that many users will not need cloud models at all.
Where it falls short
Setup takes time. Reviewers consistently mention 15 to 30 minutes of configuration to get modes, shortcuts, and models set up properly. The settings are dense and the number of options can be overwhelming for new users, though the onboarding was simplified in several 2026 updates.
The Windows version lags behind macOS. Users report crashes, freezes, and generally less polish. If you are primarily a Windows user, manage your expectations.
Audio recordings are saved to disk by default. You can now configure retention settings, but the default behaviour has surprised users who expected privacy-first defaults. Recordings can also sync to iCloud Drive if you have that enabled.
Cloud rewriting can flash raw text before the clean version. When using AI post-processing, there is a brief moment where the unprocessed transcript appears before being replaced by the rewritten version. This is a cosmetic issue, but it can be distracting.
The recording overlay is visually intrusive according to some users. Version 2.7.0 added recording window style selection, which helps, and the muffle option in version 2.19.0 lowers other apps' audio while recording, which is a nice touch, but the visual overlay remains a matter of taste.
Clipboard behaviour needs awareness. Superwhisper pastes text by writing to the clipboard and triggering a paste action, which means it overwrites whatever was previously on your clipboard. Version 2.18.3 made clipboard preservation the default, which addresses this for most use cases.
Who Superwhisper is best for
Writers and content creators who produce large volumes of text daily. The time saving from speaking rather than typing compounds quickly, and the mode system means your output arrives pre-formatted for its destination.
Developers who want to dictate code comments, commit messages, documentation, or interact with coding agents by voice. The Claude Code integration and shell command triggers open up workflows that most dictation tools cannot touch.
Multilingual professionals who switch between languages throughout the day. The automatic language detection and per-mode language settings make this smoother than any competing solution.
Knowledge workers who take meeting notes, write emails, and compose Slack messages all day. Having a mode for each context means you speak once and get appropriately formatted output every time.
If you use a tool like Fabric as your personal cloud and context layer, voice-to-text pairs well with it. You can dictate notes, thoughts, and ideas through Superwhisper and save them straight into Fabric, where they become part of your searchable, connected knowledge base. Fabric starts at $8 per month on the yearly plan ($10 monthly) with a free tier to get started, and the team plants a tree for every paying subscriber. Speaking a quick note into Superwhisper and dropping it into Fabric takes seconds and means you never lose a thought.
Who should look elsewhere
If you only need occasional, short dictation, Apple's built-in dictation on macOS is free and works well enough for casual use. The setup time and learning curve of Superwhisper are not justified if you dictate a few sentences a week.
If you primarily work on Android, Superwhisper's Android support is still early and not its strongest platform. Look at alternatives that were built for Android first.
If you need HIPAA compliance, the standard plans do not include it. You would need to explore the Enterprise tier and confirm compliance guarantees directly with SuperUltra.
If you want a simple, zero-configuration experience, the depth of Superwhisper's settings may feel like more than you need. Some people prefer tools that make fewer decisions available and just work out of the box.
Related comparisons
Frequently asked questions
Is Superwhisper free?
Yes, there is a permanent free tier. As of August 2026, all local Whisper models are free on macOS, giving you accurate voice-to-text in 100+ languages with up to three custom modes. New users also get a 3,000-word trial of Pro features including Super mode and cloud AI models. For basic dictation without AI rewriting, the free tier is a fully functional daily driver. Try it here.
Does Superwhisper work offline?
Yes. All local Whisper and Parakeet models run entirely on your device with no internet connection. This includes the full range of model sizes from tiny to Ultra V3 Turbo. Cloud models and BYOK LLM post-processing require an internet connection, but core transcription works completely offline.
Which Mac do I need?
Any Mac running macOS 14 or later. Apple Silicon (M1 or newer) is strongly recommended for local models. An M1 runs Whisper Large V3 at roughly real-time speed. An M4 runs the same model at 17 to 27 times real-time. Intel Macs can use cloud models but local model performance is poor. Macs with 8 GB of RAM should stick to the Large V3 Turbo model; 16 GB or more can run any model comfortably.
Does Superwhisper work on Windows?
Yes, but the Windows version is less mature than the Mac version. It supports Windows 10 and later on both x64 and ARM64, but users report more crashes and occasional freezes compared to the Mac experience. If you primarily use Windows, test it during the 30-day refund window before committing.
How accurate is the transcription?
On Apple Silicon Macs, accuracy is excellent. The Whisper Large V3 and V3 Turbo models achieve 99.2 to 100% accuracy in benchmarks, though real-world results depend on your accent, microphone quality, and background noise. The custom vocabulary feature helps with specialised terminology. Cloud models like ElevenLabs Scribe V2 and Superwhisper's own S1-Voice can push accuracy even higher.
What is Super mode?
Super mode sends your raw transcript through an AI language model that rewrites it according to rules you define. It can clean up grammar, change tone, add formatting, remove filler words, and restructure sentences. You choose which AI model to use (GPT-5, Claude, Gemini, and others are supported), and you pay the API provider directly for usage. Super mode requires a Pro subscription.
Can I use Superwhisper for meetings?
Yes. There is a dedicated meeting mode that records longer sessions and generates formatted notes. The app supports diarization (identifying different speakers) and can produce structured summaries. For quick internal meetings, it works well. For formal transcription of recorded proceedings, you might want a dedicated meeting tool.
Is there a lifetime deal?
Yes, the lifetime plan costs $249.99 and includes all Pro features, all future updates, and access across Mac, Windows, and iOS. It pays for itself after roughly three years compared to the annual plan. All purchases include a 30-day refund guarantee. Check the lifetime plan here.
How does Superwhisper compare to Apple Dictation?
Apple's built-in dictation is free, requires no setup, and works reasonably well for casual use. Superwhisper's advantages are accuracy (especially with specialised vocabulary), the mode system for contextual formatting, AI rewriting via Super mode, offline privacy with local models, support for 100+ languages, and the ability to customise every aspect of the dictation workflow. If you dictate frequently or need formatted output, Superwhisper is a significant step up. There is a detailed comparison here.
Does Superwhisper support my language?
Almost certainly. Whisper models support over 100 languages and dialects with automatic detection. Parakeet V3 covers 25 languages locally. You can set a specific language per mode or let the app detect automatically. Translation to English is available during transcription. The app interface itself is English only.
Can I integrate Superwhisper with coding tools?
Yes. Version 2.13.0 added integration with Claude Code and Open Code on macOS. You can speak commands or describe code changes and have them sent to a terminal-based coding agent. There is also support for shell command triggers that run scripts after transcription, Codex hooks, and several third-party coding agent plugins.
Is the student discount real?
Yes, there is a 40% student discount on Pro plans. It is not prominently listed on the main pricing page, so you may need to look for it on the Superwhisper website or contact their support to verify eligibility.
Is Fabric free?
Fabric has a free tier that lets you get started with saving and organising your content. Paid plans start at $8 per month on the yearly plan or $10 per month if you prefer to pay monthly. Fabric works well as a destination for voice-transcribed notes from Superwhisper: dictate a thought, save it to Fabric, and it becomes part of your searchable personal cloud. The team also runs a tree planting programme, planting a tree for every paying subscriber.
What happens if SuperUltra shuts down?
Local models would continue to work indefinitely since they run on your device with no server dependency. Cloud features, BYOK proxying, and the proprietary S1 models would stop working. This is a consideration for the lifetime plan in particular. The app's rapid development pace and growing user base suggest the company is in good health, but there are no guarantees with any small software company.
How do I get started with Superwhisper?
Download the app, choose a local model, and start talking. The free tier gives you everything you need for basic dictation. Spend 15 to 20 minutes setting up your first few modes and a custom vocabulary for your field. Once your modes are configured, the daily experience is simple: press your hotkey, speak, release, and your text appears. Get started with Superwhisper.
Compare similar apps and tools:
正在評估其他選擇?查看更多比較: