Skip to content

Repository files navigation

Local voice dictation. Hold a hotkey, speak, release — your words are transcribed on your own machine, cleaned up by a local AI pass (filler words removed, self-corrections resolved, grammar fixed), and typed into whatever app has focus. No cloud, no accounts, nothing leaves your PC. A free, local-first alternative to tools like Wispr Flow — with one thing most free/open dictation tools skip entirely: **the AI cleanup pass**. Most alternatives just hand you a raw transcript; Dictator fixes it before it lands in your text box. > **v1.1.** Core dictation is my own daily driver on Windows. This release > is about speed and typing reliability — text lands roughly 200 ms after > you let go of the hotkey. macOS support hasn't been run on real Mac > hardware yet — the fundamentals work. [Open an issue](../../issues) if > something breaks or you want a feature. ## Where it fits | | Cloud dictation apps (Wispr Flow, etc.) | Other open-source dictation tools | Dictator | |---|---|---|---| | Cost | Subscription | Free | Free | | Audio leaves your PC | Yes | No | No | | Cleans filler words / self-corrections | Yes | Usually not — raw transcript only | Yes, local | | Works offline | No | Yes | Yes | | App-aware tone (casual/formal/verbatim) | Rare | No | Yes | | Release-to-text latency | ~0.7–2 s (network round trip) | Usually 1 s+ (decodes after release) | ~0.2 s typical, ~0.45 s p95 | The gap this fills: free and open dictation tools exist, but almost all of them stop at the raw Whisper transcript. Dictator adds the cleanup pass — the part that actually makes a transcript usable — without sending anything to a server to do it. ## Install **Windows** — PowerShell: ```powershell irm https://raw.githubusercontent.com/IntellectDaksh/Dictator/main/scripts/bootstrap.ps1 | iex ``` **macOS** — Terminal: ```bash curl -fsSL https://raw.githubusercontent.com/IntellectDaksh/Dictator/main/scripts/bootstrap.sh | bash ``` Either one-liner clones the repo, creates a virtual environment, installs every dependency, installs Ollama if it's missing, **tunes Dictator to your hardware** (see below), pulls the matching cleanup model, then launches the app. No prompts. Safe to re-run any time; every step skips if it's already done. Requires [Git](https://git-scm.com/downloads) and Python 3.11+ (the Windows script installs Git and Ollama via `winget` if missing; the macOS script installs Homebrew, Python and Ollama if missing). ## Tuned to your machine The installer runs `hwtune.py`, which reads your RAM, GPU and CPU and picks the models that keep dictation around the same speed on any laptop — bigger models where the hardware has room, smaller ones where it doesn't: | Your hardware | Speech model (Whisper) | Cleanup model (Ollama) | |---|---|---| | NVIDIA GPU, 6 GB+ VRAM | `small.en`, float16 on GPU | `qwen2.5:3b-instruct` (~1.9 GB) | | NVIDIA GPU, 4–6 GB | `small.en`, float16 on GPU | `qwen2.5:1.5b-instruct` (~1 GB) | | NVIDIA GPU, 2–4 GB | `small.en`, int8 on GPU | `qwen2.5:1.5b-instruct` | | Apple Silicon, 16 GB+ | `base.en`, int8 on CPU | `qwen2.5:3b-instruct` | | No GPU, 12 GB+ RAM | `base.en`, int8 on CPU | `qwen2.5:1.5b-instruct` | | No GPU, under 12 GB | `base.en`, int8 on CPU | `qwen2.5:0.5b-instruct` (~0.4 GB) | It writes the choice into your config, starts Ollama, pulls the model if you don't have it, and links it — nothing to set by hand. Dictator itself never re-tunes at startup, so anything you change later in the tray menu sticks. Upgraded your hardware? Re-run it: ```bash python hwtune.py # show what it detects and would pick python hwtune.py --pull # apply it and pull the model ``` After setup: double-click `Dictator.bat` (Windows) or `Dictator.command` (macOS) to start it again. **macOS note:** the first time you dictate, macOS will prompt for Accessibility + Input Monitoring permission (needed to detect the hotkey and type text) — grant it in System Settings → Privacy & Security and try again. macOS support is new and best-effort; see `info.md` for exactly what's platform-specific if something misbehaves. ## Why this exists I dictate most of my own notes, DMs, and scripts for Uideas — typing is slower than talking, but every dictation tool I tried either shipped audio to a cloud API or left "um"s and false starts in the transcript. Built this to run fully local and clean up the mess a real voice makes, then kept using it daily until it stopped breaking. ## How to use it 1. Click into any text box. 2. Hold **Ctrl + Win** (Ctrl + Cmd on macOS) and speak — the status bar turns red. 3. Release — it turns amber while cleaning up, flashes green when your text is typed. Say it messy: "um let's meet at 12 no wait 11" becomes "Let's meet at 11." (Self-corrections like that need **Smart** cleanup — see below.) ## Cleanup modes Pick one from the tray menu (right-click the mic icon → Processing mode). | Mode | What it does | Speed | Extra memory | |---|---|---|---| | **Fast** (default) | Punctuation, capitals, filler words removed — no LLM | Instant | None | | **Smart** | Full LLM pass: self-corrections, grammar, tone | ~0.5 s more | Small Ollama model, unloaded after 2 min idle | | **Verbatim** | Exactly what Whisper heard | Instant | None | Smart mode can't invent text: output with words you never said, or that balloons in length, is rejected and the plain transcript is typed instead. ## Why it's fast - The mic stream stays open, so the first word is never clipped (0.3 s of pre-roll is kept on press — nothing outside a dictation is stored). - Whisper stays loaded, and finished sentences are transcribed *while* you're still talking — release only has to decode the last few words. - The GPU is kept out of its idle clock during recording. - Your custom vocabulary is passed to Whisper itself, not just the cleanup step, so names and brand words come out right the first time. Every dictation logs stage timings (never text) to `latency.jsonl` in the config folder, so slowdowns are visible. ## Light on memory Built to sit in the tray all day on a laptop that's already running a browser and an editor: - **~340 MB RAM** for the app with Whisper loaded (measured on Windows, `small.en` on GPU). On GPU machines the model weights live in VRAM, not RAM. - **No LLM in memory by default.** Fast mode never touches Ollama. Smart mode loads the small cleanup model on demand and Ollama drops it after 2 minutes idle, so it only costs memory while you're actively dictating. - **Small models first.** 0.5B–3B cleanup models instead of the 7B–14B most local setups default to — the cleanup job is short and structured, so a small model does it in a fraction of the time and memory. - **Tight context.** Cleanup requests use a 512-token window and cap the output length, so Ollama doesn't reserve memory it will never use. - **Memory trimmed after each dictation** — freed audio buffers are handed back to the OS instead of sitting in the process. - Audio is never written to disk, and only 0.3 s of it is buffered outside a dictation. ## Features - Hold-to-dictate or hands-free (double-tap for hold mode, single-tap toggle), with optional silence auto-stop - Configurable hotkey — pick a preset, capture any combo live, or save named hotkey profiles - Local speech-to-text (`faster-whisper`), GPU-accelerated on Windows with CPU fallback everywhere - Local cleanup LLM via Ollama — strips filler words, fixes grammar, resolves self-corrections, falls back to raw transcript if Ollama is unreachable - Language selection — 12 languages plus auto-detect - App-aware tone (casual/formal/verbatim per focused app, user-editable) and spoken tone overrides ("...make it formal") - Voice commands: "new line", "new paragraph", "bullet point" — newlines are typed as Shift+Enter so chat boxes don't send early - Safe typing — waits until you've let go of every modifier, and stops the moment focus moves to another window, so keystrokes never turn into shortcuts or land in the wrong app - Snippets/macros, custom vocabulary for names/brand words - Redaction list — sensitive words/phrases scrubbed before they're ever written to disk - Dictation history + an Insights tab (streak calendar, app usage, WPM, tone distribution, hourly activity) — opt-in logging, off by default - Dark/light theme with independently configurable accent and highlight colors, start-on-login ## Tray menu (right-click the mic icon) Enable/disable, pick microphone, pick cleanup mode (Fast/Smart/Verbatim), pick Whisper model size (base/small/medium), toggle status bar, toggle history logging, start on login, open config folder, quit. ## FAQ **Does any audio or text leave my machine?** No. Speech-to-text runs locally via `faster-whisper`, cleanup runs locally via Ollama on `localhost`. The only network calls are to Ollama itself and, during install, pulling the models. See [Privacy](#privacy) below. **Will it be as fast on my laptop as on yours?** That's what the hardware tuning is for: slower machines get lighter models so the wait after you release stays short. A laptop without a GPU uses the smaller `base.en` speech model, which is a little less accurate on unusual words than `small.en` — add those to your custom vocabulary. **Do I need a GPU?** No — CPU works fine, GPU (CUDA, Windows only) just makes transcription faster. `faster-whisper` picks whichever is available. **What if Ollama isn't running or the cleanup model isn't pulled?** Dictator falls back to the raw (optionally auto-punctuated) transcript rather than failing silently. You never lose a dictation because the cleanup step had a bad moment. **Can I use my own cleanup model instead of the default?** Yes — any model pulled in Ollama works; set it from Settings. **Windows or macOS — which is more solid right now?** Windows. It's the daily driver this was built for. macOS support is real (every platform-specific call is branched and implemented) but hasn't been run on physical Mac hardware yet — see the note in [Install](#install). ## Troubleshooting See [docs/TROUBLESHOOTING.md](docs/TROUBLESHOOTING.md) for common issues and the config file reference. ## Architecture See [docs/ARCHITECTURE.md](docs/ARCHITECTURE.md) for the stack, code map, and threading model — read that before sending a PR. `info.md` has the deeper, line-number-level map if you're picking this up cold. ## Privacy Audio lives in memory only and is discarded after transcription. The mic stream stays open while the app runs (that's what removes the start-up lag), so Windows shows the mic as in use — audio outside a dictation is dropped immediately, never buffered past 0.3 s. Turn it off with `keep_mic_warm: false` in the config if you'd rather trade the lag for the indicator. Clipboard is never touched — text is typed via simulated keystrokes. The only network traffic is to Ollama on `localhost`, plus one-time model downloads during setup. No telemetry, no accounts, no API keys. ## License MIT + Commons Clause — free to use, modify, and share; not for resale. See [LICENSE](LICENSE). # Update # Another update

About

Free, local-first voice dictation for Windows and macOS — on-device Whisper transcription, local LLM cleanup pass, hotkey-driven, no cloud, no accounts

Topics

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages