Meeting Transcriber
Real-time transcription of whatever's making noise on your machine — Teams, Zoom, or anything else — as a standalone CLI, not a plugin. Capture, voice-activity detection, and speech-to-text all run on-device. Take the notes automatically; stay focused on the meeting.
Audio In, Transcript Out
Capture and inference run on separate threads connected by a queue, so a slow transcription pass never drops or blocks audio capture.
System Audio
WASAPI loopback on Windows, ScreenCaptureKit on macOS — no virtual audio driver needed.
16 kHz Mono
Downmixed and resampled to what VAD and ASR expect, regardless of the source device's native format.
Silero VAD
Streams audio into speech segments, so silence and notification dings never reach the transcriber.
whisper.cpp
Each segment is transcribed locally and flushed to stdout and a log file immediately.
Built to Just Run in the Background
A CLI you start and forget about, not a service to operate.
Nothing Leaves the Machine
Capture, VAD, and transcription all run locally — no audio or text is ever uploaded to a server.
Works With Any App
Not tied to Teams or Zoom specifically — it captures whatever's playing through the machine, or one running app in isolation on macOS.
Crash-Safe Transcripts
The log is flushed to disk after every line, so killing the process never loses a completed segment.
Right-Sized Models
Default model is quantized and under 60 MB, auto-downloaded on first use. Swap to a bigger or multilingual one with a flag, any time.
Cross-Platform, Not Cross-Compromise
Windows and macOS each get a native capture backend suited to that OS, behind the same command-line interface.
Debuggable in Isolation
Capture and VAD are runnable as standalone scripts, so a problem is easy to narrow down before it reaches transcription.
Why a CLI, Why Local
Most meeting-notes tools are a plugin bolted onto one video app, or a cloud service that needs your audio uploaded somewhere to be useful. Neither fit what I wanted: something that works no matter which app is making noise, keeps every byte of audio on the machine it was recorded on, and doesn't ask me to trust a third party with what was said in the meeting.
That pushed the design toward a plain CLI over threads and a queue rather than a packaged app: capture and transcription are independent enough that one can never stall the other, and each stage — capture, VAD, transcription — is small enough to test and debug on its own before wiring the rest around it.
python -m transcribe — no packaging stepbase.en-q5_1Try It or Read the Source
The full source, setup instructions, and CLI flags are on GitHub. Happy to talk through the architecture or a feature you'd want added.