← igorjoshevski.com
Open source · MIT Windows & macOS Fully local

Meeting Transcriber

Real-time transcription of whatever's making noise on your machine — Teams, Zoom, or anything else — as a standalone CLI, not a plugin. Capture, voice-activity detection, and speech-to-text all run on-device. Take the notes automatically; stay focused on the meeting.

Audio In, Transcript Out

Capture and inference run on separate threads connected by a queue, so a slow transcription pass never drops or blocks audio capture.

1 · Capture

System Audio

WASAPI loopback on Windows, ScreenCaptureKit on macOS — no virtual audio driver needed.

2 · Normalise

16 kHz Mono

Downmixed and resampled to what VAD and ASR expect, regardless of the source device's native format.

3 · Detect

Silero VAD

Streams audio into speech segments, so silence and notification dings never reach the transcriber.

4 · Transcribe

whisper.cpp

Each segment is transcribed locally and flushed to stdout and a log file immediately.

Built to Just Run in the Background

A CLI you start and forget about, not a service to operate.

Nothing Leaves the Machine

Capture, VAD, and transcription all run locally — no audio or text is ever uploaded to a server.

Works With Any App

Not tied to Teams or Zoom specifically — it captures whatever's playing through the machine, or one running app in isolation on macOS.

Crash-Safe Transcripts

The log is flushed to disk after every line, so killing the process never loses a completed segment.

Right-Sized Models

Default model is quantized and under 60 MB, auto-downloaded on first use. Swap to a bigger or multilingual one with a flag, any time.

Cross-Platform, Not Cross-Compromise

Windows and macOS each get a native capture backend suited to that OS, behind the same command-line interface.

Debuggable in Isolation

Capture and VAD are runnable as standalone scripts, so a problem is easy to narrow down before it reaches transcription.

Why a CLI, Why Local

Most meeting-notes tools are a plugin bolted onto one video app, or a cloud service that needs your audio uploaded somewhere to be useful. Neither fit what I wanted: something that works no matter which app is making noise, keeps every byte of audio on the machine it was recorded on, and doesn't ask me to trust a third party with what was said in the meeting.

That pushed the design toward a plain CLI over threads and a queue rather than a packaged app: capture and transcription are independent enough that one can never stall the other, and each stage — capture, VAD, transcription — is small enough to test and debug on its own before wiring the rest around it.

Tech Snapshot
Language & Runtime
Python 3.11, plain python -m transcribe — no packaging step
Speech-to-Text
whisper.cpp via pywhispercpp, default base.en-q5_1
Voice Activity Detection
Silero VAD, streaming audio into speech segments
Windows Capture
WASAPI loopback via PyAudioWPatch
macOS Capture
ScreenCaptureKit audio capture, no virtual audio driver
License
MIT, source on GitHub

Try It or Read the Source

The full source, setup instructions, and CLI flags are on GitHub. Happy to talk through the architecture or a feature you'd want added.