A local-first audio transcription toolkit. Download, transcribe, and polish โ
from YouTube video to publication-ready transcript in three commands.
All speech recognition runs on your own GPU; audio never leaves your machine.
View source on GitHub โ
Any video or audio file
GPU-accelerated STT
LLM API formatting
Clean, readable text
All speech recognition runs on your GPU. Only the text formatting step optionally calls an external API.
Process entire folders of audio files with checkpoint support for pause and resume.
Best with 8GB+ VRAM. Falls back to int8 for smaller GPUs, or CPU mode if no NVIDIA card.
Watch tokens arrive in real time with streaming mode, or wait for the complete result.