Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
-
Updated
Aug 17, 2026 - Python
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
Open-source SenseVoiceSmall model for Mandarin, Cantonese, English, Japanese, and Korean ASR, language ID, emotion recognition, and audio event detection.
FunASR-powered video transcription, subtitle generation, and LLM-assisted clipping tool with a local Gradio UI.
Open-source LLM-based ASR model family for Chinese, dialect, accent, and multilingual speech, with FunASR, vLLM, streaming, and llama.cpp runtimes.
Terminal voice-to-text TUI — Qwen3-ASR-1.7B on the Apple GPU via MLX (mlx-speech). Fully local, no PyTorch, transcribes in ~1s. macOS Apple Silicon.
听记 (Tingji) — 本地会议录音转写与纪要,FunASR + LLM,数据不出本机。Local meeting transcription & minutes, runs entirely offline.
Local voice-to-text for macOS and iOS. Multilingual (EN/ZH/JP) with Traditional Chinese output. Runs Qwen3-ASR on Apple Silicon via MLX. No cloud, no subscription.
Talk. Ink. Push-to-talk dictation for macOS, 100% on-device. Pick your model: Qwen3-ASR, NVIDIA Nemotron or Voxtral, all via Apple MLX.
Voice dictation for the browser — free, private, MIT. Dictate into any web page, or use the pop-out to dictate for any app on your machine.
OpenAI-compatible speech-to-text server for nvidia/nemotron-3.5-asr-streaming-0.6b (NeMo). Runs on the DGX Spark / GB10.
Sono is on-device dictation for macOS. Parakeet v3 + Apple Intelligence, nothing leaves your Mac.
Free, open-source, fully-local dictation for macOS. Hold a key, speak, release: clean text at your cursor. A Wispr Flow alternative that never touches the cloud.
🎬 面向内容创作者的 AI 字幕工作流助手,基于 Gemini 2.5 Pro。独创“音频理解 + 术语人工确认”两阶段工作流,从源头解决技术口播与专有名词识别翻车问题,支持一键导出 SRT/VTT/ASS。
Live speech-to-text streaming on Apple Silicon — Qwen3-ASR + Silero VAD + MLX
VoxMinutes - free, local-first meeting assistant for Windows. Records system audio & mic together; real-time transcription, translation (13 languages) & AI summaries. Data never leaves your device.
Free Wispr Flow & Superwhisper alternative. Native macOS voice dictation & speech-to-text. Powered by Sber GigaAM v3 & Whisper. Instant direct input under cursor via ⌥+Space, push-to-talk, offline transcription of any media files.
Nói — nói ra chữ trên Mac bằng tài khoản ChatGPT của bạn. Miễn phí, mã nguồn mở.
Private, local-first meeting recorder + transcription, diarization, AI notes, voice dictation & read-aloud for Windows — runs on your own GPU.
🎬 AI subtitle generator: convert video to SRT subtitles locally with NVIDIA NeMo Parakeet-TDT speech-to-text. GPU-accelerated, word-level timestamps, VAD, LLM correction — a fast offline Whisper alternative.
Transcribe any audio file to text locally on an Apple Silicon Mac — free, private, no cloud APIs. Powered by NVIDIA Parakeet + Apple MLX.
Add a description, image, and links to the whisper-alternative topic page so that developers can more easily learn about it.
To associate your repository with the whisper-alternative topic, visit your repo's landing page and select "manage topics."