CrispASR

One C++ binary, 119 backends — 62 of them TTS engines — plus multilingual text translation, zero Python dependencies.

CrispASR started as a fork of whisper.cpp and extends that base into a unified speech engine called crispasr, backed by full ggml C++ runtimes for major open-weights ASR and TTS architectures. One build, one binary, one consistent CLI — pick the backend at the command line or let CrispASR auto-detect it from your GGUF file. See Text-to-Speech for the TTS side.

$ crispasr -m ggml-base.en.bin          -f samples/jfk.wav                    # OpenAI Whisper
$ crispasr -m parakeet-tdt-0.6b.gguf    -f samples/jfk.wav                    # NVIDIA Parakeet
$ crispasr -m canary-1b-v2.gguf         -f samples/jfk.wav                    # NVIDIA Canary
$ crispasr -m voxtral-mini-3b-2507.gguf -f samples/jfk.wav                    # Mistral Voxtral
$ crispasr --backend qwen3 -m auto      -f samples/jfk.wav                    # -m auto downloads
$ crispasr --backend kokoro -m auto --tts "Hello world" --tts-output out.wav  # TTS

No Python. No PyTorch. No separate per-model binary. No pip install. Just one C++ binary and a GGUF file.

Browser: All backends compile to WebAssembly (4.3 MB) via build-wasm.sh. Multithreaded, runs entirely client-side with COOP/COEP headers.

Demo: HuggingFace Space — live transcription + TTS + language detection, auto-deployed from hf-space/.

Ecosystem

Project What it does
CrispASR This repo — C++ speech engine. 119 backends (62 TTS), CLI + HTTP server + C-ABI + Python/Rust/Dart/Go/Ruby/Java bindings.
CrisperWeaver Cross-platform Flutter transcription app built on CrispASR. Desktop + mobile, model browser with download queue, mic capture, SRT/VTT/JSON export, diarization, batch processing. Fully offline.
CrispEmbed Text-related engine via ggml — same philosophy as CrispASR but for embeddings, retrieval, OCR and OMR, Math and Music Notation. Numerous architectures (XLM-R, Qwen3-Embed, Gemma3, ModernBERT, …), dense + sparse + ColBERT + reranking. PP-OCR, Tesseract, EasyOCR, InternVL2, etc. Python/Rust/Dart bindings.
Susurrus Python ASR GUI with 9 backends (faster-whisper, mlx-whisper, voxtral, insanely-fast-whisper, …). The Python counterpart to CrispASR’s C++ approach.

Table of contents


Start here

Everything below this section is a catalogue — 100+ backends, browse it when you need one. If you just want CrispASR working, this is the whole path. No repo clone, no Python, no model hunting.

New to the project? docs/getting-started.md walks the same path step by step, with the expected output and the three usual first-run failures.

1. Get the binary

Download one file from Releases and unzip it:

Platform Download Notes
Windows crispasr-windows-x86_64-cpu.zip Needs AVX2 (2013+ Intel / 2015+ AMD). Older CPU → …-cpu-legacy.zip
Windows + NVIDIA crispasr-windows-x86_64-cuda.zip Self-contained; a CUDA Toolkit install is not required. CUDA-13-native build: …-cuda13.zip (Turing+)
macOS crispasr-macos.tar.gz Metal GPU support built in
Linux crispasr-linux-x86_64.tar.gz …-cuda.tar.gz / …-vulkan.tar.gz for GPU

Prefer to build it yourself? See Install & build. The -hip and -vulkan builds require the matching driver and do not fall back to CPU; the Linux -cuda tarballs do fall back.

Check it runs — this should print a version banner and exit:

crispasr --version          # Windows: .\crispasr.exe --version

CUDA builds also print cuda toolkit and cuda runtime ABI, so this command distinguishes the CUDA 12 and CUDA 13 packages without inspecting DLLs.

2. Make it speak

-m auto downloads the model on first use (~135 MB here) and reuses it afterwards — nothing to find or install. It lands in ~/.cache/crispasr/ (%USERPROFILE%\.cache\crispasr on Windows).

crispasr --backend kokoro -m auto --tts "The quick brown fox jumps over the lazy dog." --tts-output hello.wav
# crispasr: TTS output written to 'hello.wav' (78000 samples @ 24000 Hz, 3.25 sec)

Play hello.wav. That is the TTS half working.

3. Transcribe it back

crispasr --backend parakeet -m auto -f hello.wav -l en
# crispasr: transcribed 3.2s audio in 0.32s (10.1x realtime)
# The quick brown fox jumps over the lazy dog.

(~467 MB on first run. -l en skips language auto-detection, which would otherwise fetch a small extra model.) Both halves now work — swap in your own .wav and you are running.

Where to go next

You want to… Go to
Clone a voice from a recording docs/tts.md — and read the consent rules first; cloning requires --i-have-rights
Pick a better ASR model Which backend should I pick?
Transcribe with word timestamps, SRT/VTT docs/cli.md
Live mic / streaming docs/streaming.md
Run it as an HTTP server docs/server.md
See every backend this binary has crispasr --list-backends

If nothing happened

Add -v to any command for verbose progress, and --dry-run-resolve to print which model files it would open (and whether they’re on disk) without loading anything.

If a command printed its banner and then simply stopped — no error, no output file — that is a crash, not a refusal, and the exit code identifies it in one step. See docs/troubleshooting.md.


Supported backends

CrispASR ships 119 backends (see the generated feature matrix for the authoritative, auto-counted list) — the majority for transcription/translation, and 62 TTS engines for synthesis. It also ships audio-to-audio S2S backends, including Sidon restoration and the VoxCPM2 AudioVAE speech upscaler; see the feature matrix for the complete capability list. Pick at the CLI with --backend NAME, or omit it to let the binary auto-detect from the GGUF metadata. Jump to the TTS table for the synthesis side.

ASR backends

Backend Model Architecture Languages License
whisper ggml-base.en.bin and all OpenAI Whisper variants Encoder-decoder transformer 99 MIT
whisper distil-whisper/distil-large-v3 Distilled Whisper: 32L encoder + 2L decoder (6.3x faster) English MIT
parakeet nvidia/parakeet-tdt-0.6b-v3 FastConformer + TDT 25 EU (auto-detect) CC-BY-4.0
parakeet nvidia/parakeet-tdt-0.6b-v2 FastConformer + TDT, original Open ASR Leaderboard topper en (mixed-case + punct) CC-BY-4.0
parakeet nvidia/parakeet-tdt-1.1b 42L FastConformer + TDT, larger English variant en (lowercase) CC-BY-4.0
parakeet nvidia/parakeet-tdt_ctc-110m 17L FastConformer + TDT+CTC hybrid; smallest variant, auto-CTC decode en CC-BY-4.0
parakeet nvidia/parakeet-tdt_ctc-1.1b 42L FastConformer + TDT+CTC hybrid; largest, mixed-case + punct en CC-BY-4.0
parakeet nvidia/parakeet-tdt_ctc-0.6b-ja FastConformer-TDT-CTC, xscaling, 80 mels Japanese CC-BY-4.0
reazonspeech reazon-research/reazonspeech-nemo-v2 FastConformer-RNNT, local attn (w=256), 80 mels, 619M params Japanese Apache-2.0
fastconformer-ctc nvidia/parakeet-ctc-0.6b 24L FastConformer + CTC, 80 mels (same arch as fc-ctc-xlarge) en CC-BY-4.0
fastconformer-ctc nvidia/parakeet-ctc-1.1b 42L FastConformer + CTC, 80 mels en CC-BY-4.0
fastconformer-ctc grider-transwithai/parakeet-ctc-1.1b-ja 42L FastConformer + CTC, 80 mels, Japanese fine-tune Japanese Apache-2.0
canary nvidia/canary-1b-v2 FastConformer + Transformer decoder 25 EU (explicit -sl/-tl) CC-BY-4.0
canary-qwen nvidia/canary-qwen-2.5b FastConformer + Qwen3-1.7B SALM en CC-BY-4.0
lfm2-audio LiquidAI/LFM2.5-Audio-1.5B FastConformer + LFM2 hybrid conv+attention backbone (ASR+TTS) en LFM Open v1.0
lfm2-audio LiquidAI/LFM2.5-Audio-1.5B-JP FastConformer + LFM2 hybrid conv+attention backbone (ASR+TTS) ja LFM Open v1.0
mini-omni2 gpt-omni/mini-omni2 Whisper-small + Qwen2-0.5B (ASR+TTS+S2S) en MIT
cohere CohereLabs/cohere-transcribe-03-2026 Conformer + Transformer 13 Apache-2.0
cohere efwkjn/cohere-asr-ja-v0.1 Japanese fine-tune of cohere-transcribe-03-2026 (TedX/JSUT-tuned) Japanese Apache-2.0
granite ibm-granite/granite-speech-{3.2-8b,3.3-2b,3.3-8b}, granite-4.0-1b-speech Conformer + Q-Former + Granite LLM (μP) (more) en fr de es pt ja Apache-2.0
granite-4.1 ibm-granite/granite-speech-4.1-2b 16L Conformer + Q-Former + Granite LLM; single ggml graph (more) en fr de es pt ja Apache-2.0
granite-4.1-plus ibm-granite/granite-speech-4.1-2b-plus 4.1 + hidden-state concat; punctuated output (more) en fr de es pt Apache-2.0
granite-4.1-nar ibm-granite/granite-speech-4.1-2b-nar Non-autoregressive: single LLM forward + slot argmax (more) en fr de es pt Apache-2.0
fastconformer-ctc nvidia/stt_en_fastconformer_ctc_large FastConformer + CTC (NeMo family, all sizes) en CC-BY-4.0
voxtral mistralai/Voxtral-Mini-3B-2507 Whisper encoder + Mistral 3B LLM 8 Apache-2.0
voxtral4b mistralai/Voxtral-Mini-4B-Realtime-2602 Causal encoder + 3.4B LLM, sliding window 13, realtime streaming Apache-2.0
qwen3 Qwen/Qwen3-ASR-0.6B Whisper-style audio encoder + Qwen3 0.6B LLM 30 + 22 Chinese dialects Apache-2.0
qwen3-1.7b Qwen/Qwen3-ASR-1.7B Whisper-style audio encoder + Qwen3 1.7B LLM 30 + 22 Chinese dialects Apache-2.0
qwen3-ja-anime jaykwok/Qwen3-ASR-1.7B-JA-Anime-Galgame-hf Qwen3-ASR-1.7B fine-tuned for Japanese anime/galgame speech ja + 30 langs Apache-2.0
mega-asr zhifeixie/Mega-ASR Qwen3-ASR-1.7B + merged robustness LoRA; always-on robust path noisy / degraded speech Apache-2.0
higgs-stt bosonai/higgs-audio-v3-stt Whisper-large-v3 encoder (4 s chunked) + Qwen3-1.7B LLM (more) en Apache-2.0
wav2vec2 jonatasgrosman/wav2vec2-large-xlsr-53-english CNN + 24L transformer + CTC head (any Wav2Vec2ForCTC) per-model Apache-2.0
wav2vec2 facebook/data2vec-audio-base-960h Data2Vec Audio (79 MB Q4_K) English Apache-2.0
wav2vec2 facebook/hubert-large-ls960-ft HuBERT Large (212 MB Q4_K) English Apache-2.0
glm-asr zai-org/GLM-ASR-Nano-2512 Whisper encoder + 4-frame projector + Llama 1.5B (GQA) Mandarin (+ Chinese dialects), English, Cantonese MIT
kyutai-stt kyutai/stt-1b-en_fr Mimi codec (SEANet + RVQ) + 16L causal LM en, fr MIT
kyutai-stt kyutai/stt-2.6b-en Mimi codec + 48L causal LM (2.6B, English-only; 3.5 s lookahead) en MIT
firered-asr FireRedTeam/FireRedASR2-AED Conformer + CTC + beam search; also LID (120 langs) Mandarin, English, 20+ Chinese dialects Apache-2.0
moonshine UsefulSensors/moonshine-{tiny,base} Conv + 6L enc + 6L dec; multilingual variants English + 6 langs MIT
moonshine‑de fidoriel/moonshine-base-de German fine-tune of moonshine-base (6.9% WER CV22) German CC‑BY‑NC‑SA‑4.0
moonshine‑tiny‑de fidoriel/moonshine-tiny-de German fine-tune of moonshine-tiny (11.4% WER CV22) German CC‑BY‑NC‑SA‑4.0
moonshine-streaming UsefulSensors/moonshine-streaming-{tiny,small,medium} Streaming: sliding-window encoder + AR decoder (34–245M) English MIT
gemma4-e2b google/gemma-4-E2B-it USM Conformer 12L + Gemma4 LLM 35L (GQA, PLE) 140+ langs Apache-2.0
gemma4-e4b google/gemma-4-E4B-it Same USM Conformer 12L + larger Gemma4 LLM 42L (GQA, PLE); runs on --backend gemma4-e2b 140+ langs Apache-2.0
omniasr omniASR-CTC-1B-v2 wav2vec2 CNN + 48L transformer + CTC (more) 1600+ Apache-2.0
omniasr‑300m omniASR-CTC-300M-v2 Same arch, 24L, ~194 MB Q4_K; auto-chunks >7 s (more) 1600+ Apache-2.0
omniasr-llm omniASR-LLM-300M-v2 Same encoder + 12L LLaMA decoder (more) 1600+ Apache-2.0
omniasr-llm omniASR-LLM-Unlimited-300M-v2 Streaming: 15s segment protocol, unlimited audio (more) 1600+ Apache-2.0
vibevoice microsoft/VibeVoice-ASR σ-VAE ConvNeXt + Qwen2.5-7B (more) 50+ MIT
vibevoice-streaming microsoft/VibeVoice-ASR-Streaming-1.5B σ-VAE ConvNeXt + Qwen2.5-1.5B; persistent-KV 2.93 s chunks with 0.53 s lookahead (more) multilingual MIT
vibevoice-bitnet VibeVoice-ASR-BitNet Same arch, TQ2_0 ternary LM (1.6 GB) (more) 7+ MIT
mimo-asr XiaomiMiMo/MiMo-V2.5-ASR 6L transformer + 36L Qwen2 LM + RVQ codec (more) Mandarin + dialects + English MIT
ark-asr ⚠️experimental/WIP cstr/ark-asr-3b-GGUF (base AutoArk-AI/ARK-ASR-3B) Whisper-large-v3 enc (partial RoPE) + Qwen2.5-3B LM (more) 19 (zh, en, de, ja, fr, ko, es, pl, it, ro, hu, cs, nl, fi, hr, sk, sl, et, lt) see base
moss-audio OpenMOSS-Team/MOSS-Audio-4B-Instruct 32L Whisper encoder + DeepStack 3-tap + 36L Qwen3 LM; audio understanding + ASR (more) zh, en Apache-2.0
moss-transcribe OpenMOSS-Team/MOSS-Transcribe-preview-2B Qwen3-Omni audio encoder (32L, windowed attn) + GatedMLP adapter + Qwen3-1.7B LM; ASR (more) zh, en Apache-2.0
moss-diarize OpenMOSS-Team/MOSS-Transcribe-Diarize-0.9B Stock Whisper encoder (24L, 80 mel) + 4x merge + VQAdaptor + Qwen3-0.6B LM; joint ASR + speaker diarization + timestamps multi Apache-2.0
whisper (tiron) ⚠️experimental Trelis/tiron (base Trelis/tiron) Whisper large-v3 with an extended vocab: emits inline <|speakerN|> markers + 20 ms timestamps for joint transcription + per-window speaker attribution, via a constrained-decoding grammar; cross-window linking to stable speakers (#295) multi (en focus) Apache-2.0
funasr FunAudioLLM/Fun-ASR-Nano-2512 70-block SANM encoder + 2-block Transformer adaptor + Qwen3-0.6B LLM zh, yue, en, ja, ko FunASR Model License v1.1 (commercial OK w/ attribution)
fun-asr-mlt-nano FunAudioLLM/Fun-ASR-MLT-Nano-2512 Same architecture, multilingual decoder 31 langs incl. de, fr, es, pt, ru, ar, hi, vi, th, ko FunASR Model License v1.1
paraformer funasr/paraformer-zh 50-block SANM encoder + CIF predictor + 16-block NAR decoder (single-pass, non-autoregressive); character-level vocab (8404); 220M params zh, en FunASR Model License (commercial OK w/ attribution)
foxnose (speaker diarization) Wespeaker/wespeaker-voxceleb-resnet34-LM Speaker diarization via --diarize-method foxnose: WeSpeaker ResNet34-LM 256-d embeddings + GMM/BIC speaker counting + spectral clustering + Viterbi temporal smoothing (more). 3.18 % DER on VoxConverse dev vs the upstream reference implementation’s 3.07 % any weights CC-BY-4.0
gigaam ai-sage/GigaAM-v3 (base ai-sage/GigaAM-v3) 16-layer rotary Conformer (220M) + CTC or RNN-T head; four revisions — e2e_rnnt / e2e_ctc emit punctuation + casing + ITN from a SentencePiece vocab, rnnt / ctc emit bare lowercase Cyrillic (more) en ru MIT
sensevoice FunAudioLLM/SenseVoiceSmall 70-block SANM encoder + CTC head; emits transcript + language ID + audio-event in one forward pass (non-AR, 15× faster than Whisper-Large); structured C ABI + -oj JSON expose the tags as separate fields. Upstream’s emotion classifier is not exposed — see EU AI Act 50+ langs; native LID + audio-event tags FunASR Model License v1.1

Speech-to-speech audio upscaling and restoration

Backend Model Architecture Input / output License
sidon KevinAHM/Sidon-GGUF (base sarulab-speech/sidon-v0.1) w2v-BERT 2.0 predictor + continuous DAC decoder (more) 16 kHz mono → restored 48 kHz mono MIT
voxcpm2-vae AudioVAE V2 from openbmb/VoxCPM2, converted with --vae-only Isolated causal AudioVAE encoder + decoder (more) 16 kHz mono → upscaled 48 kHz mono Apache-2.0
huggingface-cli download KevinAHM/Sidon-GGUF sidon-v0.1-f16.gguf --local-dir models
crispasr -m models/sidon-v0.1-f16.gguf -f input.wav --s2s --s2s-output restored.wav

python models/convert-voxcpm2-to-gguf.py --input openbmb/VoxCPM2 \
  --output models/voxcpm2-vae-f32.gguf --vae-only
crispasr -m models/voxcpm2-vae-f32.gguf -f input.wav --s2s \
  --s2s-output upscaled.wav

Text-to-Speech models

Synthesis backends, driven by the --tts flag and a --tts-output PATH.wav. See the dedicated Text-to-Speech section below for quick-start commands and engine selection guidance.

Backend Models Architecture Languages License
miotts MioTTS-0.6B Qwen3 LLM + MioCodec-v2 FSQ codec (25 Hz, 44.1 kHz output) ja, en Apache-2.0
vibevoice-tts VibeVoice-Realtime-0.5B, VibeVoice-1.5B DPM-Solver++ + σ-VAE decoder; voice presets or cloning en, zh MIT
kugelaudio kugelaudio-0-open Qwen2.5-7B LM + 4L DiT diffusion + acoustic VAE decoder; voice cloning multilingual Apache-2.0
qwen3-tts Qwen3-TTS-12Hz-0.6B-Base, 1.7B-Base, 1.7B-VoiceDesign Qwen3 talker LM + 12 Hz RVQ (more) multilingual Apache-2.0
qwen3-tts-customvoice 1.7B-CustomVoice Same talker + 9 premium built-in speakers (--voice <name>); optional style via --instruct (e.g. “spoke very slowly”) (more) multilingual Apache-2.0
moss-tts OpenMOSS-Team/MOSS-TTS-v1.5 Qwen3-8B backbone emitting 32 RVQ audio codebooks under a delay pattern, decoded by a 1.6B pure-transformer codec companion; voice cloning via --voice ref.wav; --backend moss-tts -m <backbone> --codec-model <codec> multilingual Apache-2.0
moss-tts-local OpenMOSS-Team/MOSS-TTS-Local-Transformer-v1.5 Qwen3-4B backbone; a 1-layer local/depth transformer autoregressively emits 12 RVQ codebooks per frame (RQ-Transformer, no delay), decoded to 48 kHz by MOSS-Audio-Tokenizer-v2 (downmixed to mono); --backend moss-tts-local -m <backbone> --codec-model <codec> multilingual Apache-2.0
omnivoice k2-fsa/OmniVoice Qwen3-0.6B + masked iterative 8-codebook TTS (SoundStorm-style); voice cloning; 600+ languages (more) 600+ langs Apache-2.0
melotts myshell-ai/MeloTTS EN_V2 VITS2 (6L transformer + SDP/DP + transformer coupling flow + HiFi-GAN); 44.1 kHz, 102 MB + 52 MB BERT Q4_K companion (154 MB total); neural G2P; 4 EN speakers (more) en MIT
piper rhasspy/piper community voices VITS (6L transformer + SDP + 4-block coupling flow + HiFi-GAN); 22 kHz mono, 30 MB F16 per voice; built-in G2P for EN/DE/FR/ES/RU (--g2p-dict) 30+ langs (built-in + espeak dlopen) MIT
kokoro hexgrad/Kokoro-82M + German backbones StyleTTS2 / iSTFTNet (82M); per-voice GGUF (more) en, es, fr, hi, it, ja, pt, zh, de Apache-2.0
orpheus Orpheus-3B-FT + SNAC 24 kHz Llama-3.2-3B + SNAC RVQ codec; 8 speakers (more) en, de Llama 3.2 Community License / MIT
chatterbox cstr/chatterbox-GGUF + Nano/turbo/fine-tune variants T3 AR + S3Gen flow-matching (more) 23 multilingual; separate Arabic, German, and Finnish (chatterbox-finnish-nano) fine-tunes MIT
indextts cstr/indextts-1.5-GGUF GPT-2 AR (24L/1280d) + Conformer conditioning + BigVGAN vocoder; voice cloning via reference audio zh, en Apache-2.0
voxcpm2-tts cstr/voxcpm2-GGUF Tokenizer-free CFM diffusion AR (TSLM + RALM + LocDiT) at 48 kHz native; zero-shot + voice cloning via --voice <wav> 30 languages Apache-2.0
voxtral-tts mistralai/Voxtral-4B-TTS-2603 Ministral-3B AR (26L GQA) + 3L FM acoustic transformer (7-step Euler ODE) + Voxtral codec decoder at 24 kHz; 20 preset voices; SOTA French technical text en, fr, de, es, it, pt, nl, ar, hi CC‑BY‑NC‑4.0
cosyvoice3-tts cstr/cosyvoice3-0.5b-2512-GGUF Qwen2-0.5B AR speech-token LM + DiT-CFM (10-step Euler) + HiFT (NSF + iSTFT) at 24 kHz; baked-voice zero-shot cloning via --voice <name>, or any WAV via --voice ref.wav --ref-text "<exact transcript>". --backend cosyvoice3-tts-rl selects upstream’s RL-tuned talker (same companions) 9 langs + 18 zh dialects Apache-2.0
csm cstr/csm-1b-GGUF Sesame CSM-1B conversational TTS: Llama-3.2 1B backbone + 100M depth decoder (32-codebook RVQ) + Kyutai Mimi codec at 24 kHz (more) en Apache-2.0
lfm2-audio cstr/lfm2-audio-1.5b-GGUF + jp LFM2.5-Audio ASR+TTS+S2S: FastConformer enc + LFM2 hybrid backbone + depthformer (8-codebook Mimi) + ISTFT detokenizer at 24 kHz; interleaved text+audio generation en, ja LFM Open v1.0
dia nari-labs/Dia-1.6B Byte-level text encoder (12L) + AR audio decoder (18L GQA + CFG) → 9 delayed DAC codebooks + 44.1 kHz DAC codec; dialogue style with [S1]/[S2] tags (use >100-char prompts) en Apache-2.0
zonos-tts cstr/zonos-v0.1-transformer-GGUF + cstr/dac-44khz-GGUF Zyphra Zonos-v0.1: 26L GQA AR transformer (2B) + 9-codebook DAC @ 44.1 kHz; CFG-guided; voice cloning via reference WAV (more) en Apache-2.0
bark cstr/bark-small-GGUF Suno Bark 3-stage GPT-2 TTS: text→semantic (12L) → coarse EnCodec (12L, 2 codebooks) → fine (12L, 8 codebooks) → EnCodec 24 kHz decoder; speaker conditioning via .npz prompts (--voice <file.npz>) multilingual MIT
speecht5 cstr/speecht5-tts-GGUF SpeechT5 80M: char-level encoder (12L) + AR mel decoder (6L) + 5-layer conv postnet + HiFi-GAN at 16 kHz; speaker via 512-d x-vector (--voice <xvector.bin>) en MIT
fastpitch cstr/fastpitch-en-GGUF NVIDIA FastPitch 60M: non-autoregressive parallel TTS — 6L encoder + duration/pitch predictors + 6L decoder + HiFi-GAN at 22 kHz; deterministic, single forward pass (more) en CC-BY-4.0
bananamind-tts Banaxi-Tech/BananaMind-TTS-V2.1-Preview BananaMind-TTS 13M: Tacotron-lite char-level encoder (Conv+BN+BiLSTM) + AR GRU decoder with location-sensitive attention + postnet + HiFi-GAN at 22 kHz; fixed voice per locale (more) en, de Apache-2.0
parler-tts cstr/parler-tts-mini-v1.1-GGUF Parler TTS Mini v1.1 (~900M): T5 encoder + MusicGen decoder + DAC 44.1 kHz; prompt-conditioned (describe voice in text via --instruct) en Apache-2.0
outetts cstr/outetts-0.3-1b-GGUF OLMo-1B talker + WavTokenizer single-codebook VQ-GAN at 24 kHz; voice cloning via speaker profile JSON (--voice <speaker.json>) en CC-BY-NC-SA-4.0
pocket-tts cstr/pocket-tts-GGUF Kyutai Pocket TTS 100M: continuous-latent AR at 12.5 Hz + one-step LSD flow + Mimi VAE 24 kHz; voice cloning via ref audio or official prepared .safetensors voices (more) en, de, es, it, pt; fr (24L preview) CC-BY-4.0 + gated-use conditions
tada cstr/tada-tts-1b-GGUF + HumeAI/tada-3b-ml Llama-3.2 1B/3B backbone + per-token FM diffusion head + TADA codec at 24 kHz; 1:1 text-to-acoustic alignment; default prompt via tada-ref.gguf, custom voices via --voice <tada-ref.gguf> built with models/convert-tada-ref-to-gguf.py (more) en Llama 3.2 Community License
TTS feature matrix | Backend | Voice cloning | Sampling | kHz | Auto-download | Flash attn | |---------|:---:|:---:|:---:|:---:|:---:| | vibevoice-tts | yes | temp | 24 | yes | yes | | qwen3-tts | yes* | temp | 24 | yes | yes | | omnivoice | yes | temp | 24 | — | — | | kokoro | — | — | 24 | yes | — | | orpheus | — | temp | 24 | yes | yes | | chatterbox | yes | temp | 24 | yes | yes | | outetts | yes (JSON) | temp | 24 | yes | yes | | indextts | yes | temp | 24 | yes | yes | | voxcpm2-tts | yes | — | 48 | yes | — | | cosyvoice3-tts | yes | temp | 24 | yes | yes | | f5-tts | yes | — | 24 | yes | — | | irodori-tts | yes (WAV) | VoiceDesign: `--instruct` | 48 | yes | — | | supertonic | presets F1-F5/M1-M5 | `--tts-speed`, `--tts-steps` | 44.1 | yes | — | | csm | — | temp | 24 | yes | — | | dia | — | temp | 44 | yes | — | | bark | yes (.npz) | temp | 24 | yes | — | | speecht5 | yes (xvec) | — | 16 | yes | — | | parler-tts | — | temp | 44 | yes | — | | fastpitch | — | — | 22 | — | — | | piper | — | — | 22 | — | — | | pocket-tts | yes | temp | 24 | yes | — | | tada | yes | temp | 24 | yes | — | | dots-tts | yes (`--voice ref.wav`) | 16-step CFG Euler | 48 | yes | — | | fireredtts3 | yes (`--voice ref.wav --ref-text "..."`) | 10-step CFG Euler | 24 | yes | — | | confucius4-tts | yes (`--voice ref.wav`) | 25-step CFG Euler | 22.05 | yes | — | \* CustomVoice variant only; Base uses baked speakers via `--voice `. **Output language.** `-tl ` (or `-l`) selects the language to speak; `cosyvoice3-tts`, `qwen3-tts` and `moss-tts` act on it natively. For cross-lingual **cloning** — an English reference clip speaking German, the subtitle-dubbing case — also pass `-sl ` for the language the reference is spoken in, so cosyvoice3 drops the reference transcript instead of carrying its accent. Over HTTP: `"language"` + `"source_lang"` on `POST /v1/audio/speech`. See [`docs/tts.md`](/CrispASR/docs/tts.html#output-language-and-cross-lingual-cloning--tl---sl). </details> ### Translation Text-to-text translation, distinct from the audio-side `--translate` flag (which routes audio → English text on whisper / canary / etc.). Driven by `--text "..." -sl -tl `. | Backend | Models | Architecture | Languages | License | |---|---|---|---|---| | **m2m100** | [`facebook/m2m100_418M`](https://huggingface.co/cstr/m2m100-418m-GGUF) | 12L enc + 12L dec transformer, SentencePiece 128K ([more](/CrispASR/docs/architecture.html#m2m100--wmt21)) | 100 langs, any-to-any | MIT | | **m2m100-wmt21** | [`facebook/wmt21-dense-24-wide-en-x`](https://huggingface.co/cstr/wmt21-dense-24-wide-en-x-GGUF) + [`facebook/wmt21-dense-24-wide-x-en`](https://huggingface.co/cstr/wmt21-dense-24-wide-x-en-GGUF) | Same as m2m100, scaled to 4.7B (24L enc) ([more](/CrispASR/docs/architecture.html#m2m100--wmt21)) | English ↔ 7 langs (separate `en-x` / `x-en` checkpoints) | MIT | | **madlad** | [`google/madlad400-3b-mt`](https://huggingface.co/cstr/madlad400-3b-mt-GGUF) | T5 enc-dec (12L+12L, d=2048, gated-GELU, RMSNorm) ([more](/CrispASR/docs/architecture.html#madlad)) | 419 languages | Apache-2.0 | ```bash # m2m100 base (production-ready) ./build/bin/crispasr --backend m2m100 -m auto \ --text "Hello world, how are you today?" \ -sl en -tl de # → Hallo Welt, wie bist du heute? # WMT21 dense (English ↔ X, 4.7B — auto-downloads ~2.5 GB). # Two separate checkpoints: en-x for English-source, x-en for # English-target. Pick the one matching your `-sl`/`-tl` direction # (or pass an explicit `-m ` to load the other manually). ./build/bin/crispasr --backend m2m100-wmt21 -m auto \ --text "The president said he would not attend." \ -sl en -tl de # uses wmt21-dense-24-wide-en-x ./build/bin/crispasr --backend m2m100-wmt21 \ -m models/wmt21-dense-24-wide-x-en-q4_k.gguf \ --text "Le président a dit qu'il ne serait pas présent." \ -sl fr -tl en # uses wmt21-dense-24-wide-x-en # MADLAD-400 3B (419 languages, bit-token-identical to Python SP) ./build/bin/crispasr --backend madlad -m auto \ --text "Hello world." \ -sl en -tl ta ``` For 2-stage pipelines (e.g., ASR → m2m100), use the dedicated `--tr-sl` / `--tr-tl` flags; they fall back to `-sl` / `-tl` when unset, so single-stage standalone usage is just `-sl/-tl`. ### Post-processing models Work with all backends. | Model | Task | Architecture | Languages | License | HuggingFace | |---|---|---|---|---|---| | **FireRedPunc** | Punctuation restoration | BERT-base (12L, d=768), 5 classes | Chinese + English | Apache-2.0 | [`cstr/fireredpunc-GGUF`](https://huggingface.co/cstr/fireredpunc-GGUF) | | **fullstop-punc** | Punctuation restoration | XLM-RoBERTa-large (24L, d=1024), 6 classes | EN, DE, FR, IT | MIT | [`cstr/fullstop-punc-multilang-GGUF`](https://huggingface.co/cstr/fullstop-punc-multilang-GGUF) | | **punctuate-all** | Punctuation restoration | XLM-RoBERTa-base (12L, d=768), 6 classes | 12 languages | MIT | [`cstr/punctuate-all-GGUF`](https://huggingface.co/cstr/punctuate-all-GGUF) | | **PCS** | Punc + truecase + SBD | XLM-RoBERTa-base (12L), 4 heads | 47 languages | Apache-2.0 | `--punc-model pcs` | | **truecaser‑lstm** | German truecasing (best) | BiLSTM char-level (2×150, 3.2 MB, 97.9% F1) | German | Apache-2.0 | `--truecase-model lstm` | | **truecaser‑crf** | German truecasing | CRF + context features (8.5 MB) | German | MIT | `--truecase-model crf` | | **truecaser‑de** | German truecasing (simple) | Statistical word-frequency (71K entries, 1.7 MB) | German | MIT | `--truecase-model auto` | | **CLD3** | Text language ID | Embedding-bag → FC + ReLU → softmax (~1.5 MB F32) | 109 ISO 639-1 | Apache-2.0 | [`cstr/cld3-GGUF`](https://huggingface.co/cstr/cld3-GGUF) | | **GlotLID-V3** | Text language ID | fastText supervised, flat softmax | 2102 ISO 639-3 + script | Apache-2.0 | [`cstr/glotlid-GGUF`](https://huggingface.co/cstr/glotlid-GGUF) | | **LID-176** | Text language ID | fastText supervised, hierarchical softmax | 176 ISO 639-1 | CC-BY-NC-4.0 | [`cstr/fasttext-lid176-GGUF`](https://huggingface.co/cstr/fasttext-lid176-GGUF) | ### Audio codecs Shared codec modules used by TTS backends. Also available standalone for encode/decode. | Model | Architecture | Sample Rate | Token Rate | License | HuggingFace | |---|---|---|---|---|---| | **MioCodec v2** | WavLM encoder → FSQ(12800) → Transformer decoder + AdaLN-Zero + SnakeBeta upsampler + iSTFT | 44.1 kHz | 25 Hz (341 bps) | MIT | [`cstr/miocodec-v2-44k-GGUF`](https://huggingface.co/cstr/miocodec-v2-44k-GGUF) | | **SNAC 24 kHz** | 3-codebook RVQ + decoder blocks (stride 8/8/4/2) | 24 kHz | 3×12.5 Hz | MIT | [`cstr/snac-24khz-GGUF`](https://huggingface.co/cstr/snac-24khz-GGUF) | All runtimes share ggml-based inference. The speech-LLM backends (**qwen3**, **voxtral**, **voxtral4b**, **granite**, **glm-asr**, **kyutai-stt**) inject audio encoder frames directly into an autoregressive language model's input embeddings, instead of using a dedicated CTC/transducer/seq2seq decoder. The **fastconformer-ctc** backend hosts the NeMo FastConformer-CTC standalone ASR family — `stt_en_fastconformer_ctc_{large,xlarge,xxlarge}` and the architecturally-identical `parakeet-ctc-{0.6b,1.1b}` (different training data + tokenizer, same encoder + head shape) — with greedy CTC decoding. Same C++ runtime as the canary-ctc aligner. ### Music & audio analysis Beyond speech, CrispASR runs several music/audio analysis tasks — each a small GGUF with the architecture auto-detected, no Python. See [`docs/cli.md`](/CrispASR/docs/cli.html) for the per-task flags and output formats. - **Source separation** (`--separate`) — split a mix into stems (`_.wav`) via **mel-band-roformer** (vocal/instrumental, MIT) or **htdemucs** (4-stem). `--stems vocals,drums` selects a subset; `--sep-output-dir` sets the output location. - **Piano transcription** (`--backend piano-transcription`) — piano audio → note events (88 keys @ 100 fps, ByteDance/Kong CRNN; F16 GGUF ≈ 77 MB). - **Polyphonic note events** (`--backend basic-pitch`) — Spotify Basic Pitch, any instrument → note events (~110 KB model). - **Multi-instrument transcription** (`--backend mt3`, alias `music-transcription`) — MT3's T5 encoder/decoder emits note events with per-instrument programs (F16 GGUF ≈ 96 MB). - All three take `--piano-format text|json|midi`; `midi` writes a Standard MIDI File. - **Guitar tablature** (`--tab`) — per-frame fret-per-string grid via **TabCNN** (Wiggins & Kim, ISMIR 2019; CC BY 4.0 weights). The backend emits per-string emission scores, not a decided tablature — run your own constrained Viterbi via `crispasr_session_tab_emissions()` for playable output. - **Beat / downbeat tracking** (`--beats`) — beat grid via **Beat This!** (CPJKU, ISMIR 2024; MIT for code *and* weights, no patent-encumbered DBN). - **Chord recognition** (`--chords`) — chord timeline (`.lab`) via **BTC** (ISMIR 2019). Weights are CC-BY-NC-SA, gated behind `--accept-license cc-by-nc-sa-4.0`. - **Pitch / F0 estimation** (`--pitch`) — monophonic pitch track via **CREPE** (MIT). ## Feature matrix Run `crispasr --list-backends` to see it live. Each backend declares capabilities at runtime; if you ask for a feature the selected backend does not support, CrispASR prints a warning and silently ignores the flag. **Sortable / filterable view:** [`docs/feature-matrix.html`](/CrispASR/docs/feature-matrix.html) — click any column header to sort, type to filter rows, click cap pills to require a capability. Generated from `crispasr --list-backends-json` (single source of truth — drift impossible). Regenerate via `python tools/gen-feature-matrix.py`. A Markdown twin lives at [`docs/feature-matrix.md`](/CrispASR/docs/feature-matrix.html). The static table below is a curated subset focusing on the ASR backends and the cross-cutting features that matter for ASR pipelines. The full 119-backend × 27-cap surface is in the generated views. | Feature | whisper | parakeet | canary | cohere | granite | granite‑4.1 | voxtral | voxtral4b | qwen3 | fc‑ctc | wav2vec2 | glm‑asr | kyutai‑stt | firered | moonshine | moon‑stream | omniasr | omniasr‑llm | vibevoice | gemma4‑e2b | mimo‑asr | funasr | paraformer | sensevoice | |---|:-:|:-:|:-:|:-:|:-:|:-:|:-:|:-:|:-:|:-:|:-:|:-:|:-:|:-:|:-:|:-:|:-:|:-:|:-:|:-:|:-:|:-:|:-:|:-:| | Native timestamps | ✔ | ✔ | ✔ | ✔ | | | | | | | | | ✔ | | | | | | | | | | | | | CTC timestamps | | | ✔ | | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | | Word-level timing | ✔ | ✔ | ✔ | ✔ | `-am` | ✔† | `-am` | `-am` | `-am` | `-am` | `-am` | `-am` | ✔ | `-am` | `-am` | `-am` | `-am` | `-am` | | `-am` | `-am` | `-am` | | `-am` | | Per-token confidence | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | | | ✔ | ✔ | | | | Language auto-detect | ✔ | ✔ | LID | LID | LID | LID | LID | LID | ✔ | LID | LID | ✔ | LID | LID | LID | LID | LID | LID | LID | ✔ | LID | LID | LID | ✔ | | Speech translation | ✔ | | ✔ | | ✔ | ✔ | ✔ | | ✔ | | | | | | | | | | | | | | | | | Speaker diarization | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | | Grammar (GBNF) | ✔ | | | | | | | | | | | | | | | | | | | | | | | | | Temperature sampling | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | | | ✔ | ✔ | | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | | | | Beam search | ✔ | | | | ✔ | ✔ | ✔ | | ✔ | | | ✔ | ✔ | ✔ | ✔ | | ✔ | ✔ | | | | | | | | Flash attention | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | | | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | | Punctuation toggle | | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | | | ✔ | ✔ | | ✔ | | ✔ | ✔ | | | | ✔ | | ✔ | | Punc restoration | pp | pp | pp | pp | pp | pp | pp | pp | pp | pp | pp | pp | pp | pp | pp | pp | pp | pp | pp | pp | pp | pp | pp | pp | | Source / target language | | | ✔ | | ✔ | ✔ | ✔ | | ✔ | | | | | | | | | | | | | | | | | Audio Q&A (`--ask`) | | | | | * | * | ✔ | | * | | | * | | | | | | | | * | * | | | | | Streaming | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | | Auto-download (`-m auto`) | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | | KV quant (`CRISPASR_KV_QUANT`, plus per-half `_K` / `_V`) | | | | | ✔ | ✔ | ✔ | ✔ | ✔ | | | ✔ | | | | | | ✔ | | ✔ | ✔ | ✔ | | | | mmap weights (`CRISPASR_GGUF_MMAP`) | | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | | TTS | | | | | | | | | | | | | | | | | | | ✔ | | | | | | The matrix above covers 24 ASR backends. **Additional ASR backends** not shown: `nemotron` (39-lang streaming ASR with cache-aware FastConformer + RNN-T), `lfm2-audio` (ASR + TTS + S2S in one model), `moss-audio` (audio understanding + ASR), `moss-transcribe` (Qwen3-Omni encoder + Qwen3-1.7B ASR), `mini-omni2` (ASR + TTS + S2S), `kugelaudio` (7B audio understanding). See [`docs/feature-matrix.md`](/CrispASR/docs/feature-matrix.html) for the full 109-backend matrix. **TTS-only backends** (`kokoro`, `qwen3-tts` + variants, `vibevoice-tts`, `orpheus` + DE variants, `chatterbox` / `chatterbox-turbo` / `chatterbox-nano` / `kartoffelbox-turbo` / `lahgtna-chatterbox`, `dia`, `bark`, `outetts`, `zonos`, `csm`, `f5-tts`, `irodori-tts`, `supertonic`, `parler-tts`, `speecht5`, `piper`, `fastpitch`, `pocket-tts`, `melotts`, `cosyvoice3`, `voxcpm2`, `tada-tts`) all carry the TTS, AUTO_DOWNLOAD, TEMPERATURE, and FLASH_ATTN caps; per-backend cloning + voice-pack support is documented in the [Text-to-Speech models](#text-to-speech-models) table above and [`docs/tts.md`](/CrispASR/docs/tts.html). The vibevoice and lfm2-audio columns mark dual-mode (ASR + TTS) backends. **Key:** ✔ = native/built-in, `-am` = via CTC forced aligner (`-am canary-ctc-aligner.gguf` or `-am qwen3-forced-aligner.gguf`), **LID** = via external language identification pre-step (`-l auto`), **pp** = via `--punc-model` post-processor (FireRedPunc or fullstop-punc), * = experimental or partial support, † = PLUS variant only (native `[T:N]` word timestamps with `-owts`; base uses `-am`). granite-4.1 covers both the regular and `-plus` variants; granite-4.1-nar is a non-autoregressive variant with encoder+projector only (no LLM decode features). The **KV quant** row marks backends that honor `CRISPASR_KV_QUANT={f16,q8_0,q4_0}` — CTC-style backends without a KV cache (parakeet, fc-ctc, wav2vec2, kyutai-stt, firered, moonshine variants, omniasr-CTC) don't apply. The same backends also honor the per-half `CRISPASR_KV_QUANT_K` / `CRISPASR_KV_QUANT_V` overrides (llama.cpp `--cache-type-k` / `--cache-type-v` parity) for asymmetric K-vs-V precision; common recipe `K=q8_0 V=q4_0` saves ~40 % more KV memory than symmetric Q8_0. The **mmap weights** row marks backends consuming `core_gguf::load_weights()` and therefore honoring `CRISPASR_GGUF_MMAP=1`; whisper itself uses upstream's loader and is unaffected. See [`docs/cli.md`](/CrispASR/docs/cli.html) Memory footprint for usage + recommended combos. **Speaker diarization** as a post-processing step via `--diarize`: - `energy` / `xcorr` — stereo-only, no extra deps - `foxnose` — **best accuracy, no external deps**: WeSpeaker ResNet34-LM embeddings + GMM/BIC speaker counting + spectral clustering + Viterbi smoothing. Estimates the speaker count rather than needing it up front; `--diarize-embedder auto` fetches the GGUF (24 MB, CC-BY-4.0). 7.3 % DER on VoxConverse dev where `pyannote` + TitaNet scores 7.8 %, and 3.18 % vs the upstream reference's 3.07 % when scored on turns ([more](/CrispASR/docs/architecture.html#foxnose-diarize)) - `pyannote` — native GGUF (no Python, no sherpa-onnx); add `--diarize-embedder auto` (TitaNet) or `--diarize-embedder indextts` (ECAPA-TDNN) for globally stable speaker IDs across long files - `sherpa` / `ecapa` — external [sherpa-onnx](https://github.com/k2-fsa/sherpa-onnx) subprocess; runs once globally on full audio for consistent speaker IDs (#110) - `vad-turns` — mono-friendly gap-based proxy The server endpoint supports `response_format=diarized_json` for structured speaker-labelled output with normalised speaker letters (A, B, C …) — see [`docs/server.md`](/CrispASR/docs/server.html#diarized-json-format-206). Full reference + tuning knobs (cluster threshold, max speakers, pluggable embedder adapters): see [`docs/cli.md#diarization`](/CrispASR/docs/cli.html#diarization). **Language identification** for backends without native LID: `--lid-backend whisper` (default, 75 MB ggml-tiny.bin), `--lid-backend silero` (native GGUF, 16 MB, 95 languages), or `--lid-backend firered` (FireRedLID, 1.7 GB, 120 languages — Conformer encoder + Transformer decoder). **Voice activity detection**: `--vad` uses the default Silero VAD (~885 KB, auto-downloaded). Each VAD segment is transcribed independently, producing separate SRT/VTT entries with correct timestamps. Use `--vad --split-on-punct` for best subtitle output. Four VAD backends: Silero (default), FireRedVAD (`-vm firered`, recommended), MarbleNet (`-vm marblenet`, 439 KB, 6 languages), Whisper-VAD-EncDec (`-vm whisper-vad`, experimental). **Punctuation restoration** (`--punc-model`): CTC-based backends output lowercase without punctuation. Named shortcuts: `auto`/`firered` (Chinese+English), `fullstop` (EN/DE/FR/IT, XLM-R-large), `punctuate-all` (12 languages, XLM-R-base), `pcs` (47 languages, punc + truecasing + sentence boundary detection in one model). Or pass a GGUF path directly. Also available via Python/Rust/Dart wrappers (`crispasr.PuncModel`). **Truecasing** (`--truecase-model`): Restore German noun/name capitalization in lowercase ASR output. Three options in ascending quality: `auto` (statistical, 1.7 MB), `crf` (CRF with context, 8.5 MB), `lstm` (BiLSTM char-level, 3.2 MB, **recommended** — 97.9% F1, handles adjective/noun distinction and formal "Ihnen"). All auto-download from [`cstr/truecaser-de`](https://huggingface.co/cstr/truecaser-de). Or use `--punc-model pcs` for neural punc + truecasing in one pass (47 languages).
Which backends produce punctuation natively? | Backend | Punctuation | Capitalization | Notes | |---|:-:|:-:|---| | whisper | ✔ | ✔ | Full punctuation and casing | | parakeet | ✔ | ✔ | | | canary | ✔ | ✔ | | | cohere | ✔ | ✔ | Toggleable via `--no-punctuation` | | granite | ✔ | ✔ | LLM output | | voxtral | ✔ | ✔ | LLM output | | voxtral4b | ✔ | ✔ | LLM output | | qwen3 | ✔ | ✔ | LLM output | | funasr | ✔ | ✔ | LLM output (Qwen3-0.6B decoder). Chinese chars carry full-width period; mlt-nano variant adds Latin-script casing + punctuation. | | sensevoice | ✔ | ✔ | CTC output with native ITN — toggle off via `--no-punctuation`, which controls Arabic-digit vs spelled-out numerals + comma/period emission. | | paraformer | **no** | **no** | NAR character-level output — add `--punc-model` | | gigaam | ✔ (`e2e_*`) | ✔ (`e2e_*`) | The `e2e_*` revisions carry punctuation + casing + inverse text normalization in the SentencePiece vocabulary. The charwise `ctc` / `rnnt` revisions emit lowercase Cyrillic with no punctuation — but auto-restoration is still suppressed for them, because the auto-enabled FireRedPunc is a Chinese/English model and injects full-width CJK punctuation into Russian. Use an `e2e_*` revision for punctuated output, or pass an explicit `--punc-model`. | | glm-asr | ✔ | ✔ | LLM output | | kyutai-stt | ✔ | ✔ | LLM output | | moonshine | ✔ | ✔ | Encoder-decoder output | | **fastconformer-ctc** | **no** | **no** | CTC — add `--punc-model` | | **wav2vec2** | **no** | **no** | CTC — add `--punc-model` | | **firered-asr** | **no** | **no** | CTC — add `--punc-model` | | **omniasr** (CTC) | **no** | **no** | CTC — add `--punc-model` | | **omniasr** (LLM) | ✔ | ✔ | Autoregressive decoder | Other freely-licensed alternatives that could be added: [felflare/bert-restore-punctuation](https://huggingface.co/felflare/bert-restore-punctuation) (MIT, English, includes truecasing), [xashru/punctuation-restoration](https://github.com/xashru/punctuation-restoration) (Apache-2.0, 40+ languages, BiLSTM-CRF).
**Progressive subtitle output** (`--flush-after`): By default, non-whisper backends buffer all segments and print output at the end. For real-time subtitle consumption (PotPlayer, custom media players), use `--flush-after 1` to print each SRT entry to stdout immediately after its VAD segment is transcribed: ```bash crispasr --backend parakeet -m parakeet.gguf --vad --flush-after 1 -osrt -f long_audio.wav # SRT entries appear progressively as each segment finishes ``` **JSON output with language detection**: When using `-l auto -oj`, the JSON output includes detected language info: ```json { "crispasr": { "backend": "cohere", "language": "en", "language_detected": "en", "language_confidence": 0.977, "language_source": "ecapa" }, "transcription": [...] } ``` ### Which backend should I pick? | Need | Pick | |---|---| | Battle-tested, all features exposed | **whisper** | | Lowest English WER | **cohere** | | **Fastest** (16x realtime on CPU) | **moonshine** (tiny), **fc-ctc** (10x) | | Multilingual + word timestamps + fast | **parakeet** (2.9x RT) | | Multilingual with **explicit language control** | **canary** | | **Speech translation** (X→en or en→X) | **canary**, **voxtral**, **qwen3** | | **30 languages + Chinese dialects** | **qwen3** | | **1600+ languages** | **omniasr** (CTC or LLM) | | **Realtime streaming ASR** (native incremental encoder, ~2× RT feed; sub-second-token target deferred to phase 2) | **voxtral4b** | | Highest-quality offline speech-LLM | **voxtral** | | Apache-licensed speech-LLM | **granite**, **voxtral**, **qwen3**, **omniasr-llm** | | **Lightweight CTC-only** (fast, no decoder) | **wav2vec2**, **fc-ctc**, **data2vec**, **omniasr** | | **Russian** | **gigaam** (`e2e_rnnt` — 8.4 % avg WER, punctuation + ITN), **whisper**, **qwen3** | | **Mandarin + Chinese dialects** | **firered-asr**, **qwen3**, **glm-asr**, **funasr**, **paraformer**, **sensevoice** | | **Multilingual (31 langs) speech-LLM** | **fun-asr-mlt-nano**, **qwen3**, **omniasr-llm**, **gemma4-e2b** | | **Multilingual (50+ langs) + LID + audio-event in one pass** | **sensevoice** (encoder-only CTC, non-AR, 15× faster than Whisper-Large) | ### CPU performance tips Audio-LLM backends (`qwen3`, `voxtral`, `granite`, `glm-asr`, etc.) run full transformer decoder stacks (28+ layers, 2048-dim) and are **dramatically slower on CPU** than encoder-only backends. On older dual-core hardware they can drop below 0.01× realtime. If you're on CPU-only hardware: - Prefer **moonshine** (16× RT), **fc-ctc** (10× RT), **parakeet** (2.9× RT), or **whisper** for usable speeds. - Use `--flush-after 1` to see results as each VAD slice completes instead of waiting for the entire file. - Use `-pp` / `--print-progress` for per-slice progress indicators on all backends (unified backends show slice-level progress; whisper shows encoder-level progress). - Quantize models to Q4_K or Q5_K to reduce memory and compute. ### Language detection for backends that don't do it natively Cohere, canary, granite, voxtral and voxtral4b need an explicit language code up front. If you don't know the language, pass `-l auto` and crispasr runs an optional LID pre-step before the main transcribe() call: ```bash # Downloads ggml-tiny.bin (75 MB, 99 languages) on first use crispasr --backend cohere -m $TC/cohere-transcribe-q5_0.gguf \ -f unknown.wav -l auto # crispasr[lid]: detected 'en' (p=0.977) via whisper-tiny # crispasr: LID -> language = 'en' (whisper, p=0.977) ``` These LID providers are available: - `--lid-backend whisper` (default) — uses a small multilingual ggml-*.bin model via the crispasr C API. Auto-downloads ~75 MB on first use. 99 languages. - `--lid-backend silero` — native GGUF port of Silero's 95-language classifier. 16 MB F32. Runs as a ggml graph (multi-threaded SIMD on CPU, GPU offload on Metal/CUDA; on Vulkan the graph is routed to CPU pending an upstream kernel fix). Analyzes the first 30 s of audio (`CRISPASR_SILERO_LID_MAX_S` overrides); `CRISPASR_SILERO_LID_LEGACY=1` restores the old scalar path. - `--lid-backend ecapa` — **recommended**: ECAPA-TDNN (Apache-2.0). Purpose-built for language ID. Very high accuracy on TTS benchmark. Two variants via `--lid-model`: - [`cstr/ecapa-lid-107-GGUF`](https://huggingface.co/cstr/ecapa-lid-107-GGUF) — VoxLingua107, 43 MB F16, 107 languages, ISO codes (en, de, ...). **Default.** - [`cstr/ecapa-lid-commonlanguage-GGUF`](https://huggingface.co/cstr/ecapa-lid-commonlanguage-GGUF) — CommonLanguage, 40 MB F16, 45 languages, full names (English, German, ...). - `--lid-backend firered` — FireRedLID (Conformer encoder + Transformer decoder). Q4_K (544 MB), 120 languages including Chinese dialects. Slower but covers more languages. - `--lid-backend probe` — no second model at all: ask the **ASR model itself**. Transcribes a 20 s clip once per language the model declares and keeps the best-scoring candidate (length × text-LID agreement × distinct-token ratio², the last term catching the repetitive output a wrong-language prompt produces). Currently implemented by **cohere**. The reason to prefer it is correctness, not just the saved download: an external detector knows 99 languages while Cohere Transcribe accepts 14 — and its Arabic finetune only `en`/`ar` — so external LID regularly returns a language the model was never trained on, and Cohere answers a wrong language *fluently* rather than failing. The probe cannot. Cost is one encode + one short decode per candidate, so it runs automatically only for a model with ≤ 4 languages (`CRISPASR_COHERE_PROBE_MAX_LANGS`); `CRISPASR_COHERE_PROBE_TEXTLID=0` drops the text-LID agreement term. **The ceiling is about cost, not accuracy.** Measured on the real models: the two-language Arabic finetune picks `ar` for an Arabic clip (p=0.675) and `en` for `samples/jfk.wav` (p=0.647); the 14-language base model, probed across all 14, also gets both right (`en` p=0.169, `ar` p=0.254) — it is simply slower than an external detector. The encoder output is language-independent, so the probe encodes **once** and decodes per candidate (encode is ~87 % of a pass); a 14-candidate probe measured **12 s → 4-5 s** against one-encode-per-candidate, byte-identical output. `CRISPASR_COHERE_PROBE_REUSE_ENC=0` restores the naive path. The one soft spot worth knowing: asking the model for a language it was *not* trained on can yield a clean translation rather than garbage, which a text LID then confirms — "fluent French out" is not evidence of French in. That is what makes a mismatched whitelist dangerous, and it is measurable: force a 14-language list onto the two-language Arabic finetune and its `fr` probe returns real French and wins. The real base model's `fr` probe instead code-switches ("Et so, my fellow Americans…", agreement 0.00) and loses, as it should. These VAD providers are available: - **Silero VAD** (default) — ~885 KB, auto-downloaded via `--vad`. Industry-standard, well-tested. - **FireRedVAD** — DFSMN-based, 2.4 MB, F1=97.57%. Pass `--vad -vm firered` to auto-download. Recommended. - **MarbleNet** — NVIDIA 1D separable CNN, 439 KB, 6 languages (EN/DE/FR/ES/RU/ZH). Pass `--vad -vm marblenet` to auto-download. Smallest model. ([`cstr/marblenet-vad-GGUF`](https://huggingface.co/cstr/marblenet-vad-GGUF)) - **Whisper-VAD-EncDec** *(experimental)* — Whisper-base encoder + TransformerDecoder head, 22 MB Q4_K. Trained on Japanese ASMR; may not generalise well to all domains. Pass `--vad -vm whisper-vad`. Slower than others (~1s vs ~50ms). ([`cstr/whisper-vad-encdec-asmr-GGUF`](https://huggingface.co/cstr/whisper-vad-encdec-asmr-GGUF)) Pass `--lid-backend off` to skip LID entirely. ### Text language identification (post-ASR / standalone) Audio LID (above) tags **what was spoken**; text LID tags **what was written**. Text LID runs on a transcript or any UTF-8 string and is useful for routing post-ASR pipelines (translation, punctuation, sub selection) without re-running an audio model. Three GGUF families, one binary — the dispatcher picks by `general.architecture`: | Backend | Labels | Size (F16) | License | HF repo | |---|---:|---:|---|---| | **CLD3** (Google compact language detector v3) | 109 ISO 639-1 | **440 KB** | Apache-2.0 | [`cstr/cld3-GGUF`](https://huggingface.co/cstr/cld3-GGUF) | | **GlotLID-V3** (cis-lmu fastText) | 2102 ISO 639-3 + script | 250 MB | Apache-2.0 | [`cstr/glotlid-GGUF`](https://huggingface.co/cstr/glotlid-GGUF) | | **LID-176** (Facebook fastText) | 176 ISO 639-1 | 63 MB | CC-BY-NC-4.0¹ | [`cstr/fasttext-lid176-GGUF`](https://huggingface.co/cstr/fasttext-lid176-GGUF) | ¹ LID-176 is **CC-BY-NC-4.0** — non-commercial use only. CLD3 + GlotLID-V3 are Apache-2.0 with no such constraint. Pick CLD3 for the smallest, fastest path; GlotLID for maximum coverage (low-resource languages); LID-176 only if you need its specific 176-label space and accept its non-commercial terms. **Standalone CLI** — auto-routes by GGUF arch, with auto-download: ```bash crispasr-lid -m auto --text "Bonjour le monde" # → cstr/cld3-GGUF (default, ~440 KB) crispasr-lid -m auto:glotlid --text "Bonjour le monde" -k 5 crispasr-lid -m auto:lid-fasttext176 --text "Hallo Welt" # Or pass an explicit path / canonical filename (looked up in the registry): crispasr-lid -m cld3-f16.gguf --text "你好世界" # zh 0.997816 echo "Привет мир" | crispasr-lid -m auto --quiet # ru 0.907322 ``` **Post-ASR pipeline** — `--lid-on-transcript` runs the same dispatcher on the assembled transcript (also accepts `auto[:variant]`): ```bash crispasr -m ggml-tiny.bin -f speech.wav --lid-on-transcript auto # (transcript on stdout) # lang=de conf=0.997123 backend=lid-cld3 ``` The dispatcher (`src/text_lid_dispatch.{h,cpp}`) is a thin C ABI façade — one integer compare per call; per-stage diff harness is green at cos≥0.999 across 8 multilingual smoke samples. --- ## Install & build **Don't want to build?** Prebuilt binaries for Windows, macOS and Linux are on the [releases page](https://github.com/CrispStrobe/CrispASR/releases/latest) — see [Start here](#start-here) for which file to take. The rest of this section is for building from source. ```bash git clone --recursive https://github.com/CrispStrobe/CrispASR cd CrispASR # already cloned without --recursive? initialize the bundled ggml submodule: # git submodule update --init --recursive cmake -B build -DCMAKE_BUILD_TYPE=Release cmake --build build -j$(nproc) ``` The `ggml/` submodule is required. If you cloned without `--recursive`, run `git submodule update --init --recursive` first — otherwise CMake stops with a message telling you to do exactly that. Produces `build/bin/crispasr` (main CLI), `build/bin/crispasr-quantize`, and `build/bin/crispasr-diff`. No Python, PyTorch, or pip required at runtime — just a C++17 compiler and CMake 3.14+. For GPU acceleration, add the matching ggml flag at configure time: ```bash cmake -B build -DCMAKE_BUILD_TYPE=Release -DGGML_CUDA=ON # NVIDIA cmake -B build -DCMAKE_BUILD_TYPE=Release -DGGML_METAL=ON # Apple Silicon cmake -B build -DCMAKE_BUILD_TYPE=Release -DGGML_VULKAN=ON # cross-vendor ``` **See [`docs/install.md`](/CrispASR/docs/install.html)** for the full guide: all GPU backends (CUDA / Metal / Vulkan / MUSA / SYCL), Windows convenience scripts, ffmpeg ingestion, optional BLAS, glibc notes, and the `scripts/dev-build.sh` wrapper. If a build runs but the binary exits with no output, see [`docs/troubleshooting.md`](/CrispASR/docs/troubleshooting.html). --- ## Quick start Deeper ASR examples below. If this is your first run, use [Start here](#start-here) instead. For TTS, the runnable guide is [docs/tts.md](/CrispASR/docs/tts.html) ([Text-to-Speech](#text-to-speech-models) below is the model catalogue). ### Whisper (historical path, byte-identical to upstream whisper.cpp) ```bash # Download a whisper model (same as upstream whisper.cpp) ./models/download-ggml-model.sh base.en ./build/bin/crispasr -m models/ggml-base.en.bin -f samples/jfk.wav # [00:00:00.000 --> 00:00:07.940] And so my fellow Americans ask not what your country can do for you # [00:00:07.940 --> 00:00:10.760] ask what you can do for your country. ``` ### Parakeet (multilingual, free word timestamps, fastest) ```bash # Grab the quantized model (~467 MB) curl -L -o parakeet.gguf \ https://huggingface.co/cstr/parakeet-tdt-0.6b-v3-GGUF/resolve/main/parakeet-tdt-0.6b-v3-q4_k.gguf ./build/bin/crispasr -m parakeet.gguf -f samples/jfk.wav # Auto-detected backend 'parakeet' from GGUF metadata. # And so, my fellow Americans, ask not what your country can do for you, ask what you can do for your country. # Word-level timestamps (one line per word) ./build/bin/crispasr -m parakeet.gguf -f samples/jfk.wav -ml 1 ``` ### Canary (explicit language, speech translation) ```bash # Transcription (source == target) ./build/bin/crispasr --backend canary -m canary-1b-v2-q5_0.gguf -f audio.de.wav -sl de -tl de # Translation (German speech → English text) ./build/bin/crispasr --backend canary -m canary-1b-v2-q5_0.gguf -f audio.de.wav -sl de -tl en # ...or use the familiar crispasr flag: ./build/bin/crispasr --backend canary -m canary-1b-v2-q5_0.gguf -f audio.de.wav -l de --translate ``` ### Voxtral (speech-LLM with auto-download) ```bash # First run downloads ~2.5 GB to ~/.cache/crispasr/ via curl, then runs ./build/bin/crispasr --backend voxtral -m auto -f samples/jfk.wav # Subsequent runs use the cached file ./build/bin/crispasr --backend voxtral -m auto -f samples/jfk.wav -l en ``` ### Qwen3-ASR (30 languages + Chinese dialects) ```bash # 0.6B (default, ~500 MB) ./build/bin/crispasr --backend qwen3 -m auto -f audio.zh.wav # 1.7B (higher quality, ~1.3 GB) — supports both -hf and non-hf source models ./build/bin/crispasr --backend qwen3 -m qwen3-1.7b --auto-download -f audio.wav # Japanese anime/galgame fine-tune (~1.3 GB) ./build/bin/crispasr --backend qwen3 -m qwen3-ja-anime --auto-download -f anime.wav ``` **Long audio:** the default is safe 30 s chunking. `--chunk-seconds 0` decodes the whole file in ONE pass (matches the reference model verbatim on multi-minute clips, #218) — but the encoder's full attention is O(N²) in audio length, so keep single-pass clips under ~10 minutes on 16 GB machines. For long-form use prefer the plain `-q4_k`/`-q8_0` GGUFs over the `-imatrix` variants (see the model card). ### GLM-ASR-Nano (Mandarin + dialects + Cantonese + English, 1.5B) ```bash ./build/bin/crispasr --backend glm-asr -m auto -f audio.wav # Long audio in one pass (up to 655 s — 30 s encoder windows, one LLM prompt, # same layout as the HF/zai reference; matches it verbatim on the #218 clip): ./build/bin/crispasr --backend glm-asr -m auto --chunk-seconds 0 -f long.wav ``` Note: in single-pass mode the model (like the reference) skips leading non-speech audio; the default 30 s-chunked mode transcribes more of such clips. Custom `--ask` / non-English `--language` instructions need a GGUF with baked BPE merges (re-published 2026-07; older GGUFs fall back to the default transcription prompt with a warning). ### MiMo-V2.5-ASR (Mandarin + dialects + English, 7.5B Qwen2 LM) ```bash # Download the LM + audio tokenizer (the tokenizer is a separate model) huggingface-cli download cstr/mimo-asr-GGUF mimo-asr-q4_k.gguf \ --local-dir ~/.cache/crispasr huggingface-cli download cstr/mimo-tokenizer-GGUF mimo-tokenizer-q4_k.gguf \ --local-dir ~/.cache/crispasr # Transcribe (auto-discovers tokenizer if it sits next to the LM) ./build/bin/crispasr \ --backend mimo-asr \ -m ~/.cache/crispasr/mimo-asr-q4_k.gguf \ --codec-model ~/.cache/crispasr/mimo-tokenizer-q4_k.gguf \ -f samples/jfk.wav # Output: And so, my fellow Americans, ask not what your country can do # for you. Ask what you can do for your country. ``` The 4.5 GB Q4_K is the recommended quant; F16 (14.9 GB) needs ~16 GB RAM during inference. JFK matches the upstream Python `MimoAudio.asr_sft` reference verbatim; performance on M1+Metal is ~0.3× realtime (Q4_K dequant per step is the bottleneck — F16 + KV-reuse follow-ups are queued under PLAN #51a/b/c). ### Wav2Vec2 (lightweight CTC, any HF Wav2Vec2ForCTC model) ```bash # English (Q4_K quantized, 212 MB — 6x smaller than F16) curl -L -o wav2vec2-en-q4k.gguf \ https://huggingface.co/cstr/wav2vec2-large-xlsr-53-english-GGUF/resolve/main/wav2vec2-xlsr-en-q4_k.gguf ./build/bin/crispasr -m wav2vec2-en-q4k.gguf -f samples/jfk.wav # and so my fellow americans ask not what your country can do for you ask what you can do for your country # German curl -L -o wav2vec2-de-q4k.gguf \ https://huggingface.co/cstr/wav2vec2-large-xlsr-53-german-GGUF/resolve/main/wav2vec2-xlsr-de-q4_k.gguf ./build/bin/crispasr -m wav2vec2-de-q4k.gguf -f audio.de.wav # Convert any HuggingFace Wav2Vec2ForCTC model: python models/convert-wav2vec2-to-gguf.py \ --model-dir jonatasgrosman/wav2vec2-large-xlsr-53-german \ --output wav2vec2-de.gguf --dtype f32 # Then optionally quantize: ./build/bin/crispasr-quantize wav2vec2-de.gguf wav2vec2-de-q4k.gguf q4_k ``` --- ## Streaming, TTS, and HTTP server CrispASR has three feature areas that warrant their own docs pages: - **[Streaming & live transcription](/CrispASR/docs/streaming.html)** — `--stream`, `--mic`, `--live`, sliding-window chunking, per-token confidence. - **[Text-to-Speech (TTS)](/CrispASR/docs/tts.html)** — Kokoro (multilingual, smallest), Qwen3-TTS (highest fidelity, voice cloning), VibeVoice (lowest-latency streaming), Orpheus (3 B Llama + SNAC), Chatterbox (flow-matching + HiFT vocoder, German via Kartoffelbox), IndexTTS, VoxCPM2, and CosyVoice3 (9 langs + 18 zh dialects; baked-voice bank + arbitrary-WAV cloning). Voice packs, language routing, and qwen3-tts environment switches. All TTS output is watermarked; post-embed verification warns if confidence is low. Use `--detect-watermark file.wav` to check any WAV for AI watermarks. - **[Server mode (HTTP API)](/CrispASR/docs/server.html)** — persistent model, OpenAI-compatible `/v1/audio/transcriptions` (ASR) and `/v1/audio/speech` + `/v1/voices` (TTS, automatic on any loaded CAP_TTS backend), per-request voice + speed + instructions, CORS, long-form sentence chunking, API keys, Docker Compose, prebuilt CUDA images. - **[Concurrency, parallelism & scaling](/CrispASR/docs/concurrency.html)** — one transcription already uses multiple cores; the server accepts requests concurrently but serializes inference on one model by default; `--server-workers N` runs N model instances so pure-ASR requests run concurrently; and for bulk/throughput workloads, process-level fan-out (`xargs -P` / GNU `parallel`) or N replicas behind a load balancer. Also covers what is *not* supported (batched multi-stream inference, PagedAttention) and why. Quickest taste of each: ```bash # Streaming from microphone crispasr --mic -m model.gguf # TTS via auto-downloaded VibeVoice (~636 MB on first run) crispasr --backend vibevoice-tts -m auto --tts "Hello world" --tts-output hello.wav # CosyVoice3 on GPU; companions auto-download beside the LLM crispasr --backend cosyvoice3-tts -m auto --tts "Hello world" --tts-output cosy.wav # CosyVoice3 fast mode: 5 flow steps instead of the quality-default 10 COSYVOICE3_FLOW_STEPS=5 crispasr --backend cosyvoice3-tts -m auto \ --tts "Hello world" --tts-output cosy-fast.wav # Persistent HTTP server, OpenAI-compatible crispasr --server -m model.gguf --port 8080 curl -F "file=@audio.wav" http://localhost:8080/v1/audio/transcriptions # TTS over HTTP — load a TTS backend, hit /v1/audio/speech crispasr --server --backend qwen3-tts-customvoice -m auto --voice-dir ./voices --port 8080 curl http://localhost:8080/v1/audio/speech \ -H 'Content-Type: application/json' \ -d '{"input":"Hello world","voice":"vivian"}' -o out.wav ``` CosyVoice3 uses batched classifier-free guidance and request-sized KV caching by default. Baked voices load only the LLM, flow, HiFT, and voice bank; the larger S3 tokenizer and CAMPPlus companions load lazily when a `.wav` cloning voice is first requested. --- ## CLI reference Common flags: ```bash crispasr -m auto --backend parakeet -f audio.wav --vad -osrt --split-on-punct ``` | Flag | Meaning | |---|---| | `-m FNAME` / `--backend NAME` | Model path (or `auto`) and forced backend | | `-f FNAME` | Input audio (repeatable; positional accepted) | | `--vad` | Silero VAD chunking — strongly recommended for multi-minute audio | | `-osrt` / `-ovtt` / `-otxt` / `-oj` / `-ojf` | Output formats (also `-ocsv`, `-olrc`) | | `-am FNAME` | CTC aligner GGUF for word-level timestamps on LLM backends | | `--align-only` | Standalone forced alignment: text/`.srt` + audio → timestamped SRT/JSON (no ASR needed); `.srt` input keeps its cues and gets re-timed (`--align-granularity auto\|word\|segment`) | | `-tp F` / `-bs N` | Sampling temperature / beam search width | | `-n N` / `--frequency-penalty F` | Generated-token cap / opt-in repeated-token penalty for supported autoregressive ASR backends | | `-l auto` / `--detect-language` | LID pre-step for backends without native lang detect | | `--hotwords "A,B,C"` | Contextual biasing — boost named terms during CTC/TDT decode or LLM prompt | | `-ck N` | Fallback chunk size when VAD is off (default 30 s) | | `--list-backends` | Print the capability matrix and exit | **See [`docs/cli.md`](/CrispASR/docs/cli.html)** for the full reference: every flag, VAD details, CTC alignment workflow, output JSON layout, the auto-download registry, and supported audio formats. **See [`docs/bindings.md`](/CrispASR/docs/bindings.html)** for Python / Rust / Dart / Go / Java / JavaScript / Ruby / mobile. --- ## Architecture, contributing, regression matrix CrispASR is structured as a stable C-ABI in `src/` (every algorithm: VAD, diarize, LID, alignment, cache, registry) consumed by all language wrappers, with thin presentation layers in `examples/cli/`. Per-model runtimes live in `src/{whisper,parakeet,canary,...}.cpp`, sharing primitives from `src/core/` (mel, ffn, attention, GGUF loader, FastConformer / Conformer / Granite-LLM blocks, etc.). - **[`docs/architecture.md`](/CrispASR/docs/architecture.html)** — full layered layout, file-by-file tour of `src/` and `examples/cli/`, per-backend internals table, regression discipline. - **[`docs/contributing.md`](/CrispASR/docs/contributing.md)** — adding a new backend in five files, clang-format-18 setup, the `crispasr-diff` PyTorch-ground-truth workflow, and the TTS audio-cosine-vs-reference regression target. - **[`docs/regression-matrix.md`](/CrispASR/docs/regression-matrix.html)** — `tools/test-all-backends.py` capability tiers, cache modes (`keep` / `ephemeral`), `--skip-missing` for CI. **Shared libraries** (cross-repo with CrispEmbed): - `crisp_audio/` — Whisper-shape audio encoder (Conv-stem + Transformer) - `crisp_punc/` — punctuation restoration (FireRedPunc + PCS) - `crisp_lid/` — text-based language identification (fastText + CLD3) - `crisp_truecase/` — truecasing (statistical + CRF + BiLSTM) Both are self-contained static libraries with CMakeLists.txt. CrispEmbed links them via `add_subdirectory(../CrispASR/crisp_*/)`; CrispASR uses them directly. If the shared dir is absent, both repos fall back to local copies of the source files. For benchmarks see [`PERFORMANCE.md`](/CrispASR/PERFORMANCE.html); for the session-by-session port log and the bug-class lessons, see [`LEARNINGS.md`](/CrispASR/LEARNINGS.html). --- ## Quantize models `build/bin/crispasr-quantize` is a single, model-agnostic GGUF re-quantization tool that works across all supported model families (Whisper, Parakeet, Canary, Cohere, Voxtral, Qwen3, Granite, Wav2Vec2, MiMo-ASR, GLM-ASR, Moonshine, VibeVoice, Kokoro, Qwen3-TTS, …): ```bash ./build/bin/crispasr-quantize input.gguf output.gguf q4_k ``` **See [`docs/quantize.md`](/CrispASR/docs/quantize.html)** for the full guide: supported quant types, K-quant alignment fallback, recommended quant per backend, and worked examples for each architecture. --- ## GPU backend selection All backends use `ggml_backend_init_best()` which automatically picks the highest-priority compiled backend: CUDA > Metal > Vulkan > CPU. To force a specific backend: ```bash # Force Vulkan even when CUDA is available crispasr --gpu-backend vulkan -m model.gguf -f audio.wav # Pin a specific GPU (useful on Vulkan systems with iGPU + dGPU) crispasr --gpu-backend vulkan -dev 1 -m model.gguf -f audio.wav # Force CPU (useful for benchmarking) crispasr -ng -m model.gguf -f audio.wav # CUDA unified memory (swap to RAM when VRAM exhausted) GGML_CUDA_ENABLE_UNIFIED_MEMORY=1 crispasr -m model.gguf -f audio.wav ``` Build flags: `-DGGML_CUDA=ON`, `-DGGML_METAL=ON`, `-DGGML_VULKAN=ON`. Notes: - `--gpu-backend vulkan` selects the Vulkan backend, but it does not choose which physical GPU to use. Use `-dev N` to select the Vulkan device index. - On some Windows laptops, Vulkan device `0` is the Intel iGPU and the NVIDIA GPU is `1`. If Vulkan looks unexpectedly slow, rerun with `-dev 1`. - The Windows convenience script `build-vulkan.bat` creates a separate Vulkan-capable binary at `build-vulkan\bin\crispasr.exe`. --- ## Debugging & profiling For most backends, `-v` / `--verbose` surfaces per-stage timings and device picks. For headless / library use (where the CLI flag isn't plumbed through), set `CRISPASR_VERBOSE=1` instead. ```bash # Per-stage timing breakdown (mel / encoder / prefill / decode): crispasr -v --backend gemma4-e2b -m model.gguf -f audio.wav # gemma4_e2b: mel 128x1099 (17.2 ms) # gemma4_e2b: encoder done: 1536x275 (719.0 ms) # gemma4_e2b: prefill done, first_token=3133 (1464.0 ms) # gemma4_e2b: decoded 25 tokens (7748.3 ms total) # crispasr: transcribed 11.0s audio in 7.75s (1.4x realtime) # Hugging Face access for gated models (Voxtral, Gemma4-E2B, …): HF_TOKEN=hf_xxx crispasr -m auto --backend gemma4-e2b -f audio.wav ``` The server has its own auth env: `CRISPASR_API_KEYS` (see [Server mode](/CrispASR/docs/server.html)).
Per-backend debug / bench / dump-dir env vars (developer) These are useful when porting a new backend or chasing a regression. The `*_BENCH=1` toggles emit per-stage timings even without `-v`; the `*_DEBUG=1` toggles emit per-step diagnostic prints; the `*_DUMP_DIR=` paths write per-stage F32 tensors for diff-testing against a PyTorch reference (see [Debug a new backend against PyTorch ground truth](/CrispASR/docs/contributing.md#debug-a-new-backend-against-pytorch-ground-truth)). | Env var | Purpose | | --- | --- | | `CRISPASR_VERBOSE=1` | Forces verbose mode for any backend (parallel to the `-v` flag). | | `CRISPASR_DUMP_DIR=path/` | Generic per-stage F32 tensor dump for the `crispasr-diff` harness. | | `GEMMA4_E2B_BENCH=1` | Per-stage timings for the Gemma-4-E2B backend. | | `COHERE_BENCH=1` / `COHERE_DEBUG=1` | Cohere transcribe per-stage timings / per-step diagnostics. | | `COHERE_PROF=1` | Cohere graph-level profiling (per-op timings). | | `COHERE_THREADS=N` | Override thread count for the Cohere backend. | | `COHERE_DEVICE=cpu\|cuda\|metal\|vulkan` | Force the Cohere backend onto a specific device. | | `COHERE_DUMP_ATTN=path/` | Dump attention activations for Cohere (used by the diff harness). | | `FIRERED_BENCH=1` | Per-stage timings for the FireRedASR backend. | | `FIREREDPUNC_DEBUG=1` | Per-step diagnostics for the FireRed punctuation post-step. | | `MOONSHINE_STREAMING_BENCH=1` | Per-stage timings for moonshine-streaming. | | `OMNIASR_BENCH=1` / `OMNIASR_DEBUG=1` / `OMNIASR_DUMP_DIR=` | OmniASR per-stage timings, diagnostics, and stage dumps. | | `PARAKEET_DEBUG=1` | Parakeet TDT per-step diagnostics (joint network, blank-id sanity). | | `QWEN3_TTS_BENCH=1` / `QWEN3_TTS_DEBUG=1` / `QWEN3_TTS_DUMP_DIR=` | Qwen3-TTS per-stage timings, diagnostics, and stage dumps. | | `VIBEVOICE_BENCH=1` / `VIBEVOICE_DEBUG=1` / `VIBEVOICE_DUMP_DIR=` | VibeVoice ASR per-stage timings, diagnostics, and stage dumps. | | `VIBEVOICE_REF_FEATURES=path` | Replace the live encoder with a saved feature tensor (regression harness). | | `VIBEVOICE_TTS_DUMP=path/` | VibeVoice TTS per-stage dumps (token IDs, base/TTS hidden, neg condition, frame-0 noise/v_cfg/latent/acoustic_embed) for the diff harness. | | `VIBEVOICE_TTS_DUMP_PERFRAME=1` | Per-frame VibeVoice TTS dumps written as `perframe__f.bin`. Pair with `VIBEVOICE_TTS_DUMP=path/` and `VIBEVOICE_TTS_NOISE=path` for stage-by-stage AR diff against `tools/run_official_vibevoice.py`. | | `VIBEVOICE_TTS_TRACE=1` | Extra one-line traces (negative-condition prefill rms, scaling/bias factors loaded). Same effect as library verbosity ≥ 2; there is no CLI flag for it (`-v` caps verbosity at 1). | | `VIBEVOICE_VOICE_AUDIO=path.wav` | Reference voice WAV for 1.5B-base TTS without a `.gguf` voice cache. | | `VIBEVOICE_TTS_NOISE=path` | Override the per-frame Gaussian init noise. Flat little-endian float32 `[N_frames, vae_dim]` — typically the `noise.bin` written by `tools/run_official_vibevoice.py`. | | `VIBEVOICE_VAE_BACKEND=cpu\|metal\|cuda\|vulkan` | Pin the VAE decoder onto a specific backend. | | `WAV2VEC2_BENCH=1` / `WAV2VEC2_VERBOSE=1` / `WAV2VEC2_DUMP_DIR=` | wav2vec2 per-stage timings, verbose graph traces, and stage dumps. | | `CRISPASR_VOXTRAL4B_STREAM_TIMING=1` | Per-stage timings for the voxtral4b streaming path (encoder drain / prefill / first-text-token / decode-step p50/p95). | | `CRISPASR_VOXTRAL4B_STREAM_CHUNK_MS=N` | Override the internal encoder chunk size (default 240 ms). Must be a multiple of 80 ms. Larger = faster feed (kernel-launch amortisation), longer live-caption latency floor. | | `CRISPASR_VOXTRAL4B_STREAM_BATCH_ENCODER=1` | Regression-debug: ignore the streaming encoder's audio_embeds and re-run the whole batch encoder at flush. | | `CRISPASR_VOXTRAL4B_STREAM_DEBUG=1` / `CRISPASR_VOXTRAL4B_STREAM_DIFF=1` | Per-step decode prints / side-by-side encoder cosine vs the batch encoder. | | `CRISPASR_VOXTRAL4B_STREAM_LIVE=1` | Live-captions decode-during-feed (PLAN #7 phase 3). `get_text()` polled during feed returns progressive transcript. Default OFF (PTT semantics). Wrappers: Python `Session.stream_open(live=True)`, Rust `stream_open_ex(.., live: true)`. | | `CRISPASR_VOXTRAL4B_STREAM_DECODER_THREAD=1` | Decoder worker thread (PLAN #7 phase 4, implies live mode). Lets `feed()` return between encoder chunks without waiting for the decode loop — useful for mic-driven workloads. On M1 the Metal queue serializes encoder and decoder so total wall-clock is unchanged; faster GPUs with kernel-level parallelism see real overlap. | | `CRISPASR_VOXTRAL4B_FUSED_QKV=0` | Opt out of the runtime fused-QKV LLM path (default-on, ~7-8 % decode speedup on M1 Q4_K, ~500 MB extra memory). | | `CRISPASR_QWEN3_ASR_FUSED_QKV=0` | Opt out of the runtime fused-QKV LLM path for qwen3-asr (default-on; works on F16/F32/Q4_K/Q8_0/...). | | `CRISPASR_VOXTRAL_FUSED_QKV=1` | Opt **in** to the runtime fused-QKV LLM path for voxtral 3B. Off by default (no measurable speedup on JFK-shape decodes; useful for long-form workloads where decode dominates). | | `QWEN3_TTS_FUSED_QKV=1` | Opt in to the runtime fused-QKV talker path. | | `GRANITE_DISABLE_ENCODER_GRAPH=1` | Force the granite-speech / -plus / -nar encoder back to the per-layer CPU loop (slower but kept around for debugging). The single ggml-graph encoder with per-layer Shaw RPE is the default and is bit-near-identical to the CPU loop while being ~2× faster end-to-end across all three variants. | | `CRISPASR_NO_REL_POS=1` | Ablate the relative-position bias in the Gemma-4 audio encoder (development only). | | `ECAPA_REF_FBANK=path` | Reference filterbank tensor for the ECAPA-TDNN LID model (regression harness). | | `CRISPASR_SHERPA_LID_BIN=path` | Override the auto-detected sherpa-onnx LID binary. | | `CRISPASR_ARG_DEVICE=N` | Default GPU device index when `-dev` isn't passed. | | `GGML_CUDA_ENABLE_UNIFIED_MEMORY=1` | Let CUDA swap to RAM when VRAM is exhausted. | | `GGML_VK_VISIBLE_DEVICES` / `CUDA_VISIBLE_DEVICES` | Standard ggml/CUDA device-visibility filters. | `HF_TOKEN` and `HUGGING_FACE_HUB_TOKEN` are both honoured for gated-model downloads (in that order). </details> --- ## Credits - **[whisper.cpp](https://github.com/ggml-org/whisper.cpp)** — the original ggml inference engine and Whisper runtime this fork is built on - **[ggml](https://github.com/ggml-org/ggml)** — the tensor library everything runs on - **NVIDIA NeMo** — parakeet-tdt-{0.6b-v2,0.6b-v3,1.1b}, parakeet-tdt_ctc-{110m,1.1b,0.6b-ja}, parakeet-ctc-{0.6b,1.1b}, canary-1b-v2, canary-ctc aligner, and the FastConformer-CTC family (stt_en_fastconformer_ctc_{large,xlarge,xxlarge} plus CTC branches of the stt_*_fastconformer_hybrid_large[_pc] fleet: en-pc, de, es, fr, it, nl, pl, ru, ua, hr, be, ar, fa, ka, hy, uz, kk-ru — all usable both as ASR backends and as compact ~82 MB `-am` forced aligners) - **Cohere** — cohere-transcribe-03-2026 - **Qwen team (Alibaba)** — Qwen3-ASR-0.6B, Qwen3-ASR-1.7B, Qwen3-ForcedAligner-0.6B - **Mistral AI** — Voxtral Mini 3B and 4B Realtime - **IBM Granite team** — Granite Speech 3.2-8b, 3.3-2b, 3.3-8b, 4.0-1b - **Meta / wav2vec2** — wav2vec2 CTC models (XLSR-53 English, German, multilingual via any Wav2Vec2ForCTC checkpoint) - **[sherpa-onnx](https://github.com/k2-fsa/sherpa-onnx)** — optional diarization via subprocess (ONNX models) - **[Silero](https://github.com/snakers4/silero-vad)** — VAD (native GGUF) and language identification (native GGUF, 95 languages) - **[pyannote](https://github.com/pyannote/pyannote-audio)** — speaker diarization segmentation (native GGUF port) - **[miniaudio](https://miniaud.io/)** + **[stb_vorbis](https://github.com/nothings/stb)** + **[libopus/opusfile](https://opus-codec.org/)** — embedded/linked audio decoders (WAV/MP3/FLAC/AIFF/OGG/Opus, no ffmpeg; AAC/M4A/ALAC via Apple AudioToolbox) - **[glint](https://github.com/CrispStrobe/glint)** (MIT) — in-tree clean-room codec suite: MP3/AAC-LC/Ogg-Opus output encoders (TTS `.mp3`/`.aac`/`.opus`) and cross-platform ADTS AAC-LC + Ogg Opus input decoders (no libopus needed) - **[Claude Code](https://claude.ai/claude-code)** (Anthropic) — significant portions of the crispasr integration layer, all model converters, and the FastConformer/attention/mel/FFN/BPE core helpers were co-authored with Claude --- ## License Same as upstream whisper.cpp: **MIT**. Per-model weights are covered by their respective HuggingFace model licenses (see [Supported backends](#supported-backends)). The `crispasr` binary itself links model runtimes that are mostly permissively licensed (MIT / Apache-2.0 / CC-BY-4.0 for weights).