Automatic Speech Recognition
NeMo
PyTorch
English
speech
audio
CTC
FastConformer
Transformer
NeMo
hf-asr-leaderboard
Eval Results (legacy)
Instructions to use nvidia/stt_en_fastconformer_ctc_xxlarge with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- NeMo
How to use nvidia/stt_en_fastconformer_ctc_xxlarge with NeMo:
import nemo.collections.asr as nemo_asr asr_model = nemo_asr.models.ASRModel.from_pretrained("nvidia/stt_en_fastconformer_ctc_xxlarge") transcriptions = asr_model.transcribe(["file.wav"]) - Notebooks
- Google Colab
- Kaggle
GGUF + pure-C++ runtime in CrispASR — FastConformer-CTC XXL
#2
by cstr - opened
We've added the FastConformer-CTC XXL size to CrispASR. C++ binary, GGUF — no NeMo.
Same fastconformer-ctc backend as the large/xlarge siblings — shared core/fastconformer.h encoder (conv subsampling + MHA with rel-PE), greedy CTC head. CTC = no autoregression, no native punctuation.
Punc recipe for any CTC backend in CrispASR: --punc-model fireredpunc-q8_0.gguf (BERT-base, EN+CN) or fullstop-punc-q4_k.gguf (XLM-R-large, EN/DE/FR/IT). Both ship as GGUF.
Pre-quantised GGUFs (CC-BY-4.0): cstr/stt-en-fastconformer-ctc-xxlarge-GGUF
./build/bin/crispasr --backend fastconformer-ctc \
-m stt-en-fastconformer-ctc-xxlarge-q4_k.gguf \
-f audio.wav --punc-model fireredpunc-q8_0.gguf -osrt