Instructions to use ncannings/parakeet-ultra-nemo with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- NeMo
How to use ncannings/parakeet-ultra-nemo with NeMo:
import nemo.collections.asr as nemo_asr asr_model = nemo_asr.models.ASRModel.from_pretrained("ncannings/parakeet-ultra-nemo") transcriptions = asr_model.transcribe(["file.wav"]) - TensorRT
How to use ncannings/parakeet-ultra-nemo with TensorRT:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
parakeet-ultra in NeMo format
moondream/parakeet-ultra (Moondream's post-trained
nvidia/parakeet-tdt-0.6b-v3, 25 languages) converted to a NeMo
.nemo checkpoint, so it runs in NVIDIA NeMo and in fastconformer-trt.
The weights are Moondream's, unchanged and in full precision. The conversion places every tensor into the v3 NeMo
model and refuses unless every NeMo parameter is filled exactly once with a matching shape. Moondream's small
voice-activity head is not part of the NeMo architecture and is left out, so use your own segmentation for long
audio. Conversion script: engine/ultra_to_nemo.py.
Use with NeMo
import nemo.collections.asr as nemo_asr
from huggingface_hub import hf_hub_download
path = hf_hub_download("ncannings/parakeet-ultra-nemo", "ultra.nemo")
model = nemo_asr.models.ASRModel.restore_from(path)
print(model.transcribe(["speech.wav"])[0].text)
Fast inference: fastconformer-trt
fastconformer-trt builds an FP8 TensorRT engine with custom CUDA plugins from this checkpoint. TensorRT engines are specific to the GPU and TensorRT version, so no engine is published here: build it on the machine that will run it (instructions in the repository).
| GPU | Stock NeMo | fastconformer-trt | Speed-up |
|---|---|---|---|
| NVIDIA DGX Spark (GB10) | 999x | 4,903x | 4.9x |
| NVIDIA GH200 | 3,873x | 22,776x | 5.9x |
| NVIDIA H100 SXM | 4,693x | 20,667x | 4.4x |
| NVIDIA B200 | 5,743x | 25,073x | 4.4x |
Real-time factor on LibriSpeech test-clean, timed from audio already in GPU memory to text. Accuracy matches stock NeMo on the same scorer: FLEURS mean over 25 languages 12.60% against 12.55%, full Earnings-22 10.07% against 10.03%. These are scored with the Whisper normalisers, so they differ from the Open ASR Leaderboard figures on Moondream's model card; compare only within one scorer. Full method and per-language results are in the repository.
Licence and attribution
CC-BY-4.0, as the original. Model by Moondream, based on NVIDIA parakeet-tdt-0.6b-v3. Converted by Nigel Cannings; the only change is the file format.
- Downloads last month
- 21
Model tree for ncannings/parakeet-ultra-nemo
Base model
moondream/parakeet-ultra