Step 1: Language and task
Different model families cover different languages. Some models support dozens of languages while others specialize in a narrower language set. Identify which languages you need to transcribe and whether your task is transcription, translation, or both.
The signed catalog includes families such as Whisper, Qwen3-ASR, Parakeet-TDT, FireRed-AED, SenseVoice, Moonshine, Cohere Transcribe, X-ASR, Dolphin, and Hy-MT2 speech translation, plus capability packs for diarization and Chinese punctuation.
Step 2: Hardware and memory
Model size and quantization create different memory, storage, speed, and fidelity tradeoffs; size alone does not guarantee better results. The catalog marks a recommended quantization for each model. Smaller quantizations reduce the pack size and may trade some fidelity.
Check the Models page for exact pack sizes and measured fields to confirm that a model fits your available memory and storage.
Step 3: Features
Consider whether you need word-level timestamps, automatic punctuation, speech translation, or speaker diarization. Not every model supports every feature. Check the model detail page before choosing.
Step 4: License
Models in the catalog carry different licenses. Some are permissive for commercial use; others have research or non-commercial restrictions. Review the license field on the Models page before selecting a model for production use.