Hugging Face's logo Hugging Face
  • Models
  • Datasets
  • Spaces
  • Buckets new
  • Docs
  • Enterprise
  • Pricing
    • Website
      • Tasks
      • HuggingChat
      • Collections
      • Languages
      • Organizations
    • Community
      • Blog
      • Posts
      • Daily Papers
      • Hardware
      • Learn
      • Discord
      • Forum
      • GitHub
    • Solutions
      • Team & Enterprise
      • Hugging Face PRO
      • Enterprise Support
      • Inference Providers
      • Inference Endpoints
      • Storage Buckets

  • Log In
  • Sign Up
blanchefort 's Collections
decision classifier
ML
Embedders
Medical
VLA models
Audio
Translate
OCR
OmniModels
Edge models
Video encoders
Judge
Datasets for Embodied
Ru text encoders
Text2Image
VLMs

Audio

updated 3 days ago
Upvote
-

  • nvidia/audio-flamingo-3-hf

    Audio-Text-to-Text • 8B • Updated Apr 13 • 47.8k • 196

  • facebook/sam-audio-large

    Updated Dec 30, 2025 • 11.9k • 467

  • google/medasr

    Automatic Speech Recognition • 0.1B • Updated May 26 • 17.8k • 372

  • FunAudioLLM/Fun-CosyVoice3-0.5B-2512

    Text-to-Speech • Updated Feb 3 • 174k • 658

  • facebook/sam-audio-large-tv

    Updated Dec 30, 2025 • 4.45k • 34

  • Qwen/Qwen3-TTS-12Hz-0.6B-Base

    Text-to-Speech • 0.9B • Updated Jan 29 • 585k • 324

  • lab260/spectra_0

    Audio Classification • 0.3B • Updated Jun 25 • 7

  • nvidia/Nemotron-3-Diarization

    Voice Activity Detection • 99.2M • Updated 18 days ago • 74.9k • 755

  • m-a-p/MERT-v2-FullSong

    Feature Extraction • 0.6B • Updated 12 days ago • 20.5k • 49

  • coriollon/whisper-large-v3-turbo-russian

    Automatic Speech Recognition • 0.8B • Updated 28 days ago • 768 • 5

  • SpragAI/qwen3-tts-emotion-tags

    Text-to-Speech • 2B • Updated Sep 9 • 201 • 8
Upvote
-
  • Collection guide
  • Browse collections
Company
TOS Privacy About Careers
Website
Models Datasets Spaces Pricing Docs