audio

57 projetos partilham este topic do GitHub

audio — transformers ★162.8kaudioVoxCPM — ★33.5ksrs — ★29.1kFunASR — ★19.4kspeechbrain — ★11.7kopenFrameworks — ★10.4kAudioGPT — ★10.2kspeech_recognition — ★9kpyAudioAnalysis — ★6.3kpedalboard — ★6.2kbasic-pitch — ★5.3kDeepFilterNet — ★4.5kClearerVoice-Studio — ★4.3kdistil-whisper — ★4.1kriffusion-hobby — ★3.9kMOSS-TTS — ★3.9kSimpleMem — ★3.7kaudioFlux — ★3.3kawesome-deep-learning-music — ★3kaudio — ★2.9kaeneas — ★2.9kScriberr — ★2.8kAutomatic_Speech_Recognition — ★2.8kNeuralNote — ★2.8kriffusion-app-hobby — ★2.7kui — ★2.3kaudiomentations — ★2.3kMMAudio — ★2.2kdescript-audio-codec — ★1.8kwhisper-turbo — ★1.8kRAVE — ★1.8kvocal-remover — ★1.8kSALMONN — ★1.5kAngelSlim — ★1.5kFun-ASR — ★1.4kspotlight — ★1.3kSincNet — ★1.2klhotse — ★1.1kTTS-Audio-Suite — ★1.1kCrisperWhisper — ★968Confucius4-TTS — ★691VoxCPM★ 33.5ksrs★ 29.1kFunASR★ 19.4kspeechbrain★ 11.7kopenFrameworks★ 10.4kAudioGPT★ 10.2kspeech_recognition★ 9kpyAudioAnalysis★ 6.3kpedalboard★ 6.2kbasic-pitch★ 5.3kDeepFilterNet★ 4.5kClearerVoice-Studio★ 4.3kdistil-whisper★ 4.1kriffusion-hobby★ 3.9kMOSS-TTS★ 3.9kSimpleMem★ 3.7kaudioFlux★ 3.3kawesome-deep-learning-mu…★ 3kaudio★ 2.9kaeneas★ 2.9kScriberr★ 2.8kAutomatic_Speech_Recogni…★ 2.8kNeuralNote★ 2.8kriffusion-app-hobby★ 2.7kui★ 2.3kaudiomentations★ 2.3kMMAudio★ 2.2kdescript-audio-codec★ 1.8kwhisper-turbo★ 1.8kRAVE★ 1.8kvocal-remover★ 1.8kSALMONN★ 1.5kAngelSlim★ 1.5kFun-ASR★ 1.4kspotlight★ 1.3kSincNet★ 1.2klhotse★ 1.1kTTS-Audio-Suite★ 1.1kCrisperWhisper★ 968Confucius4-TTS★ 691

Linhas conectam membros que estão mensuravelmente relacionados entre si. O tamanho do ponto reflete estrelas.

🧬 Membros
transformers
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text,…
★ 162.8k
VoxCPM
VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life…
★ 33.5k
srs
SRS is a simple, high-efficiency, real-time media server supporting RTMP, WebRTC, HLS, HTTP-FLV, HTTP-TS,…
★ 29.1k
FunASR
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker…
★ 19.4k
speechbrain
A PyTorch-based Speech Toolkit
★ 11.7k
openFrameworks
openFrameworks is a community-developed cross platform toolkit for creative coding in C++.
★ 10.4k
AudioGPT
AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head
★ 10.2k
speech_recognition
Speech recognition module for Python, supporting several engines and APIs, online and offline.
★ 9k
pyAudioAnalysis
Python Audio Analysis Library: Feature Extraction, Classification, Segmentation and Applications
★ 6.3k
pedalboard
🎛 🔊 A Python library for audio.
★ 6.2k
basic-pitch
A lightweight yet powerful audio-to-MIDI converter with pitch bend detection
★ 5.3k
DeepFilterNet
Noise supression using deep filtering
★ 4.5k
ClearerVoice-Studio
An AI-Powered Speech Processing Toolkit and Open Source SOTA Pretrained Models, Supporting Speech…
★ 4.3k
distil-whisper
Distilled variant of Whisper for speech recognition. 6x faster, 50% smaller, within 1% word error rate.
★ 4.1k
riffusion-hobby
Stable diffusion for real-time music generation
★ 3.9k
MOSS-TTS
MOSS‑TTS Family is an open‑source speech and sound generation model family from MOSI.AI and the OpenMOSS…
★ 3.9k
SimpleMem
SimpleMem: Efficient Lifelong Memory for LLM Agents — Text & Multimodal
★ 3.7k
audioFlux
A library for audio and music analysis, feature extraction.
★ 3.3k
awesome-deep-learning-music
List of articles related to deep learning applied to music
★ 3k
audio
Data manipulation and transformation for audio signal processing, powered by PyTorch
★ 2.9k
aeneas
aeneas is a Python/C library and a set of tools to automagically synchronize audio and text (aka forced…
★ 2.9k
Scriberr
Self-hosted AI audio transcription
★ 2.8k
Automatic_Speech_Recognition
End-to-end Automatic Speech Recognition for Madarian and English in Tensorflow
★ 2.8k
NeuralNote
Audio Plugin for Audio to MIDI transcription using deep learning.
★ 2.8k
riffusion-app-hobby
Stable diffusion for real-time music generation (web app)
★ 2.7k
ui
ElevenLabs UI is a component library and custom registry built on top of shadcn/ui to help you build…
★ 2.3k
audiomentations
A Python library for audio data augmentation. Useful for making audio ML models work well in the real world,…
★ 2.3k
MMAudio
[CVPR 2025] MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis
★ 2.2k
descript-audio-codec
State-of-the-art audio codec with 90x compression factor. Supports 44.1kHz, 24kHz, and 16kHz mono/stereo…
★ 1.8k
whisper-turbo
Cross-Platform, GPU Accelerated Whisper 🏎️
★ 1.8k
RAVE
Official implementation of the RAVE model: a Realtime Audio Variational autoEncoder
★ 1.8k
vocal-remover
Vocal Remover using Deep Neural Networks
★ 1.8k
SALMONN
SALMONN family: A suite of advanced multi-modal LLMs
★ 1.5k
AngelSlim
Model compression toolkit engineered for enhanced usability, comprehensiveness, and efficiency.
★ 1.5k
Fun-ASR
Open-source LLM-based ASR model family for Chinese, dialect, accent, and multilingual speech, with FunASR,…
★ 1.4k
spotlight
Interactively explore unstructured datasets from your dataframe.
★ 1.3k
SincNet
SincNet is a neural architecture for efficiently processing raw audio samples.
★ 1.2k
lhotse
Tools for handling multimodal data in machine learning projects.
★ 1.1k
TTS-Audio-Suite
A ComfyUI custom node integration for local multi-engine multi-language Text-to-Speech and Voice Conversion.…
★ 1.1k
CrisperWhisper
Verbatim Automatic Speech Recognition with improved word-level timestamps and filler detection
★ 968
Confucius4-TTS
Confucius4-TTS: a Multilingual and Cross-Lingual Zero-Shot TTS Engine
★ 691
free-spoken-digit-dataset
A free audio dataset of spoken digits. An audio version of MNIST.
★ 678
ComfyUI-VibeVoice
ComfyUI custom node for the VibeVoice TTS. Expressive, long-form, multi-speaker conversational audio
★ 590
SwiftSpeech
A speech recognition framework designed for SwiftUI.
★ 531
audio-transformers-course
The Hugging Face Course on Transformers for Audio
★ 508
ComfyUI-VoxCPM
ComfyUI node for highly expressive speech and realistic zero-shot voice cloning
★ 497
whisplay-ai-chatbot
Pocket-sized AI chatbot built using a RPI Zero 2w / 5
★ 485
ltu
Code, Dataset, and Pretrained Models for Audio and Speech Large Language Model "Listen, Think, and…
★ 478
huggingsound
HuggingSound: A toolkit for speech-related tasks based on Hugging Face's tools
★ 470
google-speech-v2
:speech_balloon: Reverse Engineering Google's Speech To Text API (v2)
★ 469
speech_dataset
The dataset of Speech Recognition
★ 463
whisper-at
Code and Pretrained Models for Interspeech 2023 Paper "Whisper-AT: Noise-Robust Automatic Speech Recognizers…
★ 421
Synthalingua
Synthalingua - Real Time Translation
★ 403
gazelle
Joint speech-language model - respond directly to audio!
★ 374
xcodec
AAAI 2025: Codec Does Matter: Exploring the Semantic Shortcoming of Codec for Audio Language Model
★ 308
LLMVoX
LLMVoX: Autoregressive Streaming Text-to-Speech Model for Any LLM
★ 308
Awesome-AVI
Awesome Audio-Visual Intelligence, Survey of Audio-Visual Intelligence
★ 83
🔗 Familias relacionadas

Medido a partir dos tópicos do GitHub compartilhados por ambos os projetos, ponderado pela raridade de cada tópico.