asr

90 progetti condividono questo topic GitHub

asr — whisperX ★23.1kasrFunASR — ★19.4kSpeech — ★17.8kvosk-api — ★15ksherpa-onnx — ★13.6kPaddleSpeech — ★12.6kspeechbrain — ★11.7kSenseVoice — ★8.9kOmniVoice-Studio — ★8.6kFunClip — ★6kwhisper-diarization — ★5.6kwenet — ★5.2kLLPlayer — ★3.9kStreamer-Sales — ★3.7kwhisper-asr-webservice — ★3.3kwhisper-standalone-win — ★3.1kfaster-whisper-GUI — ★3klingvo — ★2.9kwhisper-timestamped — ★2.8kopenless — ★2.8kSTT — ★2.6kpytorch-kaldi — ★2.4kAudioNotes — ★2.2kFireRedASR — ★1.9ksherpa-ncnn — ★1.8kbailing — ★1.7kdsnote — ★1.5kSpeech-AI-Forge — ★1.4kFun-ASR — ★1.4kSoniTranslate — ★1.4kGPA — ★1.3kStreamSpeech — ★1.3kvosk-server — ★1.3kSincNet — ★1.2kWhisper-Finetune — ★1.2kconformer — ★1.1kvosk-android-demo — ★1.1kpykaldi — ★1kspeech-swift — ★1kathena — ★969CrisperWhisper — ★968FunASR★ 19.4kSpeech★ 17.8kvosk-api★ 15ksherpa-onnx★ 13.6kPaddleSpeech★ 12.6kspeechbrain★ 11.7kSenseVoice★ 8.9kOmniVoice-Studio★ 8.6kFunClip★ 6kwhisper-diarization★ 5.6kwenet★ 5.2kLLPlayer★ 3.9kStreamer-Sales★ 3.7kwhisper-asr-webservice★ 3.3kwhisper-standalone-win★ 3.1kfaster-whisper-GUI★ 3klingvo★ 2.9kwhisper-timestamped★ 2.8kopenless★ 2.8kSTT★ 2.6kpytorch-kaldi★ 2.4kAudioNotes★ 2.2kFireRedASR★ 1.9ksherpa-ncnn★ 1.8kbailing★ 1.7kdsnote★ 1.5kSpeech-AI-Forge★ 1.4kFun-ASR★ 1.4kSoniTranslate★ 1.4kGPA★ 1.3kStreamSpeech★ 1.3kvosk-server★ 1.3kSincNet★ 1.2kWhisper-Finetune★ 1.2kconformer★ 1.1kvosk-android-demo★ 1.1kpykaldi★ 1kspeech-swift★ 1kathena★ 969CrisperWhisper★ 968

Le linee collegano membri che sono misurabilmente correlati tra loro. La dimensione dei punti riflette le stelle.

🧬 Membri
whisperX
WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
★ 23.1k
FunASR
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker…
★ 19.4k
Speech
A scalable generative AI framework built for researchers and developers working on Large Language Models,…
★ 17.8k
vosk-api
Offline speech recognition API for Android, iOS, Raspberry Pi and servers with Python, Java, C# and Node
★ 15k
sherpa-onnx
Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using…
★ 13.6k
PaddleSpeech
Easy-to-use Speech Toolkit including Self-Supervised Learning model, SOTA/Streaming ASR with punctuation,…
★ 12.6k
speechbrain
A PyTorch-based Speech Toolkit
★ 11.7k
SenseVoice
Open-source SenseVoiceSmall model for Mandarin, Cantonese, English, Japanese, and Korean ASR, language ID,…
★ 8.9k
OmniVoice-Studio
Local voice clone, video dubbing, dictation and audiobook maker. The open-source ElevenLabs alternative.
★ 8.6k
FunClip
FunASR-powered video transcription, subtitle generation, and LLM-assisted clipping tool with a local Gradio…
★ 6k
whisper-diarization
Automatic Speech Recognition with Speaker Diarization based on OpenAI Whisper
★ 5.6k
wenet
Production First and Production Ready End-to-End Speech Recognition Toolkit
★ 5.2k
LLPlayer
The media player for language learning, with dual subtitles, AI-generated subtitles, real-time translation,…
★ 3.9k
Streamer-Sales
★ 3.7k
whisper-asr-webservice
OpenAI Whisper ASR Webservice API
★ 3.3k
whisper-standalone-win
Whisper & Faster-Whisper standalone executables for those who don't want to bother with Python.
★ 3.1k
faster-whisper-GUI
faster_whisper GUI with PySide6
★ 3k
lingvo
Lingvo
★ 2.9k
whisper-timestamped
Multilingual Automatic Speech Recognition with word-level timestamps and confidence
★ 2.8k
openless
Hold a key, speak, release — AI-polished text appears at your cursor in any app. Open-source voice input…
★ 2.8k
STT
🐸STT - The deep learning toolkit for Speech-to-Text. Training and deploying STT models has never been so…
★ 2.6k
pytorch-kaldi
pytorch-kaldi is a project for developing state-of-the-art DNN/RNN hybrid speech recognition systems. The DNN…
★ 2.4k
AudioNotes
快速提取音视频内容,整理成一份结构化的markdown笔记
★ 2.2k
FireRedASR
Open-source industrial-grade ASR models supporting Mandarin, Chinese dialects and English, achieving a new…
★ 1.9k
sherpa-ncnn
Real-time speech recognition and voice activity detection (VAD) using next-gen Kaldi with ncnn without…
★ 1.8k
bailing
百聆 是一个类似GPT-4o的语音对话机器人,通过ASR+LLM+TTS实现,集成DeepSeek…
★ 1.7k
dsnote
Speech Note Linux app. Note taking, reading and translating with offline Speech to Text, Text to Speech and…
★ 1.5k
Speech-AI-Forge
🍦 Speech-AI-Forge is a project developed around TTS generation model, implementing an API Server and a…
★ 1.4k
Fun-ASR
Open-source LLM-based ASR model family for Chinese, dialect, accent, and multilingual speech, with FunASR,…
★ 1.4k
SoniTranslate
Synchronized Translation for Videos. Video dubbing
★ 1.4k
GPA
[AutoArk] GPA (General Purpose Audio) can do ASR, TTS and voice conversion with one tiny model!
★ 1.3k
StreamSpeech
StreamSpeech is an “All in One” seamless model for offline and simultaneous speech recognition, speech…
★ 1.3k
vosk-server
WebSocket, gRPC and WebRTC speech recognition server based on Vosk and Kaldi libraries
★ 1.3k
SincNet
SincNet is a neural architecture for efficiently processing raw audio samples.
★ 1.2k
Whisper-Finetune
Fine-tune the Whisper speech recognition model to support training without timestamp data, training with…
★ 1.2k
conformer
[Unofficial] PyTorch implementation of "Conformer: Convolution-augmented Transformer for Speech Recognition"…
★ 1.1k
vosk-android-demo
Offline speech recognition for Android with Vosk library.
★ 1.1k
pykaldi
A Python wrapper for Kaldi
★ 1k
speech-swift
AI speech toolkit for Apple Silicon — ASR, TTS, speech-to-speech, VAD, and diarization powered by MLX and…
★ 1k
athena
an open-source implementation of sequence-to-sequence based speech processing engine
★ 969
CrisperWhisper
Verbatim Automatic Speech Recognition with improved word-level timestamps and filler detection
★ 968
sherpa
Speech-to-text server framework with next-gen Kaldi
★ 959
espresso
Espresso: A Fast End-to-End Neural Speech Recognition Toolkit
★ 939
PPASR
★ 872
eesen
The official repository of the Eesen project
★ 834
cn2an
★ 764
PaddlePaddle-DeepSpeech
★ 761
whisper.unity
Running speech to text model (whisper.cpp) in Unity3d on your local machine.
★ 749
chinese_text_normalization
Chinese text normalization for speech processing
★ 733
MASR
★ 725
openspeech
Open-Source Toolkit for End-to-End Speech Recognition leveraging PyTorch-Lightning and Hydra.
★ 716
INTERSPEECH-2023-24-Papers
INTERSPEECH 2023-2024 Papers: A complete collection of influential and exciting research papers from the…
★ 685
whisper_android
Offline Speech Recognition with OpenAI Whisper and TensorFlow Lite for Android
★ 676
cheetah
On-device streaming speech-to-text engine powered by deep learning
★ 669
kospeech
Open-Source Toolkit for End-to-End Korean Automatic Speech Recognition leveraging PyTorch and Hydra.
★ 637
FireRedASR2S
A SOTA Industrial-Grade All-in-One ASR system with ASR, VAD, LID, and Punc modules. FireRedASR2 supports…
★ 600
neural_sp
End-to-end ASR/LM implementation with PyTorch
★ 594
WhisperS2T
An Optimized Speech-to-Text Pipeline for the Whisper Model Supporting Multiple Inference Engine
★ 577
Leaderboard
SpeechIO Leaderboard: a large, robust, comprehensive, benchmarking platform for Automatic Speech Recognition.
★ 547
vosk-browser
A speech recognition library running in the browser thanks to a WebAssembly build of Vosk
★ 526
ICASSP-2023-24-Papers
ICASSP 2023-2024 Papers: A complete collection of influential and exciting research papers from the ICASSP…
★ 525
Audar-ASR-V1
Arabic-first generative speech recognition — Audar-ASR-V1 (Flash + Turbo). #1 on the Open Universal Arabic…
★ 513
whisplay-ai-chatbot
Pocket-sized AI chatbot built using a RPI Zero 2w / 5
★ 485
leopard
On-device speech-to-text engine powered by deep learning
★ 482
Fast-Powerful-Whisper-AI-Services-API
⚡ 一款用于自动语音识别 (ASR)、翻译的高性能异步 API。不需要购买Whisper…
★ 470
huggingsound
HuggingSound: A toolkit for speech-related tasks based on Hugging Face's tools
★ 470
speech_dataset
The dataset of Speech Recognition
★ 463
docker-whisperX
Dockerfile for WhisperX: Automatic Speech Recognition with Word-Level Timestamps and Speaker Diarization…
★ 452
deepgram-python-sdk
Official Python SDK for Deepgram.
★ 450
zamia-speech
Open tools and data for cloudless automatic speech recognition
★ 448
LiveTranslate
Real-time audio translation, captures system audio + mic, runs ASR (Whisper/SenseVoice), translates via LLM…
★ 437
reverb
Open source inference code for Rev's model
★ 436
tevr-asr-tool
State-of-the-art (ranked #1 Aug 2022) German Speech Recognition in 284 lines of C++. This is a 100% private…
★ 410
PreenCut
AI-Powered Video Retrieval & Clipping Tool
★ 405
awesome-russian-speech
Russian speech technology links
★ 404
Stream-Omni
Stream-Omni is a GPT-4o-like language-vision-speech chatbot that simultaneously supports interaction across…
★ 389
xiaoniu
小牛视频翻译 是一款支持本地视频翻译、字幕翻译和 YouTube 视频翻译下载的 AI…
★ 383
wav2vec2-live
A live speech recognition using Facebooks wav2vec 2.0 model.
★ 378
hms-ml-demo
HMS ML Demo provides an example of integrating Huawei ML Kit service into applications. This example…
★ 372
parakeet-rs
very fast speech-to-text, diarization, streaming (even in CPU) with NVIDIA Parakeet in Rust
★ 367
whisper-finetune
Fine-tune and evaluate Whisper models for Automatic Speech Recognition (ASR) on custom datasets or datasets…
★ 365
LangHelper
Striving to create a great Application with full functions of learning languages by ChatGPT, TTS, STT and…
★ 349
WhisperHallu
Experimental code: sound file preprocessing to optimize Whisper transcriptions without hallucinated texts
★ 349
izwi
Voice AI runtime. Local first transcription, speaker diarization, TTS, and voice cloning with an OpenAI…
★ 349
onnx-asr
A lightweight Python package for Automatic Speech Recognition using ONNX models
★ 345
Speech-to-Text-Russian
Проект для распознавания речи на русском языке на основе…
★ 345
tensorflow_end2end_speech_recognition
End-to-End speech recognition implementation base on TensorFlow (CTC, Attention, and MTL training)
★ 314
UltraEval-Audio
Your faithful, impartial partner for audio evaluation — know yourself, know your rivals.…
★ 307
end2end-asr-pytorch
End-to-End Automatic Speech Recognition on PyTorch
★ 304
voiceai
Set of 📝 with 🔗 to help those building Voice AI agents 🎙️🤖
★ 301
🔗 Famiglie affini

Misurato dai temi di GitHub condivisi da entrambi i progetti, ponderato in base a quanto è raro ciascun tema.