inference

46 progetti condividono questo topic GitHub

inference — whisper.cpp ★51.8kinferencefaster-whisper — ★24.3kml-engineering — ★18.4knano-vllm — ★14.6kTensorRT — ★13.2kLMCache — ★10.7kargmax-oss-swift — ★6.3kMooncake — ★5.9kvllm-omni — ★5.6kzml — ★3.9kFastVideo — ★3.8kFastDeploy — ★3.7koptimum — ★3.4kopenvino_notebooks — ★3.2khuggingface.js — ★2.5kvllm-ascend — ★2.4kort — ★2.4kAI-Engineering.academy — ★2.4kany-llm — ★2.1kDeepSpeed-MII — ★2.1kaici — ★2.1ktensorflow_template_application — ★1.9kagibot_x1_infer — ★1.8kpicolm — ★1.7ktransformer-deploy — ★1.7kuzu — ★1.7knvidia_gpu_exporter — ★1.5kxllm — ★1.5krtp-llm — ★1.3kkvpress — ★1.1kims — ★938bark.cpp — ★866tensorrt-cpp-api — ★807vidur — ★642Deepdive-llama3-from-scratch — ★631distill-sd — ★618optimum-intel — ★606YOLO-Patch-Based-Inference — ★553SwiftInfer — ★478isaac_ros_pose_estimation — ★475JetStream — ★451faster-whisper★ 24.3kml-engineering★ 18.4knano-vllm★ 14.6kTensorRT★ 13.2kLMCache★ 10.7kargmax-oss-swift★ 6.3kMooncake★ 5.9kvllm-omni★ 5.6kzml★ 3.9kFastVideo★ 3.8kFastDeploy★ 3.7koptimum★ 3.4kopenvino_notebooks★ 3.2khuggingface.js★ 2.5kvllm-ascend★ 2.4kort★ 2.4kAI-Engineering.academy★ 2.4kany-llm★ 2.1kDeepSpeed-MII★ 2.1kaici★ 2.1ktensorflow_template_appl…★ 1.9kagibot_x1_infer★ 1.8kpicolm★ 1.7ktransformer-deploy★ 1.7kuzu★ 1.7knvidia_gpu_exporter★ 1.5kxllm★ 1.5krtp-llm★ 1.3kkvpress★ 1.1kims★ 938bark.cpp★ 866tensorrt-cpp-api★ 807vidur★ 642Deepdive-llama3-from-scr…★ 631distill-sd★ 618optimum-intel★ 606YOLO-Patch-Based-Inferen…★ 553SwiftInfer★ 478isaac_ros_pose_estimatio…★ 475JetStream★ 451

Le linee collegano membri che sono misurabilmente correlati tra loro. La dimensione dei punti riflette le stelle.

🧬 Membri
whisper.cpp
Port of OpenAI's Whisper model in C/C++
★ 51.8k
faster-whisper
Faster Whisper transcription with CTranslate2
★ 24.3k
ml-engineering
Machine Learning Engineering Open Book
★ 18.4k
nano-vllm
Nano vLLM
★ 14.6k
TensorRT
NVIDIA® TensorRT™ is an SDK for high-performance deep learning inference on NVIDIA GPUs. This repository…
★ 13.2k
LMCache
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
★ 10.7k
argmax-oss-swift
On-device Speech AI for Apple Silicon
★ 6.3k
Mooncake
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
★ 5.9k
vllm-omni
A framework for efficient model inference with omni-modality models
★ 5.6k
zml
Any model. Any hardware. Zero compromise. Built with @ziglang / @openxla / MLIR / @bazelbuild
★ 3.9k
FastVideo
A unified inference and post-training framework for accelerated video generation.
★ 3.8k
FastDeploy
High-performance Inference and Deployment Toolkit for LLMs and VLMs based on PaddlePaddle
★ 3.7k
optimum
🚀 Accelerate inference and training of 🤗 Transformers, Diffusers, TIMM and Sentence Transformers with…
★ 3.4k
openvino_notebooks
📚 Jupyter notebook tutorials for OpenVINO™
★ 3.2k
huggingface.js
Use Hugging Face with JavaScript
★ 2.5k
vllm-ascend
Community maintained hardware plugin for vLLM on Ascend
★ 2.4k
ort
Fast ML inference & training for ONNX models in Rust
★ 2.4k
AI-Engineering.academy
Mastering Applied AI, One Concept at a Time
★ 2.4k
any-llm
Communicate with an LLM provider using a single interface
★ 2.1k
DeepSpeed-MII
MII makes low-latency and high-throughput inference possible, powered by DeepSpeed.
★ 2.1k
aici
AICI: Prompts as (Wasm) Programs
★ 2.1k
tensorflow_template_application
TensorFlow template application for deep learning
★ 1.9k
agibot_x1_infer
The inference module for AgiBot X1.
★ 1.8k
picolm
Run a 1-billion parameter LLM on a $10 board with 256MB RAM
★ 1.7k
transformer-deploy
Efficient, scalable and enterprise-grade CPU/GPU inference server for 🤗 Hugging Face transformer models…
★ 1.7k
uzu
A high-performance inference engine for AI models
★ 1.7k
nvidia_gpu_exporter
Nvidia GPU exporter for prometheus using nvidia-smi binary
★ 1.5k
xllm
A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators.…
★ 1.5k
rtp-llm
RTP-LLM: Alibaba's high-performance LLM inference engine for diverse applications.
★ 1.3k
kvpress
LLM KV cache compression made easy
★ 1.1k
ims
📚 Introduction to Modern Statistics - A college-level open-source textbook with a modern approach…
★ 938
bark.cpp
Suno AI's Bark model in C/C++ for fast text-to-speech generation
★ 866
tensorrt-cpp-api
TensorRT C++ API Tutorial
★ 807
vidur
Accurate, large-scale, and extensible simulator for LLM inference Systems
★ 642
Deepdive-llama3-from-scratch
Achieve the llama3 inference step-by-step, grasp the core concepts, master the process derivation, implement…
★ 631
distill-sd
Segmind Distilled diffusion
★ 618
optimum-intel
🤗 Optimum Intel: Accelerate inference with Intel optimization tools
★ 606
YOLO-Patch-Based-Inference
Python library for YOLO small object detection and instance segmentation
★ 553
SwiftInfer
Efficient AI Inference & Serving
★ 478
isaac_ros_pose_estimation
Deep learned, NVIDIA-accelerated 3D object pose estimation
★ 475
JetStream
JetStream is a throughput and memory optimized engine for LLM inference on XLA devices, starting with TPUs…
★ 451
nanoowl
A project that optimizes OWL-ViT for real-time inference with NVIDIA TensorRT.
★ 440
KIVI
[ICML 2024] KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
★ 418
super-rag
Super performant RAG pipelines for AI apps. Summarization, Retrieve/Rerank and Code Interpreters in one…
★ 395
gpmp2
Gaussian Process Motion Planner 2
★ 357
swiftLLM
A tiny yet powerful LLM inference system tailored for researching purpose. vLLM-equivalent performance with…
★ 329
🔗 Famiglie affini

Misurato dai temi di GitHub condivisi da entrambi i progetti, ponderato in base a quanto è raro ciascun tema.