cuda

72 progetti condividono questo topic GitHub

cuda — vllm ★86.7kcudavoicebox — ★44.1ksglang — ★30.6kinstant-ngp — ★17.5kburn — ★15.6kkaldi — ★15.4kTensorRT-LLM — ★14.2kGPU-Puzzles — ★12.3kLMCache — ★10.7kcutlass — ★10.1kcog — ★9.4koneflow — ★9.4kgocv — ★7.5kflashinfer — ★6kchainer — ★5.9kgpustack — ★5.4kcuml — ★5.2knccl — ★4.9kCTranslate2 — ★4.6ktiny-cuda-nn — ★4.5kiree — ★3.8kSageAttention — ★3.5kTransformerEngine — ★3.4kLichtFeld-Studio — ★3.4kjittor — ★3.2khow-to-optim-algorithm-in-cuda — ★3.1kheavydb — ★3.1kTensorRT — ★3kramalama — ★3kCVprojects — ★2.6ktorchrec — ★2.6kpykeen — ★2konediff — ★2kppq — ★1.8ksonar — ★1.8kawesome-yolo-object-detection — ★1.8kbeta9 — ★1.7kcurobo — ★1.7ktt-metal — ★1.6kgpu-hot — ★1.6k3d-ken-burns — ★1.6kvoicebox★ 44.1ksglang★ 30.6kinstant-ngp★ 17.5kburn★ 15.6kkaldi★ 15.4kTensorRT-LLM★ 14.2kGPU-Puzzles★ 12.3kLMCache★ 10.7kcutlass★ 10.1kcog★ 9.4koneflow★ 9.4kgocv★ 7.5kflashinfer★ 6kchainer★ 5.9kgpustack★ 5.4kcuml★ 5.2knccl★ 4.9kCTranslate2★ 4.6ktiny-cuda-nn★ 4.5kiree★ 3.8kSageAttention★ 3.5kTransformerEngine★ 3.4kLichtFeld-Studio★ 3.4kjittor★ 3.2khow-to-optim-algorithm-i…★ 3.1kheavydb★ 3.1kTensorRT★ 3kramalama★ 3kCVprojects★ 2.6ktorchrec★ 2.6kpykeen★ 2konediff★ 2kppq★ 1.8ksonar★ 1.8kawesome-yolo-object-dete…★ 1.8kbeta9★ 1.7kcurobo★ 1.7ktt-metal★ 1.6kgpu-hot★ 1.6k3d-ken-burns★ 1.6k

Le linee collegano membri che sono misurabilmente correlati tra loro. La dimensione dei punti riflette le stelle.

🧬 Membri
vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
★ 86.7k
voicebox
The open-source AI voice studio. Clone, dictate, create.
★ 44.1k
sglang
SGLang is a high-performance serving framework for large language models and multimodal models.
★ 30.6k
instant-ngp
Instant neural graphics primitives: lightning fast NeRF and more
★ 17.5k
burn
Burn is a next generation tensor library and Deep Learning Framework that doesn't compromise on flexibility,…
★ 15.6k
kaldi
kaldi-asr/kaldi is the official location of the Kaldi project.
★ 15.4k
TensorRT-LLM
TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and…
★ 14.2k
GPU-Puzzles
Solve puzzles. Learn CUDA.
★ 12.3k
LMCache
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
★ 10.7k
cutlass
CUDA Templates and Python DSLs for High-Performance Linear Algebra
★ 10.1k
cog
Containers for machine learning
★ 9.4k
oneflow
OneFlow is a deep learning framework designed to be user-friendly, scalable and efficient.
★ 9.4k
gocv
Go package for computer vision using OpenCV 4 and beyond. Includes support for DNN, CUDA, OpenCV Contrib, and…
★ 7.5k
flashinfer
FlashInfer: Kernel Library for LLM Serving
★ 6k
chainer
A flexible framework of neural networks for deep learning
★ 5.9k
gpustack
A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU…
★ 5.4k
cuml
cuML - RAPIDS Machine Learning Library
★ 5.2k
nccl
Optimized primitives for collective multi-GPU communication
★ 4.9k
CTranslate2
Fast inference engine for Transformer models
★ 4.6k
tiny-cuda-nn
Lightning fast C++/CUDA neural network framework
★ 4.5k
iree
A retargetable MLIR-based machine learning compiler and runtime toolkit.
★ 3.8k
SageAttention
[ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to…
★ 3.5k
TransformerEngine
A library for accelerating Transformer models on NVIDIA GPUs, including using 8-bit and 4-bit floating point…
★ 3.4k
LichtFeld-Studio
Train, inspect, edit, automate, and export 3D Gaussian Splatting scenes from a single native application.
★ 3.4k
jittor
Jittor is a high-performance deep learning framework based on JIT compiling and meta-operators.
★ 3.2k
how-to-optim-algorithm-in-cuda
how to optimize some algorithm in cuda.
★ 3.1k
heavydb
HeavyDB (formerly MapD/OmniSciDB)
★ 3.1k
TensorRT
PyTorch/TorchScript/FX compiler for NVIDIA GPUs using TensorRT
★ 3k
ramalama
RamaLama is an open-source developer tool that simplifies the local serving of AI models from any source and…
★ 3k
CVprojects
computer vision projects | 计算机视觉相关好玩的AI项目(Python、C++、embedded system)
★ 2.6k
torchrec
Pytorch domain library for recommendation systems
★ 2.6k
pykeen
🤖 A Python library for learning and evaluating knowledge graph embeddings
★ 2k
onediff
OneDiff: An out-of-the-box acceleration library for diffusion models.
★ 2k
ppq
PPL Quantization Tool (PPQ) is a powerful offline neural network quantization tool.
★ 1.8k
sonar
Large-scale LLM inference engine
★ 1.8k
awesome-yolo-object-detection
🚀🚀🚀 A collection of some awesome public YOLO object detection series projects and the related object…
★ 1.8k
beta9
Ultrafast serverless GPU inference, sandboxes, and background jobs
★ 1.7k
curobo
CUDA Accelerated Robot Library
★ 1.7k
tt-metal
:metal: TT-NN operator library, and TT-Metalium low level kernel programming model.
★ 1.6k
gpu-hot
🔥 Real-time NVIDIA GPU dashboard
★ 1.6k
3d-ken-burns
an implementation of 3D Ken Burns Effect from a Single Image using PyTorch
★ 1.6k
stable-fast
https://wavespeed.ai/ Best inference performance optimization framework for HuggingFace Diffusers on NVIDIA…
★ 1.3k
InferenceX
Open Source Continuous Inference Benchmark Research Platform — Kimi K2.7-Code, MiniMax M3, DeepSeekv4, GLM5…
★ 1.3k
cupoch
Robotics with GPU computing
★ 1.1k
ZhiLight
A highly optimized LLM inference acceleration engine for Llama and its variants.
★ 905
gprMax
gprMax is open source software that simulates electromagnetic wave propagation using the Finite-Difference…
★ 861
UniLab
UniLab: A Heterogeneous Architecture for Robot RL Beyond GPU-Dominant Paradigms
★ 842
Savant
Python Computer Vision & Video Analytics Framework With Batteries Included
★ 837
GPUMD
Graphics Processing Units Molecular Dynamics
★ 811
surogate
Training/Fine-tuning at the speed of light
★ 806
ServerlessLLM
Serverless LLM Serving for Everyone.
★ 692
chatterbox-tts-api
Local, OpenAI-compatible text-to-speech (TTS) API using Chatterbox, enabling users to generate voice cloned…
★ 625
vins-application
VINS-Fusion, VINS-Fisheye, OpenVINS, EnVIO, ROVIO, S-MSCKF, ORB-SLAM2, NVIDIA Elbrus application of different…
★ 607
attorch
A subset of PyTorch's neural network modules, written in Python using OpenAI's Triton.
★ 606
atlas
Pure Rust Inference Engine
★ 601
llm_training_handbook
An open collection of methodologies to help with successful training of large language models.
★ 563
radarsimpy
Radar Simulator built with Python and C++
★ 560
willow-inference-server
Open source, local, and self-hosted highly optimized language inference server supporting ASR/STT, TTS, and…
★ 508
large_language_model_training_playbook
An open collection of implementation tips, tricks and resources for training large language models
★ 502
popsift
PopSift is an implementation of the SIFT algorithm in CUDA.
★ 498
cucim
cuCIM - RAPIDS GPU-accelerated image processing library
★ 463
hoomd-blue
Molecular dynamics and Monte Carlo soft matter simulation on GPUs.
★ 444
dynamicfusion
Implementation of Newcombe et al. CVPR 2015 DynamicFusion paper
★ 413
splatad
SplatAD: Real-Time Lidar and Camera Rendering with 3D Gaussian Splatting for Autonomous Driving
★ 406
MFC
Exascale multiphase flow solver — 2025 Gordon Bell Prize Finalist | 200T grid points on 43K+ GPUs
★ 386
Dia-TTS-Server
Self-host the powerful Dia TTS model. This server offers a user-friendly Web UI, flexible API endpoints…
★ 352
swiftLLM
A tiny yet powerful LLM inference system tailored for researching purpose. vLLM-equivalent performance with…
★ 329
dynamic-occupancy-grid-map
Implementation of "A Random Finite Set Approach for Dynamic Occupancy Grid Maps with Real-Time Application"
★ 310
sndr_core_engine
SNDR Core Engine (Genesis) — vLLM runtime patch-overlay for Qwen3.6 + Gemma4 on consumer NVIDIA (Ampere…
★ 125
self-hosted-ai-stack
Deploy a complete self-hosted AI stack with Docker Compose: Ollama, LiteLLM, AnythingLLM, Whisper,…
★ 125
dgx-spark-inference-stack
Serve the home! Inference stack for your Nvidia DGX Spark aka the Grace Blackwell AI supercomputer on your…
★ 50
trellis2.c
Generate textured, segmented and rigged GLB assets entirely on your own GPU.
★ 45 · GitHub ↗
🔗 Famiglie affini

Misurato dai temi di GitHub condivisi da entrambi i progetti, ponderato in base a quanto è raro ciascun tema.