llm-inference

15 projets partagent ce topic GitHub

llm-inference — gpt4all ★77.4kllm-inferencellm-action — ★24.8kPowerInfer — ★9.7kGenerativeAIExamples — ★4.1kMedusa — ★2.8kEAGLE — ★2.5kllama2-webui — ★1.9kreact-native-executorch — ★1.7kblast — ★777LLM-PowerHouse-A-Curated-Guide-for-Large-Language-Models-with-Custom-Training-and-Inferencing — ★730Star-Attention — ★392embedding_studio — ★382NanoLLM — ★379syncode — ★337MoE-Infinity — ★319llm-action★ 24.8kPowerInfer★ 9.7kGenerativeAIExamples★ 4.1kMedusa★ 2.8kEAGLE★ 2.5kllama2-webui★ 1.9kreact-native-executorch★ 1.7kblast★ 777LLM-PowerHouse-A-Curated…★ 730Star-Attention★ 392embedding_studio★ 382NanoLLM★ 379syncode★ 337MoE-Infinity★ 319

Les traits relient les membres réellement apparentés entre eux. La taille des points suit les étoiles.

🧬 Membres
gpt4all
GPT4All: Run Local LLMs on Any Device. Open-source and available for commercial use.
★ 77.4k
llm-action
★ 24.8k
PowerInfer
High-speed Large Language Model Serving for Local Deployment
★ 9.7k
GenerativeAIExamples
Generative AI reference workflows optimized for accelerated infrastructure and microservice architecture.
★ 4.1k
Medusa
Medusa: Simple Framework for Accelerating LLM Generation with Multiple Decoding Heads
★ 2.8k
EAGLE
Official Implementation of EAGLE-1 (ICML'24), EAGLE-2 (EMNLP'24), and EAGLE-3 (NeurIPS'25).
★ 2.5k
llama2-webui
Run any Llama 2 locally with gradio UI on GPU or CPU from anywhere (Linux/Windows/Mac). Use `llama2-wrapper`…
★ 1.9k
react-native-executorch
Declarative way to run AI models in React Native on device, powered by ExecuTorch.
★ 1.7k
blast
Open-source VMs-as-a-service
★ 777
LLM-PowerHouse-A-Curated-Guide-for-Large-Language-Models-with-Custom-Training-and-Inferencing
LLM-PowerHouse: Unleash LLMs' potential through curated tutorials, best practices, and ready-to-use code for…
★ 730
Star-Attention
Efficient LLM Inference over Long Sequences
★ 392
embedding_studio
Embedding Studio is a framework which allows you transform your Vector Database into a feature-rich Search…
★ 382
NanoLLM
Optimized local inference for LLMs with HuggingFace-like APIs for quantization, vision/language models,…
★ 379
syncode
Efficient and general syntactical decoding for Large Language Models
★ 337
MoE-Infinity
PyTorch library for cost-effective, fast and easy serving of MoE models.
★ 319
🔗 Familles voisines

Mesuré à partir des topics GitHub communs aux deux projets, pondérés par leur rareté.