Home Projects tiny-vllm
tiny-vllm

tiny-vllm

by jmaczan · GitHub

Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM

# ai# attention# batching# course
View on GitHub
⭐ Stars
938
🍴 Forks
66
🔥 Trending
+1 today
📜 License
Apache-2.0
Commercial use OK
📅 Created
2026
🔄 Last commit
18 days ago
🏷️ Category
ai
💻 Language
You maintain this project?

Claim its page: indexed whatever its rank, translated into six languages, and enriched with what you write yourself.

Claim this page →
tiny-vllm — GitHub preview card
📈 Star history
938937
2026-07-202026-07-21
📄 About

Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM

Frequently asked questions

What is tiny-vllm?

Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM

Is tiny-vllm open source?

tiny-vllm is an open-source project. It is released under the Apache-2.0 license.

Is tiny-vllm free?

Yes. tiny-vllm is free and open source — you can use, modify and self-host it.

🏅 Maintainer of this project?
OpenSourceAI badge — tiny-vllm

Add this live badge to your README — your GitHub stars and directory rank, refreshed daily.

[![OpenSourceAI](https://opensourceai.tech/badge.php?tool=jmaczan-tiny-vllm)](https://opensourceai.tech/project/jmaczan-tiny-vllm.html)
More badge options →
🧬 Related projects🧬 View the DNA map →