Welcome to vllm-lite! In this tutorial we'll go from a fresh clone to a working build.
-
Rust 1.88+ (we use edition 2024 — matches
rust-versionin the rootCargo.tomland therust-toolchain.tomlpinned by the repo). Install via rustup:curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh
-
Git — for cloning and version control
-
~5 GB disk — for Rust toolchain, dependencies, and build artifacts
-
(Optional) CUDA 11.8+ — only needed for GPU inference (Qwen3, Llama with CUDA kernels). CPU-only inference works without it.
git clone https://github.com/pplmx/vllm-lite.git
cd vllm-lite
cargo build --workspaceThe first build downloads ~300 crates and takes 5-15 minutes. Subsequent builds are incremental (seconds).
# Run all unit + integration tests (skips #[ignore] slow tests)
just nextest
# Or with cargo directly
cargo test --workspace --no-fail-fastExpected: all tests pass, 0 failures. The exact count changes
every batch — run just nextest for the current number; don't
trust hard-coded numbers in older docs.
We use just (a Make alternative) for
common workflows:
cargo install just --lockedVerify: just --version → 1.x+
Before committing, run:
just fmt-check # cargo fmt --all --check
just clippy # cargo clippy with workspace lints
just doc-check # cargo doc --no-deps (0 warnings)All three must pass for CI to merge your PR.