What’s new here
vllmops spawns vLLM processes directly, scrapes /metrics into an in-memory ring buffer, and renders the result in a Textual TUI. It requires no Docker, no Prometheus, and no Grafana. LiteLLM is installed as a separate tool to avoid version conflicts with vLLM itself.
What it does
vllmops is a control plane for running vLLM on bare metal. You declare each model in a YAML file (one file per model, one project config for profiles), then use the CLI or TUI to start, stop, restart, and inspect models individually or as a group.
The vllmops proxy subcommand generates a LiteLLM config from whatever models are running and keeps it in sync: every start or stop regenerates the config and respawns the gateway if it changed. The gateway exposes one OpenAI-compatible URL; clients pick the model by name. A respawn drops requests in flight and takes the gateway down for a second or two, and a model that dies on its own keeps its entry until you run vllmops proxy restart.
Per-project vLLM versions are pinned via uv. A vllmops doctor command checks Python, the venv, GPU visibility, and open ports before you try to start anything.
Who it’s for
Engineers running one or a few vLLM servers on a machine they own, who want reproducible config in Git and don’t want to maintain Docker or a Prometheus stack. It works on Linux and macOS with Python 3.10+.
Try it
Install with pipx or uv:
pipx install vllmopsuv tool install vllmopsThen initialize a workspace and start a model:
mkdir my-llms && cd my-llms
vllmops init
uv sync # creates .venv with vLLM installed
vllmops create-model # interactive: name, HF model, GPUs, port
vllmops start qwen3 # blocks on /health by default
vllmops tui # live metricsHow mature is it
88 stars, 12 forks, 5 releases (latest v0.5.0 on 2026-10-02) plus a rolling nightly pre-release, 19 commits in the last 90 days, 1 contributor. Licensed Apache-2.0. The README mentions 380+ tests and mypy strict type-checking. Created June 2026, so under six months old.