Deployment Bundle Generator

Open model → running endpoint, one script.

Paste any Hugging Face link — LaunchKit reads the real architecture off the Hub, recommends an engine + tensor-parallel config, checks the VRAM math, and generates a bundle you copy to any bare GPU box over SSH. The script preflights the machine (docker or a python venv with uv — it handles both), downloads weights, starts an OpenAI-compatible server, and smoke-tests it. Text and multimodal models supported.

⛨ 100% client-side — model lookups go straight from your browser to huggingface.co; your HF token is only ever asked for on your own machine, never here
01

Model

02

Hardware

TP ranks must sit on NVLink for decent inter-token latency — check nvidia-smi topo -m shows NV# between them.
03

Engine & runtime

auto = use docker if present, else install into a python venv with uv. TensorRT-LLM requires docker.
04

Serving shape

05

Advanced (optional — production hygiene)

Flags are mapped per engine (e.g. chunked prefill → vLLM --max-num-batched-tokens, SGLang --chunked-prefill-size, TRT-LLM --max_num_tokens). Revision pinning and parsers apply to vLLM/SGLang; TRT-LLM ignores what it doesn't support.
Feasibility — live VRAM math
checking…
weightsKV cacheoverhead

Your bundle