Deployment Bundle Generator
Open model → running endpoint, one script.
Paste any Hugging Face link — LaunchKit reads the real architecture off the Hub, recommends an engine + tensor-parallel config, checks the VRAM math, and generates a bundle you copy to any bare GPU box over SSH. The script preflights the machine (docker or a python venv with uv — it handles both), downloads weights, starts an OpenAI-compatible server, and smoke-tests it. Text and multimodal models supported.
⛨ 100% client-side — model lookups go straight from your browser to huggingface.co; your HF token is only ever asked for on your own machine, never here
01
Model
02
Hardware
TP ranks must sit on NVLink for decent inter-token latency — check
nvidia-smi topo -m shows NV# between them.03
Engine & runtime
auto = use docker if present, else install into a python venv with uv. TensorRT-LLM requires docker.
04
Serving shape
05
Advanced (optional — production hygiene)
Flags are mapped per engine (e.g. chunked prefill → vLLM
--max-num-batched-tokens, SGLang --chunked-prefill-size, TRT-LLM --max_num_tokens). Revision pinning and parsers apply to vLLM/SGLang; TRT-LLM ignores what it doesn't support.Feasibility — live VRAM math
checking…
weightsKV cacheoverhead