Quantized LLM ZeroGPU Playground

Download and test one Hugging Face LLM at a time. Auto routes standard Transformers checkpoints to Transformers and GGUF files to the CUDA-enabled llama.cpp backend. GGUF repositories are inspected first so you can select only the Q4/Q5/Q8 (or newer) quant you want.

Backend
GGUF quant file (choose one)

Disk cache: no model selected.

Status: inspect a repository, then Download and Load it.

Detected format: not inspected yet
Backend: Auto

1 8192
0 2
0.05 1