Where Does llama.cpp Store Models?
llama.cpp reads GGUF models from a models folder or a Hugging Face cache. Point the server at the file with the model flag.
Last updated
llama.cpp has no registry of models. Every run points at a GGUF file: a local path via the model flag, a URL, or a Hugging Face repo via the hf flag. The classic default path is a models folder beside the binaries.
Hub downloads cache separately. The cache root follows the LLAMA_CACHE variable, so shared machines can redirect it to big storage. Server deployments mount a models volume into the container and reference it by container path. Quantize choices trade RAM against quality per file.
Where llama.cpp stores this, by platform
/home/<username>/llama.cpp/models/model.gguf
Conventional local model path from the docs and examples. Any folder works when passed explicitly.
$LLAMA_CACHE/huggingface/hub
Hub download cache root. Set LLAMA_CACHE to move it off a small home partition.
C:\Tools\llama.cpp\models\model.gguf
Same layout on Windows builds. Server examples use models backslash versioned name under the build tree.
Frequently asked questions
How do I load a model in llama.cpp?
Pass a repo with the hf flag and it downloads plus caches automatically. For manual control, drop the gguf anywhere and pass its path with the model flag.
Where is the Hugging Face model cache?
Set LLAMA_CACHE in the environment. Cached hub downloads land there instead of the default spot. The server default models dir is just a convention.
Notice an outdated path? Let us know.