Install llama.cpp on Windows 11 and download to run the first local model

Table of Contents

    I want to try running a translation model locally on a Windows computer using llama.cpp, to translate some Chinese content into English. After working on it all morning, I finally managed to install it. I had used Ollama on Windows before; referring to the previous article, Ollama Installing the Google Gemma 3n Model. Overall, llama.cpp is a little more complicated than Ollama, but the difference isn’t significant.

    “Llama” means alpaca 🦙….

    Download llama.cpp

    Since I want to specify the installation directory, as there isn’t much space left on the system drive C of this old computer. Therefore, I didn’t use a PowerShell command for installation; instead, I directly downloaded the exe executable file. Download link:

    https://github.com/ggml-org/llama.cpp/releases

    I found the Windows x64 (CPU) version and downloaded it directly. The biggest surprise was that this zip package is only 19M in size, much smaller than the installation package for ollama. I remember ollama’s installation package being over 700M in size. After extracting it, I can use it right away. There are several exe files inside; I’ll explain their functions later.

    Then, I set the extraction directory of llama.cpp to the system’s PATH environment variables, so it can be called directly from the command line.

    Download the model

    Just downloading llama.cpp doesn’t do anything; a model is still missing. The official recommended command for installing a model is:

    # Download and run a model directly from Hugging Face
    llama cli -hf ggml-org/Qwen3.5-0.8B-GGUF
    
    # Launch OpenAI-compatible API server
    llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF
    

    However, it’s important to note that Hugging Face isn’t directly accessible in China. Therefore, you’ll need to find a mirror site.

    https://modelscope.cn

    Once you open the site, search for the model you want to install. For example, if you want to install Hy-MT2-1.8B, simply search for:

    https://modelscope.cn/models/Tencent-Hunyuan/Hy-MT2-1.8B-GGUF

    Click to download the model. Follow the prompts to install the Python dependencies first:

    pip install modelscope
    

    Then use the installed modelscope to download the model:

    modelscope download --model Tencent-Hunyuan/Hy-MT2-1.8B-GGUF README.md --local_dir ./dir
    

    Replace README.md with the file name of your .gguf file you want to download. This model weighs around 1GB.

    Run It

    Execute

    llama-server -m Hy-MT2-1.8B-Q6_K.gguf --port 8080
    

    You will see the following messages:

    0.00.003.588 I srv  llama_server: initializing ...
    0.00.022.599 I cmn  common_param: common_params_print_info: verbosity = 3 (adjust with the `-lv N` CLI arg)
    0.00.023.015 W srv  llama_server: security: no API key is set and CORS allows all origins (see https://github.com/ggml-org/llama.cpp/pull/25655)
    0.00.029.825 I srv    load_model: loading model 'Hy-MT2-1.8B-Q6_K.gguf'
    0.00.851.910 W load: special_eos_id is not in special_eog_ids - the tokenizer config may be incorrect
    0.04.102.178 I cmn          init: llama threadpool init, n_threads = 6
    0.38.440.620 I srv    load_model: initializing, n_slots = 4, n_ctx_slot = 216064, kv_unified = 'true'
    0.38.547.017 I srv  llama_server: model loaded
    0.38.547.623 I srv  llama_server: listening on http://127.0.0.1:8080
    

    When you see “listening on http://127.0.0.1:8080”, it means it’s running (wait a few seconds). Now you can use it in your browser.

    llama.cpp WEB UI

    The interface is similar to the DeepSeek web version. I’ll test how to call it with Python later, so you can automatically run some translation tasks locally.

    Continue reading

    Python calls the model run by llama.cpp