ChiChieh HuangFOUNDER · AI ENGINEER

Run Llama 2 on Your Phone: A llama.cpp Tutorial

Date2024.01.15
Length808 words
Reading~4 min
Figures8 figures
ChiChieh HuangFounder · AI engineer

Translated from the Chinese original · Read the original

LocatorLLM Plateau
2 wks
Overview

Run Llama on your laptop! A llama.cpp tutorial

llama.cpp has changed a lot since this was written. If you’re interested, see the new article I wrote on 2025/2/17.

With AI developing so quickly, large language models (LLMs) like Llama 2 and 3 have become a hot spot at the frontier of technology. But the smallest Llama 2 model is 7B (Llama 3’s smallest is 8B) and needs about 14 GB of memory, which ordinary consumer graphics cards can’t handle. So there’s a lot of work on reducing their resource use, like llama.cpp, which claims to run inference on a Raspberry Pi, and pyllama. After processing, inference can run on as little as 4 GB of RAM.

This article shows you how to quantize a model with llama.cpp and set up a server, so you can use your own LLM.

Step 1. Clone the llama.cpp project

Clone the llama.cpp repo

git clone https://github.com/ggerganov/llama.cpp.git

Step 2. Compile with the Makefile

Compile the llama.cpp project with make. This generates a lot of files, such as:

  • ./llama-cli: for command-line use
  • ./llama-quantize: for quantizing binary files
  • ./llama-server: starts a server

The approach above only runs on the CPU. If you want to use a GPU:

  • Windows/Linux users: compiling with BLAS (or cuBLAS if you have a GPU) is recommended and speeds up prompt processing
make GGML_CUDA=1 # LLAMA_CUBLAS is deprecated
  • macOS users: nothing extra needed. llama.cpp is already optimized for ARM NEON, and BLAS is enabled automatically. On M-series chips, using Metal to enable GPU inference is recommended and significantly increases speed.
LLAMA_METAL=1 make

— — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — —

If you get the following error during make:

: not foundld-info.sh: 2:
: not foundld-info.sh: 4:
: not foundld-info.sh: 9:
scripts/build-info.sh: 31: Syntax error: end of file unexpected (expecting "then")
make: *** [Makefile:676: common/build-info.cpp] Error 2

it’s caused by different line endings. If you clone with Git on Windows (CRLF) and then upload the folder to a Linux system (LF), you get this error. You can fix it like this:

sudo apt install dos2unix
find -type f -print0 | xargs -0 dos2unix

Step 4. Prepare the environment

Managing it with conda is recommended

conda create --name llamaCpp python=3.9
conda activate llamaCpp
python3 -m pip install -r requirements.txt

Step 5. Quantize the model

Use convert.py to quantize the model. Here’s how to quantize Llama 2 chat 7B:

a). Download Meta’s Llama 2 7B chat

See my earlier article for this; you can also download Llama 3 depending on your needs. If you followed my earlier steps, you’ll have a cloned llama project with llama-2-7b-chat/ downloaded inside it. Here are the files in that folder:

b). Convert to GGUF with convert-hf-to-gguf.py

The model file format llama.cpp can run is called GGUF. You can use the convert-hf-to-gguf.py tool that comes with llama.cpp to convert the llama-2-7b-chat model to GGUF.

convert-hf-to-gguf.py was installed in the environment step above, so you just need to run the command below.

Note: older versions of llama.cpp use convert.py

  • Be sure to use --vocab-dir to specify where tokenizer.model is; if you don’t, you’ll get the error FileNotFoundError: spm tokenizer.model not found.
  • Use the --outtype parameter to set the data type; it’s usually detected automatically, so you can leave it out. Note that this is not quantization!
  • Use --outfile to set where the converted model goes
cd llama.cpp
python3 convert-hf-to-gguf.py models/llama-2-7b-chat --vocab-dir ../llama --outtype fP16
 --outfile models/llama-2-7b-chat/llama-2-model.gguf

When I ran it, I got an error:

ValueError: The model's vocab size is set to -1 in params.json. Please update it manually. Maybe 32000?

To fix it, edit params.json in llama-2-7b-chat/ and change vocab_size from -1 to 32000.

{"dim": 4096, "multiple_of": 256, "n_heads": 32, "n_layers": 32, "norm_eps": 1e-06, "vocab_size": -1} # original
{"dim": 4096, "multiple_of": 256, "n_heads": 32, "n_layers": 32, "norm_eps": 1e-06, "vocab_size": 32000} # changed

Run it again:

python3 convert.py models/llama-2-7b-chat --vocab-dir ../llama --outtype fP16
 --outfile models/llama-2-7b-chat/llama-2-model.gguf

and llama-2-model.gguf will be generated at the --outfile location.

c). Quantize with llama-quantize

You can quantize the model with ./llama-quantize, at levels anywhere from 8-bit and 6-bit down to 2-bit. Usage:

Note: older versions of llama.cpp use quantize

./llama-quantize models/llama-2-7b-chat/llama-2-model.gguf models/llama-2-7b-chat/llama-2_q4.gguf Q4_K

The file size goes from the original 12.5 GB to 6.66 GB at Q8 and 3.79 GB at Q4. If you don’t want to quantize yourself, there are plenty of ready-made quantized models online that you can download and use directly.

Step 6. Load the model

Then you can use llama.cpp’s llama-cli for some basic usage:

Note: older versions of llama.cpp use main

./llama-cli -m models/llama-2-7b-chat/llama-2_q4.gguf --color -f prompts/alpaca.txt -ins -c 2048 --temp 0.2 -n 256 --repeat_penalty 1.3
./llama-cli -m models/llama-2-7b-chat/llama-2_q4.gguf -n 256 --repeat_penalty 1.0 --color -i -r "User:" -f prompts/chat-with-bob.txt
  • -c controls the context length; the bigger the value, the longer the conversation history it can refer to. Default 512
  • -f specifies a prompt template
  • -n controls the maximum length of the reply. Default 128
  • -b controls the batch size. Default 512
  • -t controls the number of threads. Default 8
  • --repeat_penalty controls how strongly repeated text is penalized
  • --temp is the temperature; the lower it is, the less random the reply, and vice versa
  • --top_p, top_k control decoding and sampling
  • -m specifies the model path
  • -ngl specifies how many layers to load on the GPU

Step 7. Server

llama.cpp can also set up a server so you can access the model over an HTTP API. Just run ./llama-server to get it up quickly. By default it opens an LLM service at http://127.0.0.1:8080/. The parameters are much the same as for .main/, and you can also adjust them once you’re in the server.

Note: older versions of llama.cpp use server

./llama-server -m models/llama-2-7b-chat/llama-2_q4.gguf -ngl 100 # GPU version

Support

If this article helped you, or you’d like to encourage me to keep writing, you can clap for it or buy me a coffee through the link below. Thank you for your support!

#llama-2#llm#llama-cpp#chatbots#tutorial

End of the trail

808 words, and you made it to the end.

Newsletter

Get the next essay by email.

One email when a new essay goes up, nothing else. Unsubscribe in one click.

ChiChieh Huang
Surveyor

ChiChieh Huang

I build generative AI products and write about them, first in Chinese. Lately I’ve been researching agent memory and testing the ideas in Cairn.