ChiChieh HuangFOUNDER · AI ENGINEER

Build Your Own LLM: A Llama 2 Tutorial

Date2024.01.14
Length381 words
Reading~2 min
Figures9 figures
ChiChieh HuangFounder · AI engineer

Translated from the Chinese original · Read the original

LocatorLLM Plateau
2 wks
Overview

In this article I’ll show how to set up and use Llama 2 and 3. The basic requirement is a GPU with 28 GB of VRAM, or at least 14 GB. If you don’t have that, see my llama.cpp tutorial.

Step 1. Request a license on Meta’s website

Go to Meta’s website and request to download the Llama models. You can request Llama 2, Llama Guard 3 and Code Llama at the same time. It usually takes 1–2 days, but in my recent experience I got approval within 10 minutes.

Step 2. Receive Meta’s commercial license

Once you’ve applied, you’ll receive an email with a commercial license.

Step 3. Download the model

I downloaded through GitHub. Meta also offers downloads through Hugging Face, where there’s an extra kind of model with -hf in the name, meaning it has been converted to a Hugging Face checkpoint.

If you’re also downloading through GitHub, follow the steps below:

a). Clone the project

First go to the Llama 2 GitHub or the Llama 3 GitHub and clone the project. The example below uses Llama 2.

b). Run download.sh

If you’re using WSL, the Windows Subsystem for Linux, you’ll hit the problem below. After some digging, it turned out to be caused by not using sudo.

c). Enter the URL from the email

Next, paste the URL you received by email (note that the URL in the example has expired).

Fig. 6The email Meta sends you

d). Enter the models you want to download

I chose 7B and 7B-chat. Once entered, the download starts automatically.

Step 4. Prepare the environment

Next, install the Python packages Llama 2 needs. Create an environment that can use CUDA. I used conda on WSL2:

conda create --name llama7B python=3.9b
conda activate llama7B
pip install -e .

Note: run the commands above inside the cloned folder.

Step 5. Quick test

Enter the following command to quickly test whether your Llama 2 works.

  • --nproc_per_node 1 is the model-parallel (MP) value, which differs by model: 7B-1, 13B-2, 70B-8
  • You can replace example_chat_completion.py with your own .py file
  • --ckpt_dir is the folder you downloaded the model into
torchrun --nproc_per_node 1 example_chat_completion.py \
    --ckpt_dir llama-2-7b-chat/ \
    --tokenizer_path tokenizer.model \
    --max_seq_len 512 --max_batch_size 6

Support

If this article helped you, or you’d like to encourage me to keep writing, you can clap for it or buy me a coffee through the link below. Thank you for your support!

#llm#llama-2#meta#installation#tutorial

End of the trail

381 words, and you made it to the end.

Newsletter

Get the next essay by email.

One email when a new essay goes up, nothing else. Unsubscribe in one click.

ChiChieh Huang
Surveyor

ChiChieh Huang

I build generative AI products and write about them, first in Chinese. Lately I’ve been researching agent memory and testing the ideas in Cairn.