How to Fine-Tune Llama 3 on Custom Data with Unsloth

Added:

Model Overview
Setup & Load
Data Prep
Fine-Tune
Inference & Save

Model Overview

0:02
Playing Section
  • 1

    Meta released Llama 3 with 8B and 70B variants.

  • 2

    Trained on 15T tokens with an 8K context length.

  • 3

    Outperforms Mistral and Gemma on benchmarks.

Basic understanding of Transformer-based Large Language Models (LLMs) and tokenization.
Familiarity with Python programming and navigating Jupyter notebooks, specifically Google Colab.
Conceptual knowledge of Supervised Fine-Tuning (SFT) and how it differs from pre-training.
Awareness of Parameter-Efficient Fine-Tuning (PEFT) methods, specifically Low-Rank Adaptation (LoRA) and QLoRA.
Advanced alignment methods such as Direct Preference Optimization (DPO) and Reinforcement Learning from Human Feedback (RLHF).
Model evaluation techniques, including the use of LLM-as-a-judge frameworks and standard benchmark evaluations.
Model quantization (converting weights to GGUF, AWQ, or EXL2 formats) and serving models efficiently using inference engines like vLLM or Ollama.
Advanced dataset curation strategies, including synthetic instruction generation and data cleaning pipelines.
30.6K views324likes8:31@fahdmirzaOriginal Release: 2024-04-18

This tutorial demonstrates how to fine-tune the Llama 3 model (8B or 70B parameter variants) on custom datasets using Google Colab, leveraging the unslot library for efficient 4-bit model loading and the Hugging Face TRL library for parameter-efficient fine-tuning with LoRA adapters that update only 1-10% of model parameters.