Fine-Tuning using LoRA QLoRA with Unsloth

Table of Contents

  1. Introduction: You Can Train Your Own Model
  2. Requirements: Hardware, Software, Knowledge, and Cost
  3. How Fine-Tuning Actually Works: LoRA, QLoRA, and What Unsloth Does
  4. Why Fine-Tune Instead of Using Frontier Models?
  5. Ten Practical Use Cases, With Examples
  6. Datasets: The Part That Actually Matters
  7. Reinforcement Learning vs Fine-Tuning: GRPO and Friends
  8. How Unsloth Compares to the Alternatives
  9. Deploying Your Trained Models Online
  10. Running Everything Locally: Colab, Ollama, and Apple MLX
  11. The Practical Checklist
  12. Appendix: Terminology

1. Introduction: You Can Train Your Own Model

For most of the last decade, “training your own AI model” was something that happened at companies with hundred-GPU clusters and research budgets. That era is over. In 2025–2026, a developer with a free Google Colab account, a laptop with a modest GPU (or no GPU at all beyond the cloud), and a couple of hours can take an open-weight language model like Qwen3 and specialize it: teach it a persona, a domain, an output format, or a skill. The tool that made this practical for hundreds of thousands of developers is Unsloth; an open-source fine-tuning library that makes the process roughly 2× faster and use about 70% less memory than the standard Hugging Face training stack, according to the project’s own benchmarks.

This article is about the whole journey: what fine-tuning is, what hardware you actually need, why you would do it instead of just calling GPT-class APIs, ten concrete use cases, how to think about datasets, when to reach for reinforcement learning (GRPO) instead of supervised fine-tuning, how Unsloth compares to competitors like Axolotl, LLaMA-Factory, torchtune, and raw TRL/PEFT, and finally how to put your trained model online; on managed paid platforms, on your own server with vLLM, or on your laptop with Ollama and Apple MLX.

What fine-tuning is, in one paragraph

A large language model is a giant stack of numbers (weights) learned during pre-training on trillions of tokens of text. Pre-training gives the model general language ability; it costs millions of dollars. Fine-tuning takes those pre-trained weights and continues training them, gently, on a much smaller dataset of examples specific to your task: thousands (not trillions) of examples, minutes or hours (not months) of compute, dollars (not millions) of cost. The result is a model that keeps everything it knew about language and the world, but is now reliably good at your thing; answering support tickets in your company’s voice, emitting valid JSON for your schema, tutoring physics, or behaving like a specific character.

Why “Unsloth” exists

Fine-tuning has always been possible with the standard Hugging Face stack (transformers + trl + peft). The problem was efficiency and friction: naive QLoRA training of a 7–8B model needed 16–24 GB of VRAM and ran slowly; tutorials broke with library versions; and free Colab’s 16 GB T4 GPU was a tight fit. Unsloth, built by the Unsloth AI team (Daniel and Michael Han), rewrites the hot paths of the training loop: hand-derived forward and backward passes in Triton (GPU kernels written at a lower level than PyTorch’s automatic differentiation), fused optimizer steps, fused RoPE (rotary position embeddings), fused cross-entropy loss, and careful memory management that offloads at the right moments. The outcome, per the project’s benchmarks: about 2× faster training, ~70% less VRAM, longer usable context lengths, and near-zero code changes; because Unsloth works as a drop-in layer over the same transformers/trl APIs you would use anyway.

In simple terms: Fine-tuning is “sending your smart intern to a one-day course on your company.” They arrive already fluent; the course just teaches them your procedures. Unsloth makes the course cheaper (less GPU memory), faster (custom math), and easier to enroll in (fewer setup problems).

How to read this guide

If you are new to fine-tuning, read linearly; each section assumes the vocabulary of the previous one. If you have a specific question, jump to it:
Section 2 for will my machine run this?
Section 4 for should I even do this?
Section 5 for what would I use it for?
Section 6 for how do I build the dataset?
Section 7 for what is GRPO and why does everyone mention it?
Section 8 for why Unsloth and not the others?
Section 9 for how do I ship the result?
Sections close with an “in simple terms” summary you can skim if the details get dense, plus reference links when a topic goes deeper than this article can. Every concept in Part 1 has a corresponding runnable demo in Part 2; when you see (Demo 0N) in the text, that’s your pointer to code you can execute today on a free Colab GPU. The Terminology Appendix at the end of this article collects every technical word, Python function, and parameter used here.

References: Unsloth docs, GitHub README (installation, benchmarks, feature list), Unsloth blog.


2. Requirements: Hardware, Software, Knowledge, and Cost

The official requirements page is at unsloth.ai/docs/get-started/fine-tuning-for-beginners/unsloth-requirements; read it for the live table, since exact numbers shift as new model families are supported. This section summarizes what matters and adds the practical context the docs assume you already know.

Hardware: what GPU do you need?

Unsloth supports NVIDIA GPUs with CUDA compute capability 7.0 or higher; that means V100, T4, RTX 20/30/40/50 series, A100, H100, and newer data-center cards.

It also runs on AMD and Intel GPUs, on Apple Silicon, and even CPU-only (slow, but it works; useful for debugging your training script before you pay for GPU time). NVIDIA remains the best-supported and fastest path.

The number that decides everything is VRAM (video RAM, the GPU’s own memory). Thanks to QLoRA (explained fully in the next section), the requirements are startlingly small. Approximate VRAM for QLoRA fine-tuning with Unsloth, from the official docs:

Model sizeQLoRA (4-bit) VRAMLoRA (16-bit) VRAMExample model
~1B~3–4 GB~8–10 GBQwen3-0.6B / 1.7B
~3B~3.5–6 GB~16–20 GBPhi-4-mini class
~7–8B~5–6 GB~24–32 GBLlama 3.1 8B, Qwen3-8B
~14B~8.5–12 GB~48–58 GBQwen3-14B
~32B~20–24 GB~128 GBQwen3-32B
~70B~41–48 GB~2× A100 80GBLlama 3.3 70B

Context for those numbers:

  • A free Google Colab T4 (16 GB) can QLoRA fine-tune up to ~14B parameters with Unsloth; including Qwen3-14B. Without Unsloth’s memory savings, that same T4 tops out around 7–8B and often out-of-memory’s below that. This single fact is why so many Unsloth tutorials run on free Colab.
  • A gaming laptop with an RTX 3060/4060 (8 GB) comfortably trains 7–8B models; 4 GB cards can still train 1–3B models like the Qwen3 0.6B/1.7B models this article’s demos use.
  • M-series Macs (16 GB+ unified memory) are supported through Unsloth’s Mac/MLX paths; more in Section 10.
  • VRAM scales with context length too. A 7B model at 2,048-token context might need 6 GB; the same model at 32k context needs substantially more (Unsloth’s docs on long-context training show techniques like RoPE scaling that keep this manageable; up to 3.5× longer context in the same memory on supported models).

One more hardware note: RAM and disk. Have at least 16 GB of system RAM and ~50 GB of free disk; models are several GB each, and you will accumulate more of them than you expect in your Hugging Face cache.

Software: Python, CUDA, and the one-line install

  • Python 3.10–3.13 (recent 3.x versions are supported; the docs’ quickstart is the source of truth).
  • Linux and Windows are first-class; macOS works with Apple Silicon (Metal/MLX); WSL2 is the smoother Windows path for CUDA work.
  • CUDA: you do not need to install the CUDA toolkit yourself for QLoRA; PyTorch ships with everything needed. You only need a working NVIDIA driver. (You’d install the full toolkit only for compiling things like vLLM or flash-attention from source, which Unsloth deliberately avoids requiring.)

The recommended install (from the docs), using the uv package manager for speed:

uv venv
source .venv/bin/activate
uv pip install unsloth --torch-backend=auto

The --torch-backend=auto flag pulls the correct CUDA-enabled PyTorch build for your machine; this sidesteps the single most common “unsloth is broken” complaint, which is a CPU-only PyTorch accidentally installed. The plain pip install unsloth also works.

On Colab, pip install unsloth at the top of the notebook is all you need.

When something misbehaves, the docs’ universal repair command is:

pip install --upgrade --force-reinstall --no-cache-dir unsloth unsloth_zoo

And the classic sanity check:

import torch
print(torch.cuda.get_available())  # or torch.cuda.is_available() on NVIDIA
import unsloth
print(unsloth.__version__)

If import unsloth succeeds and your GPU is visible, every tutorial in Part 2 will run.

Knowledge: what to know before you fine-tune

Hardware and software are the easy requirements to tick off. The third requirement is knowledge, and its bar is lower than most engineers fear; Unsloth’s notebooks handle the deep plumbing, but you cannot operate what you do not understand. Here is the honest inventory, in three tiers.

Tier 1: non-negotiable (you will hurt without these).

  • Working Python. Not expert-level: you can read a traceback, write a 50-line script, manage a virtual environment (uv venv, pip), and edit nested dictionaries. Every knob in a training script is a Python argument; every failure surfaces as a Python exception.
  • Command line and notebooks. Running scripts, installing packages, navigating folders, and feeling at home in a Jupyter or Colab cell. Fine-tuning day-to-day is “edit a cell, run it, read the log,” repeated.
  • Training vocabulary. Loss, step, epoch, batch size, learning rate, train/eval split; plus the ability to read a loss curve. You need the vocabulary, not the calculus behind it. If “the loss fell from 1.5 to 1.0 and then plateaued” forms a picture in your head, you are ready; Section 3’s anatomy of a training run supplies the rest.
  • Data handling. Opening JSON/JSONL/CSV files, inspecting rows (pandas or plain Python), spotting malformed entries. This previews Section 6’s thesis: the dataset is the product, and most of your engineering time on a fine-tuning project is data engineering.

Tier 2: concept-level (one good paragraph each is enough; Section 3 covers all of them).

  • How text becomes numbers. Tokens, tokenizers, and context length; why max_seq_length exists and why it scales memory.
  • Chat templates and special tokens. What <|im_start|> is, and why a train/serve template mismatch silently breaks exports; the field’s #1 failure mode is a literacy problem, not a math problem (Section 6; Demo 02).
  • The LoRA/QLoRA knobs. What r, lora_alpha, target_modules, and 4-bit quantization do at the “which dial does what” level; Section 3’s analogies are the required depth, and the low-rank mathematics behind them are optional reading.
  • The Hugging Face surface. What the Hub is (GitHub, but for models and datasets), and what transformers, peft, trl, and Unsloth’s FastModel each contribute; Section 3 defines all of them in one table.
  • Experiment hygiene. Fixed random seeds, checkpoints, a held-out eval set, and a note of what you changed between runs; without these you cannot tell learning from luck.

Tier 3: explicitly NOT required.

  • No CUDA, Triton, or GPU-kernel knowledge; Unsloth exists precisely so you never write GPU code.
  • No distributed-training experience; everything in this article is single-GPU.
  • No research-level math or paper reading; the LoRA and QLoRA papers are linked for the curious, not as a gate. Each in one line: LoRA trains small side-matrices instead of all weights; QLoRA does that on a 4-bit-compressed frozen base.
  • No prior model-training experience and no ML degree; first-time fine-tuners routinely ship usable models in their first week.

Cost: what this actually costs to try

  • $0; free Google Colab (T4, 16 GB) trains Qwen3 0.6B/1.7B/4B/8B and even 14B with QLoRA. Sessions are time-limited and can disconnect, so save checkpoints and push to the Hugging Face Hub early and often.
  • ~$10/month; Colab Pro removes most session limits and gives better GPUs; this is the realistic “hobbyist” tier.
  • $1–$5 per training run; renting an RTX 4090/A100 by the hour (RunPod, Vast.ai, Lambda, Modal) for a few hours. An SFT run on 5–10k examples with a small model takes well under an hour on modern hardware.
  • Deployment; from $0 (run GGUF locally in Ollama) to a few cents per hour for a managed endpoint; see Section 9.

In simple terms:
If you can open a free Colab notebook, you can fine-tune a model today. A 16 GB GPU handles models up to ~14B parameters with Unsloth’s QLoRA; your laptop’s gaming GPU handles 7–8B; the whole software setup is one pip install unsloth; and the knowledge bar is working Python, basic data wrangling, and the training vocabulary this article teaches as it goes. No GPU programming, no research math.

References: Unsloth requirements, installation quickstart, GitHub README; both linked from the docs’ Getting Started section; for knowledge gaps, the Hugging Face LLM course and 3Blue1Brown’s neural-network video series are the fastest fillers.


3. How Fine-Tuning Actually Works: LoRA, QLoRA, and What Unsloth Does

To make good decisions (which model, how much data, which knobs), you need a working mental model of what fine-tuning does mechanically. This section builds it in three steps: full fine-tuning → LoRA → QLoRA → where Unsloth’s speedup comes from.

Fine-tuning is hiring a trained professional, not raising a child. Picture a conservatory-trained pianist joining your jazz band. You do not teach them what a piano is, how scales work, or how to read music; that education took two decades and you cannot afford it. Instead you rehearse your setlist for a few weekends, and the same pianist now plays your songs the way your band plays them.

Fine-tuning is exactly that: it takes a model whose general abilities were bought with millions of dollars of pre-training (the conservatory years) and rehearses it briefly on your examples (the setlist) until its default behavior matches your needs. In technical terms: fine-tuning means continuing to adjust the model’s weights (the billions of numbers inside it that encode everything it knows) against your dataset, instead of building those weights from scratch.

LoRA is margin notes, not a rewritten book. The brute-force way to change a professional’s behavior is to rewire their muscle memory entirely; that is expensive, slow, and they may forget Chopin along the way. LoRA (Low-Rank Adaptation) takes a gentler route: it leaves every original weight frozen and untouched, and bolts a small set of new, adjustable numbers alongside chosen layers of the network.

The analogy: you own a 900-page professional cookbook you are not allowed to rewrite, so you add a small card of margin notes to each chapter (“more lemon, less salt”), and every dish leaving the kitchen now reflects your notes. The book (the base model) stays pristine; the cards (the LoRA adapte-r, usually a few dozen megabytes) are the only thing you trained. Want a different cuisine tomorrow? Swap one set of cards for another; the expensive book is reused, not repurchased.

QLoRA is the travel edition of the book. The trouble is counter space: the full-size cookbook occupies the entire kitchen. QLoRA asks a practical question: while we are rehearsing, why not keep the frozen base model in a compressed form so it fits on a much smaller counter, and write our margin notes in full quality beside it? That compression is quantization: storing each weight with fewer bits, the way a JPEG stores a photograph at a fraction of the raw file’s size with barely visible loss.

QLoRA stores the frozen base at 4 bits per weight instead of 16, a roughly fourfold shrink. The counter, in real terms, is VRAM: the GPU’s own memory, the scarcest resource in this entire field. A 16 GB laptop-class GPU cannot hold a 16 GB model at full precision, but the 4-bit “travel edition” fits with room to spare, and the original QLoRA paper’s finding is the reassuring part: fine-tuning quality barely moves.

The ten-word vocabulary those three pictures rely on (parameters, pre-training, fine-tuning, training, inference, token, VRAM, adapter, quantization, overfitting) used to live here as a table. It has moved to the Terminology Appendix at the end of this article and grown along the way: every machine-learning term. Sections link to the appendix instead of pausing to define words mid-sentence; treat it as the dictionary for the whole series.

Hold those three pictures in mind: a professional you rehearse, notes you add to a book you cannot rewrite, and a travel edition that fits the counter you actually have. Everything else in this section, including what Unsloth does, is a refinement of these three ideas, not a new concept.

Full fine-tuning

The obvious way to fine-tune is to continue training all the model’s weights on your dataset. For a 8B-parameter model in 16-bit floats, that means: 16 GB just to store the weights, another 16 GB for gradients (one float per weight), and ~64–130 GB more for the AdamW optimizer’s per-parameter statistics. Total: well over 100 GB of VRAM for an “8 GB” model. Fine; if you have an 8×A100 node. Not fine on a T4. (LLaMA-Factory’s hardware table gives the same picture from the other direction: full fine-tuning a 7B model wants ~80–120 GB.)

Full fine-tuning also destroys the original model’s general abilities more easily and produces multi-GB artifacts. For almost every reader of this article, the answer is LoRA.

LoRA: train 1% of the weights

LoRA (Low-Rank Adaptation) rests on an empirical finding from the original 2021 paper: when you fine-tune a model for a specific task, the change in the weights has low “intrinsic rank”; it can be approximated by a much smaller matrix than the full weight matrix it modifies.

LoRA exploits that: freeze the original weights entirely, and bolt small trainable “adapter” matrices alongside them.

Concretely, inside a transformer, attention and feed-forward layers multiply activations by big weight matrices (for example a 2048×2048 projection; 4M parameters). LoRA adds, in parallel, two small matrices: A (2048×16) and B (16×2048); just 65k parameters, about 1.6% of the original. The layer’s output becomes original_output + B(A(x)).

During training, only A and B update. At inference you can either keep the adapters separate (tiny files, swappable per task) or merge them into the base weights so the model runs at full speed with zero overhead.

The knobs you’ll see in every Unsloth script:

  • r (rank); the “16” above. Bigger = more capacity, more memory, more overfitting risk. r=8 to r=16 is the sweet spot for most tasks; r=16 or 32 for style/persona work; higher mainly helps when datasets are large.
  • lora_alpha; a scaling factor for the adapter outputs, effectively “how loudly the adapter speaks.” Common heuristics: alpha = r or alpha = 2r. When tutorials set r=16, lora_alpha=16, that’s why.
  • target_modules; which layers get adapters. Modern practice (and Unsloth’s default for text models): all of q_proj, k_proj, v_proj, o_proj (attention) and gate_proj, up_proj, down_proj (the MLP blocks). Targeting only attention is the old way and underperforms.
  • lora_dropout; regularizer; Unsloth’s own notebooks run 0 because their faster kernels don’t benefit and dropout slows learning.

QLoRA: freeze the base model and reduce it to 4-bit

QLoRA (Dettmers et al., 2023) is LoRA plus one more trick: store the frozen base model in 4-bit precision (instead of 16-bit), with a technique called NF4 (NormalFloat4; a 4-bit format whose discrete levels match the bell-curve distribution of neural network weights) plus double quantization (quantize the quantization constants, saving a bit more memory).

The base model becomes ~4× smaller and is never updated during training; it’s read-only; while the LoRA adapters stay in 16-bit and train normally. Computation even happens in 16-bit (weights are dequantized on the fly), so quality holds up remarkably well: the original paper showed 4-bit QLoRA matches 16-bit LoRA fine-tuning quality on standard benchmarks.

The practical consequence is the VRAM table from Section 2: an 8B model that eats ~16 GB in 16-bit drops to ~5–6 GB in 4-bit, leaving room on a T4 for activations, optimizer states, and a decent batch size.

When to choose what:

MethodWhen to use
QLoRA (4-bit)Default for ≤24 GB GPUs; what all Part 2 demos use
LoRA (16-bit)You have 48 GB+ and want maximum quality; slightly better on very narrow domains
Full fine-tuningYou have a GPU cluster, a 100k+ dataset, and want to teach substantial new knowledge
Continued pre-trainingLarge raw-text corpora (domain absorption); same machinery, different data format; Unsloth supports it

What Unsloth actually does differently

Unsloth’s value proposition is not a new algorithm; QLoRA is QLoRA. It is an implementation upgrade of the whole training pipeline: the same mathematics, executed on the GPU with far less waste. There are six upgrades. Each is explained below in plain English, with its jargon defined as it appears.

All kernels rewritten in Triton, OpenAI’s programming language designed for GPU operations. A kernel is the small block of GPU code that performs one operation: a matrix multiply, a normalization step, a slice of the loss calculation. Stock PyTorch ships a generic implementation of every kernel, written to cope with any shape, size, and hardware combination a user might ever produce; that generality costs speed and memory on every single call.

Triton is a programming language created by OpenAI specifically for writing GPU operations: it reads like Python but compiles down to raw, optimized GPU code. Unsloth’s move rewrote all of the performance kernels in the training pipeline in Triton rather than reusing PyTorch’s generic ones and that’s what makes it fast. It results in about 2x faster training utilizing ~70% less VRAM.

Hand-written math for the backward pass. Every training step has two halves: the forward pass, where the model reads a batch of examples and produces predictions, and the backward pass, where arithmetic works out, weight by weight, how much each one contributed to the error; those per-weight signals are the gradients that tell the model how to adjust.

PyTorch normally computes backward passes with automatic differentiation, a general-purpose machine that can differentiate any calculation you can assemble; being universal is its strength and its tax, like a GPS that always finds a legal route but never knows the local shortcut. Unsloth’s team derived the backward math for the supported architectures (Llama/Qwen/Mistral/Gemma/Phi families) by hand and wrote it as Triton kernels.

The hand-derived route stops off in memory far less often; fewer round-trips per step means faster steps and lower peak memory, which is the whole game on a 16 GB card. Both of the first two points appear as simplified pseudo-code just below the list.

Fused operations: one pan instead of six. A GPU pays a small fixed toll every time it launches a kernel, and the standard stack launches a lot of them. Per training step, at least four operations each run as their own kernel, writing intermediate results to VRAM and reading them back: RoPE (rotary position embeddings; the mechanism that tells the model what order the tokens come in), the cross-entropy loss (the “how wrong was my prediction” score each step tries to reduce), RMSNorm (a numeric stabilizer applied between layers so numbers do not blow up), and the AdamW optimizer step (the algorithm that decides the size of each weight’s nudge, per the learning rate).

Unsloth fuses each into a single kernel, the way a cook who knows the recipe by heart uses one pan for the whole dish instead of washing a fresh pan for every ingredient. Fewer tolls paid and fewer intermediate numbers parked in VRAM; this is where a healthy slice of the “~70% less memory” claim comes from.

A “repair layer” over the ecosystem: same sockets, better wiring. Normally you wire LoRA training together from three Hugging Face libraries: transformers (models and tokenizers), peft (the LoRA adapters), and trl (the training loops). Unsloth does not ask you to learn new ones. When you run import unsloth, it patches those libraries’ objects at import time; think of an electrician who rewires the house behind the same wall sockets.

Everything still plugs in exactly where it did (SFTTrainer, GRPOTrainer, get_peft_model, …), but the current now flows through the fast Triton kernels from points 1 through 3. Your script reads like ordinary Hugging Face code; the physics underneath is different. (The table below defines each of those names in one line.) One honest consequence: because every model family needs its own hand-derived backward pass, Unsloth historically lags a few days behind brand-new architectures while the team does the math.

Ship-ready extras baked in. The steps where do-it-yourself tutorials usually break are at the edges of the workflow, and Unsloth ships those ready-made: chat-template handling (correctly rendering your data with each model family’s special tokens; a train/serve template mismatch is what makes exports babble, per Section 6), one-call GGUF export (turning the trained model into the compressed “travel edition” file format that Ollama and llama.cpp run; mechanics in Section 9), vision-language support (fine-tuning models that read images as well as text: Llama 3.2 Vision, Qwen2.5-VL, Gemma 3), and one-click Colab notebooks for ~40+ models in the notebooks repo and the docs notebook gallery, so a known-working reference for your exact model is one click away. A faster pen for reinforcement learning (vLLM-powered rollouts).

In GRPO training (Section 7’s main event), most wall-clock time is not spent learning; it is spent writing. Before any scoring happens, each prompt needs a whole group of candidate answers generated by the model itself; with the typical group size of 8, that is eight full completions per question.

Those generations are the rollouts, and producing them with plain model.generate is like hand-copying a newsletter eight times per subscriber. Unsloth can route generation through vLLM, an industrial serving engine (the same one from Section 9) that batches generation the way a printing press batches pages. Rollouts get roughly an order of magnitude faster (Unsloth reports large gains; the RL guide has current numbers), which is often the difference between a GRPO experiment that finishes overnight and one you abandon on Tuesday.

Below are simplified pseudo-code versions of each (heavily trimmed; real kernels also handle masking, strides, and dtypes), followed by an ASCII diagram of where they sit inside a training step.

Pseudo-code 1: the shape of a Triton kernel (point 1).
The example is a fused RMSNorm-style operation (RMSNorm, the numeric stabilizer, is defined in point 3): scale each row of numbers by its own average energy. The pattern to notice: exactly one read from VRAM at the top, exactly one write at the bottom, and everything in between happens in fast on-chip memory:

@triton.jit                          # "compile this function into raw GPU code"
def rmsnorm_kernel(x_ptr, out_ptr, N, BLOCK: tl.constexpr):
    row  = tl.program_id(0)          # the GPU runs THOUSANDS of copies of this
                                     # function at once, each copy owning one row
    cols = tl.arange(0, BLOCK)       # the slice of the row this copy handles
    x    = tl.load(x_ptr + row*N + cols)     # one read from (slow) VRAM
    rms  = tl.sqrt(tl.mean(x * x) + 1e-6)    # all the math happens in
    y    = x / rms                           # (fast) on-chip memory,
    tl.store(out_ptr + row*N + cols, y)      # then one write back to VRAM

The stock path would do the same job as a chain of separate kernels (square, then mean, then sqrt, then divide), each writing its intermediate result to VRAM for the next kernel to read back. The Triton version touches VRAM twice.

Pseudo-code 2: the hand-derived backward pass (point 2).
Toy operation: z = a*b + c. Watch where the three gradients come from, and count the kernel launches and VRAM trips each style needs:

# --- Stock PyTorch: automatic differentiation (the general-purpose GPS) ---
z    = a * b + c                # forward: silently RECORDS this op on a graph
loss = z.square().mean()
loss.backward()                 # replay the graph in reverse using generic
                                # kernels, one per recorded op, each paying a
                                # VRAM round-trip:
                                #   grad_z -> kernel -> grad_a   (VRAM trip)
                                #   grad_z -> kernel -> grad_b   (VRAM trip)
                                #   grad_z -> kernel -> grad_c   (VRAM trip)

# --- Unsloth: calculus done on paper, then ONE kernel per layer ---
@triton.jit
def addmul_backward(g_z, a, b, c, g_a, g_b, g_c):   # g_ means "gradient of"
    # derived by hand, once, by a human:
    #   dL/da = g_z * b
    #   dL/db = g_z * a
    #   dL/dc = g_z
    g_a = g_z * b      # all three gradients computed in the SAME pass,
    g_b = g_z * a      # in on-chip memory, from a single load of a, b, c:
    g_c = g_z          # no graph to record, no replay, no VRAM intermediates

Autodiff must stay general, so it builds and replays a graph; the hand-derived kernel collapses the whole chain into one launch.

Now scale the toy up: 2048×2048 matrices, dozens of layers, thousands of steps. The per-step savings compound into the ~2x speedup and ~70% VRAM reduction on the project’s benchmarks.

The same training step, drawn. Where the two execution paths sit in a full forward/backward cycle:

        ONE TRAINING STEP: the same mathematics, two execution paths

STOCK STACK (generic kernels + autodiff graph replay)
------------------------------------------------------
forward: x -> [k1] -> [k2] -> [k3] -> ... -> [kN] -> loss
| | | |
v v v v every arrow is a
VRAM VRAM VRAM VRAM round-trip through
(write) (read) (write) (write) the slow shared pool
backward: loss -> [kN'] -> ... -> [k2'] -> [k1'] -> nudge weights
(replays the recorded graph, one small kernel per op)

UNSLOTH (fused Triton forward + hand-derived backward)
------------------------------------------------------
forward: x -> [ FUSED K1 ] --> [ FUSED K2 ] -> loss
| |
v v intermediates live in fast
registers registers on-chip memory; VRAM is
touched rarely
backward: loss -> [ HAND-DERIVED GRAD K1 ] -> ... -> nudge LoRA A, B only
(one kernel produces a whole layer's gradients at once;
the frozen 4-bit base model is never updated)

Point 4 name-drops the libraries Unsloth sits on top of, so here they are in plain English.

NameWhat it actually is
transformersHugging Face’s foundational library; defines how models and tokenizers are downloaded, run, and trained. Everything else in the stack builds on it
peft“Parameter-Efficient Fine-Tuning”; Hugging Face’s adapter toolkit. LoRA is implemented here
trl“Transformer Reinforcement Learning”; Hugging Face’s library of ready-made training loops that sit on top of transformers and peft
get_peft_modelA peft function: takes a plain loaded model and returns the same model with LoRA adapters bolted on, ready to train
SFTTrainerA trl object: the supervised fine-tuning loop; give it a model and a dataset, call .train()
GRPOTrainerA trl object: the GRPO reinforcement-learning loop (Section 7); generates groups of attempts per prompt and reinforces the rewarded ones
FastModelUnsloth’s own front door; one object that wraps all of the above and routes them onto the fast kernels

In kitchen terms from earlier in this section: transformers built the kitchen, peft supplies the margin-note cards, trl runs the rehearsal schedule, and Unsloth’s FastModel is the sous-chef who makes all of them move faster.

The costs and caveats, stated honestly: Unsloth’s acceleration applies to the model families it supports. Some hyperparameters of the underlying stack are pinned to known-good versions (the FAQ advises specific settings when things break). And because it patches the stack, mixing Unsloth with heavy custom modifications of trl internals can be brittle.

For standard SFT/GRPO workloads; which is 95% of what readers of this article will ever do; the trade is overwhelmingly favorable.

The anatomy of a training run

Every SFT script you will ever write with Unsloth has the same five-act structure, and knowing it makes every tutorial legible:

Act 1: Load
FastModel.from_pretrained(...) fetches the base model (4-bit by default via load_in_4bit=True) plus its tokenizer. You set max_seq_length here; it caps how long each training example can be and directly scales memory.

Act 2: Attach adapters
FastModel.get_peft_model(...) wraps the frozen base with trainable LoRA matrices; the r, lora_alpha, and target_modules from above. Print trainable_parameters here: you should see something like “8.5M/1,700M parameters (0.5%) trainable”; that number is why your memory footprint is small.

Act 3: Format the data.
Apply the chat template to your dataset (Section 6) so every row becomes one text string in the model’s native format, with only the assistant’s tokens actually learned (Unsloth’s templates mask the prompt portion automatically; training on the questions themselves would waste capacity).

Act 4: Train
An SFTTrainer loop over steps: each step, the model sees a batch of examples, predicts token-by-token, measures cross-entropy loss on the assistant tokens, and nudges the LoRA weights (via the optimizer) to reduce it. The learning rate controls nudge size; 2e-4 is the QLoRA default, roughly 100× larger than full fine-tuning rates because the adapters are small and the base is frozen. Epochs = passes over the dataset; 1–3 is normal, more invites overfitting (the model memorizes your examples instead of the pattern; you’ll see loss keep falling while held-out outputs get worse).

Act 5: Export
Merge/save adapters, merged weights, or GGUF (Section 9).

Freeze the giant model, bolt on tiny trainable matrices, show it thousands of “here’s the prompt, here’s the good answer” pairs, and let gradient descent tune only the small matrices until the model’s default behavior matches your data.

References: LoRA paper (Hu et al., 2021), QLoRA paper (Dettmers et al., 2023), Unsloth GitHub; the README’s benchmark section details the kernel work; the docs’ fine-tuning guide covers the API.


4. Why Fine-Tune Instead of Using Frontier Models?

The obvious question, asked by every stakeholder: “GPT-5/Claude/Gemini are already smarter than anything I can train. Why bother?” It’s a fair question, and the honest answer is that most projects should not fine-tune; prompt engineering, few-shot examples, and RAG (retrieval-augmented generation) come first, always. Fine-tuning wins in specific, well-understood situations. Here they are, with the reasoning you can defend in a design review.

Reason 1: Cost at scale

Frontier APIs price per token, forever. A fine-tuned small model prices per GPU-hour (or $0 if self-hosted on existing hardware). Do the arithmetic on a real workload: a support classifier processing 10M tickets/month at 1,000 tokens each is 10B tokens/month. At even $0.15 per million input tokens, that’s $1,500/month, every month, forever. A 1.7B model fine-tuned to do the same classification runs on a single modest GPU $50–150/month reserved, or a fraction of a serverless GPU’s time), and does the task with lower latency. Fine-tuning converts a linear-forever cost into a flat one. The break-even is usually somewhere between “thousands of daily requests” and “tens of thousands”; past it, the small model is dramatically cheaper.

The counter-consideration: frontier models can be cheaper for spiky or rare workloads; no idle GPU to pay for. And the newest small API models (Gemini Flash class, GPT mini class) have compressed the gap. The math must be done per workload.

The counter-consideration: frontier models can be cheaper for spiky or rare workloads; no idle GPU to pay for. And the newest small API models (Gemini Flash class, GPT mini class) have compressed the gap. The math must be done per workload.

Reason 2: Latency and locality

Every API call is a network round trip to someone else’s datacenter, sharing capacity with the world. A 0.6B–1.7B model quantized to 4-bit GGUF runs on a laptop CPU at usable speed, or on a $5/month micro-VM, or at the network edge. First token in tens of milliseconds, no queueing, no rate limits. Use cases where this is decisive: autocomplete-style UX, on-device features (a desktop app that works offline), robots/drones/industrial equipment, voice interfaces where 300 ms feels broken, and high-frequency batch jobs. Also: the data never leaves the machine; relevant for legal, medical, and defense contexts even beyond formal compliance rules.

Reason 3: Consistency and control

Frontier models drift. Providers silently update models, deprecate versions, and change behavior mid-contract; a prompt that worked Tuesday fails the following month. A fine-tuned open-weight model is frozen physics: same weights in, same outputs out, forever. You control when you upgrade, what it says, what it refuses, and how it formats. For regulated products, reproducible behavior is not a nice-to-have; an audit that asks “exactly which model produced this output in March?” must have an answer.

Reason 4: Format and style lock-in

A large fraction of production “prompt engineering” is really a wrestling match to force a general model into a rigid output contract: exact JSON schema, a house tone, a fixed set of categories, refusal rules, a persona. Prompts doing this are long (tokens = money, again), brittle, and leak. Fine-tuning moves that contract into the weights: the model’s default behavior becomes “emit this schema in this voice,” the prompt shrinks to the actual user input, and reliability on the narrow task typically rises from “mostly” to “consistently.” This is the single most common commercial fine-tune: not making a model smarter; making it dependable on one job.

Reason 5: Distillation; capability you can’t prompt into small models

Few-shot prompting has a floor: a 1.7B model given five examples in context still lacks the raw capability for some tasks. Distillation breaks the floor: use a frontier model (or an ensemble) to generate a training dataset; e.g., 50k carefully reasoned solutions in your exact format; then fine-tune the small model on it. The small model learns the pattern permanently rather than renting it via context. This is how many production “small but specialized” models are built: the frontier model is the teacher consulted once at dataset-construction time; the student serves millions of requests cheaply.

Reason 6: Privacy and compliance

Some data cannot be sent to any external API, full stop: patient records, banking documents, source code under license, defense data, EU data-residency regimes. Open weights + open tooling means training and inference happen wherever your compliance officer says they must: your VPC, your country, an air-gapped machine. Frontier APIs increasingly offer “zero data retention” enterprise tiers, but the strongest guarantee is architectural: the weights are yours and the network path doesn’t exist.

Reason 7: Ownership and independence

API models can be deprecated, repriced, rate-limited, or policy-changed out from under your product. A fine-tuned open-weight model is an asset you keep: it can be re-deployed anywhere, quantized for new hardware, merged with other adapters, or handed to a different vendor’s serving stack. Startups have learned this lesson the hard way when a pricing change turned their unit economics upside down overnight.

Prompting vs RAG vs fine-tuning

Prompt engineeringRAGFine-tuning
What changesInstructions at ask-timeDocuments provided at ask-timeThe model’s weights
Great forPrototyping; broad tasks; anything occasionalFresh, changing, or private factsConsistent behavior, format, style, narrow skills
Cost to buildMinutesHours–daysDays (mostly data work)
Marginal cost per requestHighest (long prompts, big models)Medium (retrieval + generation)Lowest (small model, short prompt)
Fails whenContract is rigid or prompt gets longThe task needs behavior, not factsThe task needs knowledge, not behavior
Update latencyInstantInstant (swap the index)Re-train (minutes–hours)

The production pattern is usually a combination: a fine-tuned small model for the format/persona/guardrails, RAG for the facts, and a frontier API behind an escalation path for the hard tail.

The honest decision procedure

  1. Can plain prompting do it? Then don’t fine-tune.
  2. Is the problem knowledge retrieval? Then RAG (give the model the documents at ask-time) usually beats fine-tuning; fine-tuning is poor at storing facts (it’s great at behavior). See Section 6’s caveat.
  3. Is the problem format/style/classification/persona/skill-on-known-domain? → fine-tuning shines.
  4. Is it cost/latency/privacy at volume? → fine-tune a small model, likely distilling from a big one.
  5. Is it math/reasoning/verifiable tasks where you want above-baseline reliability? → fine-tune first, then RL (Section 7).

In simple terms: Frontier models are brilliant generalists rented by the word. Fine-tuning builds you a specialist you own. Rent the genius while designing; deploy the specialist when the job is narrow, the volume is high, the format must be exact, or the data must stay home.

References: Unsloth FAQ (use-case guidance), Hugging Face’s PEFT/TRL docs for the underlying mechanics, and the Unsloth blog for real-world cost/latency case studies.



5. Ten Practical Use Cases With Examples

Below are the ten fine-tuning use cases that come up frequently in the Unsloth FAQ, in community threads, and in production systems. Each includes a concrete example you can imagine, pointers to free starting datasets, and a public example on Hugging Face you can open today: a shipped, documented model a business can learn from (and often deploy) directly. A common thread: fine-tuning changes behavior and format superbly, and absorbs domain style well; it is weak at cramming in raw facts; pair it with RAG when facts matter.

Use case 1: The custom assistant persona

You want a consistent character; a patient astrophysics tutor for a learning app, a noir detective for a game, a brand-voiced concierge; and system prompts drift, break jailbreaks, and get ignored under load.

The fine-tune: 1–5k multi-turn conversations demonstrating the persona (written by hand or synthesized by a frontier model and reviewed), fine-tuned with QLoRA on Qwen3-1.7B. Persona adherence becomes the model’s default rather than an instruction it must remember.

Unsloth’s own docs feature the “Making a Custom Chatbot with a Nickname” recipe; the model learns its name, tone, and boundaries from a tiny dataset, and the behavior survives hostile users far better than a system prompt does.

Example: dphn/dolphin-2.9-llama3-8b (Dolphin; the model card names the exact curated datasets used to give a base Llama 3 one consistent, steerable default persona) and HuggingFaceH4/zephyr-7b-beta (a small assistant tuned by the Hugging Face team, with the full recipe documented on the card). Both are evidence that a distinctive persona is a dataset problem, not a research problem.

Free starting datasets: yahma/alpaca-cleaned (52k instruction pairs), HuggingFaceH4/no_robots (10k high-quality human-written chats), OpenAssistant/Guanaco.

Use case 2: Domain specialization (legal, medical, finance, internal docs)

A general model knows a little law/medicine/your company’s product line and hallucinates confidently in all of them.

The fine-tune: Instruction pairs from your domain; bar-exam Q&A, clinical-note summarization, product-support CBTs; 1–10k examples minimum. The model absorbs vocabulary, reasoning patterns, and citation habits of the field. (For retrieving specific facts; statute numbers, this patient’s history, this quarter’s numbers; combine with RAG; the fine-tune teaches the domain’s style of answering, RAG supplies the facts.)

The classic community build is a medical QA assistant: fine-tune Qwen3-4B on MedQA-style question/answer/explanation triples; the model speaks “clinical” and follows the exam’s reasoning format. LLaMA-Factory’s showcase list includes several published medical fine-tunes (CareGPT, Chinese-LLaVA-Med) built exactly this way.

Example: epfl-llm/meditron-7b (EPFL’s medical model: Llama-2 continued-pre-trained on a curated corpus of PubMed abstracts, full articles, and clinical guidelines) and ProsusAI/finbert (a finance fine-tune that has run in production sentiment pipelines for years). Note what both cards publish: corpus composition and benchmark deltas. That is the level of documentation a business fine-tune should aim for.

Free starting datasets: openlifescienceai/medmcqa, GBaker/MedQA-USMLE-4-options, lighteval/legal_summarization.

Use case 3: Classification and structured extraction

Route support tickets, tag product reviews, extract invoice fields, classify transactions. You need a label, reliably, millions of times, cheaply.

The fine-tune: This is fine-tuning’s home turf; the task is narrow, the gold standard is cheap to label (or already labeled in your ticketing system’s history), and the required “reasoning” is modest. Train on prompt→label pairs; a 0.6B model often beats a prompted frontier model on this task because it does nothing else. Output as strict JSON for free plumbing.

“Classify this email: billing | bug | feature | sales | other → {"category":"bug","severity":2}” trained on 2k historical tickets; served as a 4-bit GGUF at ~zero marginal cost per message.

Example: cardiffnlp/twitter-roberta-base-sentiment-latest (social-post sentiment triage; millions of downloads) and SamLowe/roberta-base-go_emotions (a 28-emotion multi-label classifier with ~2M downloads, embedded in real products from forum tooling to monitoring dashboards). Classifier fine-tunes are the single most-deployed category of NLP model on the Hub, which tells you how business-proven the pattern is.

Free starting datasets: SetFit/tweet_sentiment_extraction, zeroshot/twitter-financial-news-topic, ibm/claim-review (or export your own history; usually better).

Use case 4: Text summarization and rewriting in a house style

Summaries that are fine but wrong-shaped: executive briefs must hit three bullets and a risk line; medical notes must follow SOAP; news must match the outlet’s voice.

The fine-tune: 1–5k document→summary pairs in exactly your shape. The model’s default output becomes your template. Works especially well for long documents with LoRA on a long-context base (Qwen3 supports long sequences well; Unsloth’s long-context tricks keep memory sane).

A legal-tech fine-tune turns 40-page contracts into a fixed eight-field risk memo; precision matters more than eloquence, so a distilled small model beats a generic one prompted to “summarize.”

Example: facebook/bart-large-cnn (the canonical news summarizer, fine-tuned on CNN/DailyMail and the default starting point in thousands of summarization projects) and its smaller, faster sibling sshleifer/distilbart-cnn-12-6. The pattern to copy: start from a public summarizer, then fine-tune the last mile into your house template.

Free starting datasets: EdisonLeeeee/Summarization, ccdv/cnn_dailymail (news), big_patent, or wikisum.

Use case 5: Code generation for your stack (and your DSLs)

Frontier models write plausible code; in last decade’s framework versions, or in generic SQL, not your internal DataQuery DSL.

The fine-tune: Instruction pairs of (natural language → code in your exact idiom), pulled from your real code review history or synthesized. Small code-focused bases (Qwen3-Coder, DeepSeek-Coder, StarCoder2) respond dramatically to even 1–3k curated examples of your patterns.

Text-to-SQL on your schema: fine-tune Qwen3-Coder-1.7B on question→SQL pairs over your tables; the model memorizes your column naming conventions and join graph quirks that a generic model guesses wrong at.

Example: defog/sqlcoder-7b-2 and defog/llama-3-sqlcoder-8b (Defog’s text-to-SQL fine-tunes, built as actual products for enterprise analytics over Postgres, Snowflake, and Redshift schemas; the cards publish the eval harness and benchmark comparisons against GPT-class models), plus Qwen/Qwen2.5-Coder-7B-Instruct as a strong open coding base to fine-tune further on your own DSL.

Free starting datasets: b-mc2/sql-create-context, JTiger/StackOverflow-question-python (or HuggingFaceH4/CodeAlpaca_20K for general instruction-code pairs).

Use case 6: Multilingual and low-resource language work

The model must operate in Swahili, Tagalog, Icelandic, or a dialect; where frontier coverage is shallow and open models are weaker still.

The fine-tune: Continue pre-training or SFT on corpora in the target language; even modest data (100k+ tokens of text for CPT, a few thousand instruction pairs for SFT) measurably shifts fluency, tone, and cultural framing. This is a flagship Unsloth community use case; multiple African- and Asian-language fine-tunes started as Unsloth notebooks.

Fine-tune Qwen3-1.7B on translated instruction data for a local-language ed-tech product; serve it in-region on one GPU instead of shipping every student’s words overseas.

Example: CohereForAI/aya-101 (instruction-tuned across 101 languages, roughly half of them low-resource such as Wolof, Shona, and Hausa; the card documents the human-annotated data program behind it) and the Helsinki-NLP organization, whose hundreds of per-language-pair translation fine-tunes show the “one model per locale” shipping pattern an in-region product would follow.

Free starting datasets: Helsinki-NLP datasets, allenai/nllb-adjacent corpora, mc4 slices per language, and the unsloth notebooks repo’s language fine-tuning examples.

Use case 7: Reliable tool calling and routing

Your agent’s planner must emit exactly one of 14 tool calls with exactly valid arguments, or the whole pipeline dies. General models get this 90–98% right; production wants 99.9%.

The fine-tune: Thousands of (situation → correct tool call JSON) pairs, including tricky negatives (“none of the tools apply → say so”). The small tuned model becomes a fast, deterministic-feeling router; frequently paired with a big model for the doing, small model for the routing.

A smart-home skill router fine-tuned on 3k utterance→intent pairs runs on a Raspberry Pi-class device (4-bit 0.6B model) with sub-100 ms responses and near-zero invalid calls.

Example: NousResearch/Hermes-2-Pro-Llama-3-8B (its function-calling format is used in production agent stacks, and the card documents the training data behind it) and gorilla-llm/gorilla-openfunctions-v2 (UC Berkeley’s model dedicated to emitting valid API calls). Both are templates for a business router: small, single-purpose, and validated against an executable benchmark.

Free starting datasets: glaiveai/glaive-function-calling-v2, Salesforce/xlam-function-calling-60k.

Use case 8: Guardrails and safety specialization

Your public-facing bot must never discuss off-topic subjects, must refuse a specific class of requests, and must do so in character, without lecturing; and every prompt-layered jailbreak that trends on social media becomes your incident report.

The fine-tune: Mix “refuse gracefully” and “stay on task” examples into your SFT set, including adversarial phrasings. Weights-level behavior is far more robust than instruction-level behavior, because there is no instruction to overwhelm. Teams routinely ship a “narrow, safe, cheap” tuned model as the public front door, with a frontier model behind it for allowed deep tasks.

A children’s-learning assistant that answers only math/science questions in age-appropriate tone, trained on 800 on-topic and 800 refusal dialogues; off-topic requests get a friendly redirect baked into the weights.

Example: meta-llama/Llama-Guard-3-8B (Meta’s industry-standard safety classifier that labels prompts and responses against a configurable policy; gated but free to request access). It demonstrates the professional pattern: safety behavior shipped as a versioned, auditable model, not as a paragraph of prompt text.

Free starting datasets: Anthropic/hh-rlhf (helpful/harmless pairs; also the base for DPO), OpenAssistant conversations.

Use case 9: Math and reasoning with reinforcement learning (GRPO)

You need the model to actually get arithmetic/word problems/logic right; and SFT teaches it to sound right, not to be right.

The fine-tune (RL): When answers are verifiable (math has ground truth; code compiles or doesn’t), GRPO training; generating many candidate answers per question and rewarding the correct ones; measurably improves accuracy beyond what SFT alone reaches. This is Section 7’s main event, and Unsloth has dedicated GRPO notebooks (GSM8K is the canonical starter: Qwen3-1.7B, ~300–500 steps, watch accuracy climb on the training chart).

Unsloth’s docs walk through GRPO on GSM8K with reward functions for correct answers, correct format, and integer-only answers; the exact recipe Part 2’s Method 4 demo miniaturizes.

Example: deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B (the public face of the R1 reasoning line, distilled into a model small enough for a laptop) and Hugging Face’s open-r1 project, an open reproduction of R1-style reinforcement learning with public datasets, training code, and logs; the closest thing to a published GRPO recipe book.

Free starting datasets: openai/gsm8k (grade-school math, 7.5k), meta-math/MetaMathQA, HuggingFaceH4/MATH-500 (eval).

Use case 10: Distillation; your own “mini-frontier” for one job

You want frontier-quality output on one task (say, generating product descriptions from specs) at small-model cost.

The fine-tune (2-stage): (1) Generate 20–100k task examples with a frontier model; including its reasoning where useful; and human-review a sample. (2) QLoRA your small model on them. The result: 80–95% of the teacher’s quality at 1–2% of the serving cost, with your format guaranteed by construction. This pattern; teacher generates, student learns; is arguably the industrial fine-tuning pattern of the decade, and it’s also how most synthetic-dataset vendors make their products.

Distill Qwen3-1.7B on 30k frontier-generated “spec → marketing copy” pairs; ship it as a 4-bit GGUF in your e-commerce plugin; per-description cost rounds to zero.

Example: teknium/OpenHermes-2.5-Mistral-7B (trained on ~1M largely GPT-4-generated examples; both the model and its source dataset teknium/OpenHermes-2.5 are public, making it a textbook teacher-to-student case study) and microsoft/Phi-3-mini-4k-instruct (a small model trained heavily on synthetic “textbook-quality” data that punches far above its size class; the technical report explains the data engine).

Free starting datasets: any high-quality instruct set (teknium/OpenHermes-2.5, HuggingFaceH4/ultrachat_200k, mlabonne/orca-agentinstruct-1M-v1 slice); and your frontier API account for the task-specific generation.

Fine-tune when the job is narrow and repetitive: one persona, one domain’s voice, one label scheme, one schema, one language, one router, one guardrail, one math skill, one distilled task. Don’t fine-tune to add knowledge; retrieve. Don’t fine-tune to add brilliance; rent it. Fine-tune to make behavior exact, cheap, and yours.

References: each example dataset lives on the Hugging Face Hub (search the names above); the “Public example” links in each use case point to open model cards and datasets on the Hub; the Unsloth notebooks gallery has ready-made Colab notebooks for persona work, medical/finance domains, multilingual training, and GRPO.


6. Datasets: The Part That Actually Matters

Ask anyone who ships fine-tuned models: the dataset is the product; the training code is plumbing. Unsloth (and its competitors) have made the plumbing nearly free; which moves 100% of your design effort to the data. Unsloth maintains an excellent datasets guide (see the docs’ dataset section); this is the working summary.

The five data formats

FormatShapeUsed for
Raw corpusjust textContinued pre-training (CPT); absorbing a domain’s prose
Instruction (Alpaca)instruction, input, outputSFT on task→answer pairs
Conversation (ShareGPT)conversations: [{from: human/gpt, value: ...}]SFT on multi-turn chats
ChatML / messagesmessages: [{role, content}]The modern native format most models expect
Preference (RLHF/DPO)prompt + chosen + rejected responsesDPO / preference tuning

Each format in plain English:

  • Raw corpus. Rows of plain text with no labels or roles; the model simply reads them, the way a student reads a library. Use it for continued pre-training, where the goal is absorbing a domain’s vocabulary and prose style rather than performing a specific task.
  • Instruction (Alpaca). One row per task: an instruction, an optional input the instruction applies to, and the correct output. It is the easiest format to author by hand or generate synthetically, and it covers most “do X to Y” single-shot tasks.
  • Conversation (ShareGPT). One row per dialogue: an ordered list of alternating human and assistant turns rather than a single question and answer. It is how you teach multi-turn rhythm; follow-up questions, corrections, and knowing when to stop talking.
  • ChatML / messages. The same dialogues expressed in the modern role-based structure (messages: [{role, content}]) that chat models and their tokenizers natively understand. Most current pipelines convert every other format into this one before applying the chat template.
  • Preference (RLHF/DPO). One row per prompt plus two completions: a chosen answer humans preferred and a rejected answer they did not. The model is not imitating either; it is learning to score the better one higher, which is the mechanism behind DPO training.

Unsloth’s unsloth.chat_templates module handles the fiddly part; turning your format into the exact special-token dance (<|im_start|>user\n...<|im_end|> for Qwen, etc.) that each model family expects:

from unsloth.chat_templates import get_chat_template
tokenizer = get_chat_template(tokenizer, chat_template="qwen3")   # or llama-3.1, gemma-3, ...

def formatting_prompts_func(examples):
    convos = examples["conversations"]
    texts = [tokenizer.apply_chat_template(c, tokenize=False, add_generation_prompt=False) for c in convos]
    return {"text": texts}

dataset = dataset.map(formatting_prompts_func, batched=True)

The single most common fine-tuning failure mode; “my model outputs gibberish after export”; is almost always a chat-template mismatch between training and inference. Same template both places, always.

How much data? (Less than you fear, more than a demo)

Unsloth’s guidance, echoed by practitioners: ~100 rows minimum to see any effect; 1,000+ rows for a usable model; 5–10k for something production-shaped. Quality beats quantity at every scale: 500 hand-reviewed examples routinely beat 20k auto-scraped ones. Multi-turn conversations help the model learn when to stop as much as what to say.

Where good datasets come from

Your own logs; support tickets, past edits, review-approved PRs. The gold standard: real examples of the exact behavior you want, in your exact context. A row from a cleaned ticketing export: 
{"instruction": "Draft a first reply to: 'I was double-charged for the March invoice.'", "output": "Hi Ana, I found the duplicate charge on invoice INV-2043 and have issued a refund; it will arrive within 3-5 business days."}
Public datasets; the Hub (links in Section 5) for standard tasks; great for learning the mechanics. A typical row from yahma/alpaca-cleaned:
{"instruction": "Rewrite the sentence using the passive voice.", "input": "The chef cooked the meal.", "output": "The meal was cooked by the chef."}
Synthetic generation; the modern default: a frontier model writes your training examples from a spec; you review samples and filter (the "constitutional" loop). A row generated from the spec "write polite refusal dialogues for out-of-scope questions":
 {"conversations": [ {"role": "user", "content": "Can you fix my internet connection?"}, {"role": "assistant", "content": "That one is outside my lane; I only handle billing. Want me to point you to tech support?"}]}

Tools exist to automate this, and the Unsloth datasets guide covers generating synthetic data, merging dataset columns (to_sharegpt with {column} and [[optional]] patterns), and conversation extension (splitting long threads into multiple training rows).

Distillation captures; log your production system’s (prompt, frontier-completion) pairs where the frontier answer was accepted; those pairs are a fine-tuning set. A captured pair from production logs:

{"messages": [ {"role": "user", "content": "Turn these release notes into three bullets."}, {"role": "assistant", "content": "- Adds SSO login\n- Fixes the sync crash on iOS\n- New dark theme"}], "accepted": true}

Dataset hygiene: the five checks before you train

Each check with a concrete example of what it catches:

  • Deduplicate near-identical rows (they over-weight themes and overfit). Example: "Summarize this contract" at row 120 and "Summarize this contract." at row 4,451, with identical outputs; a trailing period difference is not a difference.
  • Balance refusal/answer, short/long, categories; the model learns your proportions as priors. Example: a support set counting refund: 900 rows, upgrade: 60, cancellation: 40 trains a model that reflexively answers “refund”; resample until proportions match reality.
  • Decontaminate against your eval set (no train/test leakage; it will lie to you). Example: your eval question “If a train travels 60 km/h for 2.5 hours, how far does it go?” appears verbatim in train row 882; delete it from training or the eval score is fiction.
  • Sanity-read 50 random rows yourself, every time. Typos, truncations, and wrong answers in data become model behavior. Example: row 3,071’s output reads "Sure! Here is the summary: [TRUNCATED AT 512 TOKENS]"; until you catch it, that is the model’s idea of a summary.
  • Match inference-time format exactly; same chat template, same delimiters, same system prompt presence/absence. Example: trained on <|im_start|>user\n...<|im_end|> but served with User: ...\nAssistant:; that mismatch alone makes exports babble.

The model is a mirror of the dataset. 500 excellent examples of the behavior you want beat 50,000 mediocre ones, and the chat template you train with must be the one you serve with or the export will babble.

References: Unsloth datasets guide (formats, get_chat_template, to_sharegpt merging, synthetic data, vision formats), chat-templates docs for the template zoo.


7. Reinforcement Learning vs Fine-Tuning: GRPO and Friends

Supervised fine-tuning (SFT) teaches by demonstration: “here is a good answer; be like this.” Reinforcement learning (RL) teaches by consequence: “here is a question; try answers; good ones earn reward.” Modern open-model RL for LLMs centers on GRPO (Group Relative Policy Optimization); the algorithm DeepSeek popularized for R1-style reasoning training; and Unsloth ships one of the most accessible GRPO implementations, runnable from ~5 GB VRAM on ≤1.5B models. The official Unsloth RL guide is the deep reference; here is the working mental model.

Why SFT hits a ceiling

SFT is imitation. It pushes probability mass onto tokens in your demonstrations; and if your demonstrations contain the correct answer plus a flawed explanation, SFT dutifully learns both. It cannot exceed the data, cannot prefer one of its own outputs over another, and cannot practice. Three failure modes follow:

  • Sounding right ≠ being right. On math, logic, and code, fluent wrong answers train as happily as right ones.
  • Passive learning. The model never tries and fails on its own, so it never learns which of its own habits pay off.
  • Catastrophic imitation of noise. Every typo, hedge, and mediocre example in the dataset is curriculum.

How GRPO works, without the equations

For each training prompt (say, a math word problem), GRPO:

  1. Samples a group of candidate completions from the current policy; e.g., 8 different attempts at the problem (this is where Unsloth’s vLLM-accelerated generation pays off; rollouts dominate RL training time).
  2. Scores each with a reward function you write in plain Python. For GSM8K: does the extracted final number match the gold answer (correctness)? Is the answer an integer? Is the format right (#### 42 or the required JSON)? Rewards can be exact-match, format-regex, or continuous.
  3. Normalizes within the group (subtract the group mean, divide by its std); so each prompt’s relative winners are reinforced. This is the “Group Relative” trick, and it’s why GRPO needs no separate value/critic model; roughly halving memory versus PPO.
  4. Updates the policy with a clipped objective (the “PO” family trait, inherited from PPO) that limits how far any update can move the model in one step, plus a KL penalty tethering it to a reference model so it doesn’t forget its language.

Intuition: it’s practice with a coach. SFT is watching game film of champions; GRPO is taking shots and being told which went in.

RLVR: why verifiable rewards changed everything

The 2024–2025 breakthrough was RLVR; Reinforcement Learning with Verifiable Rewards: on tasks where correctness is mechanically checkable (math answers, unit tests passing, exact-match facts, JSON schema validity), the reward function is trivial, objective, and unlimited in supply; no reward model to train, no human preferences to elicit. That is why GRPO results on math/code are dramatic and on “vibes” tasks are modest: RL amplifies what you can measure. Choose tasks with checkable answers first.

When to use which

SituationTool
Teach style, persona, format, domain voiceSFT (LoRA/QLoRA)
Add knowledge/factsRAG (not training)
Raise accuracy on verifiable tasks (math, code, extraction)SFT first (bootstraps), then GRPO
Prefer among your own model’s outputs (no gold answer)DPO on preference pairs (cheaper than RL, weaker)
Teach a new skill from zeroSFT; RL only sharpens abilities the base model already has

The standard modern pipeline for a reasoning task: SFT → GRPO. SFT installs the format and basic competence; GRPO then buys real accuracy points by practicing against the verifier. Unsloth’s GRPO notebooks demonstrate exactly this on GSM8K, with guidance distilled in their RL docs: small models (≤1.5B) train with ~5 GB VRAM; expect to run 300+ steps (with 500+ prompt rows) before judging results; tune beta (KL strength) to 0 for pure task focus or higher to protect general ability; use mask_truncated_completions so over-long, truncated generations don’t pollute the loss; and lean on fast_inference (vLLM) unless memory forces otherwise.

Example code

from trl import GRPOConfig, GRPOTrainer

def correctness_reward(prompts, completions, answer, **kwargs):
    # +2.0 if the extracted final number equals the gold answer
    return [2.0 if extract_num(c) == int(a) else 0.0
            for c, a in zip(completions, answer)]

trainer = GRPOTrainer(
    model=model,                  # Unsloth FastModel + LoRA
    reward_funcs=[correctness_reward, format_reward, int_reward],
    args=GRPOConfig(
        max_completion_length=256, num_generations=8,   # the "group"
        learning_rate=5e-6, beta=0.0, temperature=1.0,
    ),
    train_dataset=gsm8k_dataset,
)
trainer.train()

Three reward functions scored per completion, group of 8 attempts per prompt, updates reinforce within-group winners. Part 2’s Method 4 runs a miniature of this end-to-end on Qwen3-0.6B.

The GRPO knobs that matter

KnobWhat it doesSensible start
num_generationsGroup size; attempts per prompt8 (lower if memory-tight)
max_prompt_length / max_completion_lengthTruncation bounds; long completions cost rollout time256 / 256 for GSM8K-class tasks
learning_ratePolicy update size; RL wants far smaller than SFT2e-6 to 5e-6
betaKL penalty to the reference model; 0 = pure task optimization, higher = protect general ability0.0 (raise if the model becomes repetitive)
temperatureRollout diversity; too low starves the group of variety, too high wastes it1.0
epsilonClipping range on policy updates0.2
mask_truncated_completionsExclude cut-off generations from the loss (they carry no signal about quality, only length)True
fast_inferencevLLM-backed generation for rolloutsTrue where VRAM allows

Two failure patterns to recognize early: rewards that look great on training prompts but collapse on held-out ones mean your reward function is gameable (usually rewarding format over correctness), and a model that drifts into repetition usually wants beta raised or the learning rate halved.

SFT says “copy this.” RL (GRPO) says “try, get graded, keep what worked.” Use SFT to teach manners and format; use GRPO; on tasks with machine-checkable answers; to actually raise the score.

References: Unsloth RL/GRPO guide (the definitive practical how-to: VRAM, hyperparameters, reward-function library), GRPO as introduced in DeepSeek-Math, the notebooks repo GRPO examples.


8. How Unsloth Compares to the Alternatives

Fine-tuning frameworks cluster into a handful of serious options. All of them wrap the same underlying research (LoRA, QLoRA, GRPO/DPO); they differ in speed, memory efficiency, ergonomics, breadth of models, and scale. Here is the field as of late 2026.

The contenders in one table

UnslothAxolotlLLaMA-FactorytorchtuneTRL + PEFT (raw HF)
Core identitySpeed/memory layer over HF stackConfig-file training harnessAll-in-one toolkit + Web UIPyTorch-native recipesThe foundation libraries themselves
Best atSingle-GPU efficiency; notebooks; Colab-first UXMulti-GPU/multi-node breadth; YAML reproducibilityBreadth of methods/models; GUI for non-codersHackability; research controlMaximum flexibility; it’s the substrate
Model coverage~40+ popular families, curated; day-quality supportVery broad (any HF model, mostly)100+ models incl. day-0 new releasesPopular Llama/Qwen/Mistral/Gemma familiesEverything HF supports (it is HF)
Training methodsSFT, CPT, DPO, ORPO, KTO, GRPO/RL, vision, text-to-speechSFT, DPO, PPO, pretraining, multimodalSFT, CPT, reward modeling, PPO, DPO, KTO, ORPO, SimPO + advanced optimizers (GaLore, BAdam, APOLLO, Muon…)SFT, DPO, PPO variants, full fine-tuneEverything, if you wire it
Multi-GPUSupported (newer), strongest single-GPUFirst-class (DeepSpeed/FSDP baked in)First-class (DeepSpeed, FSDP+QLoRA 70B on 2×24GB, Megatron backend)Native PyTorch distributedManual but unlimited
Config stylePython API (drop-in over trl)YAMLYAML + optional Web UI (llamafactory-cli webui)Python/yaml recipesPure Python
Speed on 1 GPU~2× faster, ~70% less VRAM (their benchmarks)Baseline (can integrate Unsloth)Baseline; integrates Unsloth (use_unsloth: true → ~170% speed in their tests)Near-baselineBaseline
Killer featureGGUF/Ollama export; free Colab notebooks for everythingBattle-tested configs; community-proven recipesLLaMA Board GUI; Ollama modelfile exportZero magic; you own every lineUniversality

Unsloth, honestly

Strengths. The efficiency claims are real and reproducible: on supported models, the same QLoRA job finishes in about half the wall-clock time and roughly a third less memory than vanilla HF; the difference between “fits on a free T4” and “doesn’t.” The notebook gallery (every major model, SFT/RL/vision/TTS, one click to Colab) is unmatched as a learning tool. Export tooling (save_pretrained_gguf, push_to_hub_gguf, merged-16bit, vLLM-ready safetensors) removes the most error-prone step of the workflow. And because it is the HF API underneath, everything you learn transfers.

Weaknesses. It accelerates supported architectures; a brand-new or exotic model may run (in degraded/baseline mode) or need a wait for kernels. Historically single-GPU-first: multi-GPU support exists now but DeepSpeed-scale distributed training is Axolotl/LLaMA-Factory territory for big jobs. It patches the stack; occasionally a specific trl/transformers version pairing is required (the FAQ documents the pins; the force-reinstall command fixes 90% of issues). And its opaqueness cuts both ways: when you want to modify the training loop at a low level, you may end up disabling the fast path.

Notably, the ecosystem treats Unsloth as a component, not a rival: LLaMA-Factory ships use_unsloth: true to run Unsloth’s kernels inside its own harness (their benchmarks: ~170% speedup on LoRA), and Unsloth itself builds on TRL/PEFT. Your choice is rarely exclusive.

Choosing

  • Learning, single GPU, Colab, fastest path from zero to fine-tuned model → Unsloth (this article’s demos use it).
  • Serious multi-GPU/multi-node training with reproducible YAML configs → Axolotl.
  • Team with non-coders, or you want a GUI, or the broadest method/model matrix → LLaMA-Factory (optionally with Unsloth enabled inside it).
  • Research control, custom algorithms, PyTorch maximalist → torchtune.
  • You already know the HF stack and need total flexibility → raw TRL/PEFT (accept baseline speed, or bolt Unsloth on for its kernels).

In simple terms: Unsloth is the turbocharger for one-GPU training; fastest to learn, fastest to run, best notebooks. Axolotl and LLaMA-Factory are the fleet-truck harnesses for bigger jobs and teams; torchtune and raw TRL/PEFT are for people who want to build their own engine. They’re not enemies: LLaMA-Factory literally embeds Unsloth as a speed option.

References: Unsloth GitHub (benchmarks, supported models), LLaMA-Factory GitHub (feature matrix, hardware tables, changelog with day-0 support entries), Axolotl docs, torchtune, TRL / PEFT.


9. Deploying Your Trained Models Online

Training produces an artifact; a LoRA adapter directory, a merged model, or a GGUF file; and deployment means putting that artifact where users can reach it. This section covers the menu: managed paid platforms (per the scope of this article), your own server with vLLM, and personal/local serving with Ollama and MLX. Unsloth’s docs cover export mechanics in depth (saving/loading models and GGUF export); here we pair mechanics with platform strategy.

Step zero: get the artifact right

Three artifact forms, three purposes:

# 1) LoRA adapters only (~50–500 MB); for serving stacks that compose adapters
model.save_pretrained("my_lora")

# 2) Merged 16-bit model; for vLLM/cloud platforms that want full weights
model.save_pretrained_merged("my_model", tokenizer, save_method="merged_16bit")
model.push_to_hub_merged("yourname/my-qwen3-1.7b-tuned", tokenizer,
                         save_method="merged_16bit", token="hf_...")

# 3) GGUF quantized; for Ollama/llama.cpp/local & CPU serving
model.save_pretrained_gguf("my_gguf", tokenizer, quantization_method="q4_k_m")
model.push_to_hub_gguf("yourname/my-qwen3-1.7b-gguf", tokenizer,
                       quantization_method="q4_k_m", token="hf_...")

GGUF quantization guidance from Unsloth’s docs: q4_k_m is the recommended balance (~4 bits/weight, minimal quality loss); q8_0 for near-16-bit quality at half size; f16 for exact; the q2_k/iq2_xxs rungs for extreme shrinking at real quality cost. Unsloth also publishes “Dynamic 2.0 GGUFs”; quantization that protects the most-sensitive layers; under names like UD-Q4_K_XL; look for the UD- prefix when you grab their pre-quantized models. If a model babbles after export, the docs’ troubleshooting list says: it’s the chat template or the eos token; train and serve with the identical template (this echoes Section 6).

The managed paid platforms

These run your model on infrastructure someone else operates, billed per hour or per use. Prices shift constantly; treat any number here as an order of magnitude and check current pricing.

Hugging Face; Inference Endpoints & Spaces. The natural first stop because your model is already on the Hub after push_to_hub_merged. Inference Endpoints are dedicated, autoscaled, API-compatible deployments: pick a GPU (small ones start around a few cents to ~$1/hour range), get an OpenAI-compatible URL in minutes, scale to zero when idle. Spaces host demos (Gradio/Streamlit UIs) with free CPU tiers and paid GPU tiers; ideal for showing stakeholders. Best for: quickest path from Hub repo to production URL; teams already in the HF ecosystem.

Together AI. Offers serverless fine-tuning and serving of open models, including uploading your own fine-tunes (LoRA or merged) to a dedicated endpoints system; strong OpenAI-compatible API, per-token billing, high throughput. Best for: product teams wanting per-token economics without GPU ops; also useful for distillation workloads (teacher + student under one API).

Replicate. Push a model (typically as a Cog container, or use their LLM tooling that accepts HF-mergeable weights/GGUFs); get an HTTPS API with per-second GPU billing and cold-start management. Best for: developers who want “a URL that runs my model” with minimal ceremony and are fine paying per use; excellent for demos and spiky workloads.

Modal (or RunPod Serverless, Lambda, Baseten; same category). You write a small Python function that loads your merged model into vLLM and Modal wraps it in an autoscaling, scale-to-zero container, billed per second. This is the most flexible managed pattern; you control the serving engine exactly (vLLM version, quantization, KV-cache config); at the cost of writing ~30 lines of infra code. Best for: engineers who want managed hardware but unmanaged software; the pattern also ports unchanged to your own k8s cluster later.

Rough monthly economics for a Qwen3-1.7B-tuned model at modest production traffic (think: a few hundred thousand requests/day, short prompts/outputs): any of these lands in the tens-of-dollars to low-hundreds band on small GPUs or serverless billing; versus API costs for a frontier model that can be 10–100× higher for the same tokens. The smaller the task, the bigger the multiplier.

Self-managed: vLLM on your own box

For real throughput, the open-source standard is vLLM; a serving engine with paged-attention KV caching and continuous batching that typically extracts several times the tokens/sec/GPU of naive transformers pipelines. Unsloth’s docs recommend exporting merged_16bit (or their pre-quantized dynamic quants) and serving with vLLM or SGLang; both speak an OpenAI-compatible API:

vllm serve yourname/my-qwen3-1.7b-tuned --max-model-len 8192
# then: POST /v1/chat/completions with the usual OpenAI-style payload

Any cloud VM (RunPod, Vast, Lambda, your own rack) becomes your inference server. You own uptime, autoscaling, and security; the trade is control for responsibility.

Personal deployment: Ollama and llama.cpp

For laptops, desktops, and homelab boxes, GGUF + Ollama is the endgame, and it is gloriously simple with Unsloth-pushed models (Ollama can pull HF GGUF repos directly):

# If you pushed with push_to_hub_gguf earlier:
ollama run hf.co/yourname/my-qwen3-1.7b-gguf:Q4_K_M

# Or any of Unsloth's pre-quantized models:
ollama run hf.co/unsloth/Qwen3-1.7B-GGUF:UD-Q4_K_XL

Your fine-tuned 1.7B at q4_k_m is a ~1.1 GB download that runs briskly on any modern laptop CPU and instantaneously on any recent GPU. Ollama also exposes a local OpenAI-compatible API (http://localhost:11434/v1), so your laptop becomes a zero-cost dev/staging endpoint. Unsloth’s Qwen3 guide notes one practical wrinkle: if a model loops or rambles in Ollama, raise the context window (e.g., num_ctx to ~32k); defaults are often too small for thinking-mode models.

A deployment decision tree

  1. Just me / my team / a demo? GGUF + Ollama on a laptop. $0. Done in five minutes.
  2. Public product, low/medium traffic, want zero ops? HF Inference Endpoint or Together on a small GPU; or Replicate per-second.
  3. Public product, real traffic, cost-sensitive? Modal/RunPod serverless with vLLM, scale-to-zero; or a reserved GPU box with vLLM once utilization justifies it.
  4. Compliance/air-gap? vLLM on hardware you control, weights shipped as safetensors.
  5. Whatever you choose: load-test with realistic prompts, watch tokens/sec and time-to-first-token, and keep the chat template byte-identical to training.

In simple terms: Export merged weights for cloud servers, GGUF for laptops. Hugging Face/Together/Replicate/Modal will host your fine-tune for roughly the price of a coffee per day at small scale, each trading control for convenience in a different ratio; and Ollama gives you a free private endpoint on every machine you own.

References: Unsloth GGUF export guide (quant table, troubleshooting), Unsloth Qwen3 guide (Ollama/llama.cpp commands), vLLM docs, platform pricing pages for HF Inference Endpoints, Together, Replicate, Modal.


10. Running Everything Locally: Colab, Ollama, and Apple MLX

The three environments this article’s demos target, and how they fit together.

Google Colab: the free training floor

Colab’s free tier provides a T4 with 16 GB VRAM; enough, with Unsloth QLoRA, for Qwen3-0.6B through Qwen3-14B. The workflow that works: pip install unsloth in the first cell, train with checkpoints pushed to the Hub, and export GGUF inside the same session (Colab’s CPU export can be slow but works; restart between heavy steps if memory fragments). Sessions disconnect after idle/time limits; treat everything on the VM as ephemeral and the Hub as home. Colab Pro (~$10/month) buys longer sessions and better GPUs; ZeroGPU slots give bursts of very fast hardware. All Part 2 demos are Colab-first.

Ollama: the free inference floor

Ollama (Section 9) is the simplest way to use models; base or fine-tuned, yours or Unsloth’s pre-quantized ones; on your own machines. It runs on macOS, Linux, and Windows, serves an OpenAI-compatible local API, and pulls directly from Hugging Face GGUF repos. Development loop that feels like magic: fine-tune on Colab → push_to_hub_gguf → ollama run hf.co/you/model:Q4_K_M on your laptop ninety seconds later.

Apple MLX: the Mac-native path

For Apple Silicon, the modern stack is MLX; Apple’s array framework tuned for unified memory; plus Unsloth’s Mac offerings. As of 2025–2026 Unsloth ships: Unsloth Desktop (a Mac app bundling fine-tuning with MLX and GGUF support, GPU-accelerated, with a UI), Unsloth Studio (installable on your own hardware, including Macs, via curl -fsSL https://unsloth.ai/install.sh | sh and launched with unsloth studio), and the pip unsloth package itself, which supports Apple Metal for many operations. The mlx-community org on the Hub hosts MLX-format builds of most popular models, and Unsloth’s Mac documentation covers the details (memory guidance: unified memory is shared with the OS, so leave headroom; a 16 GB Mac trains the 0.6B/1.7B sizes comfortably, 32 GB+ opens 4–8B). Part 2’s final method points Mac readers to the MLX route.

Why MLX matters on a Mac: NVIDIA-style CUDA does not exist on Apple Silicon, so the usual bitsandbytes/QLoRA path is unavailable; MLX reimplements the whole training stack (4-bit quantized LoRA included) natively for Apple’s unified-memory GPUs, and in practice feels comparable to a mid-range CUDA card for small models. The trade is ecosystem maturity; MLX covers fewer models and fewer cutting-edge techniques than the CUDA world; but for learning, prototyping, and small personal fine-tunes, a MacBook is now a legitimate training machine, not just an inference box.

In simple terms: Train in Colab (free GPUs). Serve with Ollama (free, everywhere). If you live on a Mac, Unsloth’s Desktop/MLX path trains right on the machine. The demos in Part 2 work through all three.

References: Unsloth Mac installation docs, MLX, mlx-community on HF, Ollama, Colab.


11. The Practical Checklist

A condensed, prioritized checklist for your first fine-tune; the same order Part 2’s demos follow.

  1. Decide if you should. Prompting first, RAG for knowledge, fine-tuning for behavior/format/cost/latency/privacy (Section 4’s decision procedure).
  2. Pick a small base model; Qwen3-0.6B/1.7B for learning and many production tasks; 4–8B when quality demands and hardware allows. Start from unsloth/ pre-quantized repos on the Hub.
  3. Get the environment; free Colab T4 + pip install unsloth; verify with import unsloth and a GPU check. (Demo 00.)
  4. Build the dataset before anything else; 100 rows minimum to feel it, 1,000+ to use it; your logs > public sets > synthetic; read 50 random rows yourself; dedupe, balance, decontaminate. (Demo 02.)
  5. Lock the chat template; get_chat_template(tokenizer, chat_template="qwen3"); the same template at train and serve time, or the export babbles.
  6. Train with QLoRA; load_in_4bit=True, r=16, alpha=16, all seven target modules, lr 2e-4, adamw_8bit, gradient checkpointing "unsloth". Short run first (60 steps), confirm loss falls and samples improve, then scale. (Demo 03.)
  7. Test before exporting; FastModel.for_inference(model), sample 10–20 task prompts, compare against the base model side by side.
  8. Export the right artifact; LoRA for adapter-aware servers; merged_16bit for vLLM/cloud; GGUF q4_k_m for Ollama/local. (Demo 05.)
  9. Deploy somewhere real; Ollama locally ($0) or a managed endpoint (HF/Together/Replicate/Modal) for a public URL. (Demos 06–07.)
  10. Only then consider GRPO; if the task is verifiable (math/code/format), RL after SFT buys real accuracy. Budget 300+ steps, write the reward function first, and measure on a held-out set. (Demo 04.)
  11. Keep an eval set forever. 50–200 fixed prompts with expected behaviors; run it after every change. Fine-tuning without evaluation is driving with your eyes closed.
  12. Save everything to the Hub; adapters, merged, GGUF, plus the dataset and training script. Your future self (and your team) inherit a reproducible artifact, not a mystery folder.

After the first fine-tune

The second fine-tune is where the leverage appears. By then you have an eval set you trust, a data pipeline you can rerun, and a feel for how loss curves map to output quality; the marginal cost of the next variant (a new persona, another domain, refreshed data) drops to an afternoon. Fine-tuning rewards iteration, and the entire point of the tooling surveyed in this article is to make that iteration cheap. Train small, measure honestly, keep what works, and discard the rest.


Appendix: Terminology

Every technical word used in this article and the Part 2 demos defined in plain English. Read this appendix as a dictionary: the body of the article links here instead of stopping to define words mid-sentence.

Technical wordPlain-English meaning
Parameters / weightsThe billions of numbers inside a model that encode its abilities; an “8B model” has eight billion of them
Pre-trainingThe enormously expensive first education (trillions of text fragments) that produces the base model
Fine-tuningContinuing to adjust the weights, gently and briefly, on your task’s examples
TrainingThe loop of showing examples and nudging weights until outputs match the data
InferenceRunning the finished model to generate text; no learning happens
TokenA word-piece; models read and write tokens, about three-quarters of an English word each
VRAMThe GPU’s own random access memory
AdapterThe small trainable add-on LoRA produces; a swappable file, not a whole model
QuantizationStoring weights with fewer bits (16 to 4) to shrink memory
OverfittingMemorizing your training examples instead of the underlying pattern

Training mechanics

Technical wordPlain-English meaning
LossThe “how wrong was the model” number every training step tries to shrink; lower is better
Cross-entropy lossThe specific loss for next-token prediction; high when the model assigned low probability to the token that actually came next
StepOne training update: show a batch, measure loss, nudge weights
EpochOne complete pass over the training dataset; 1-3 epochs is the normal fine-tuning range
Batch / batch sizeHow many examples the model sees per step; bigger learns smoother but eats more memory
Gradient accumulationFaking a bigger batch by summing gradients over several small batches before updating once (effective batch = batch size x accumulation steps)
Learning rateThe size of each weight nudge per step; 2e-4 for QLoRA, roughly 100x smaller for RL
GradientThe per-weight signal from the backward pass: “adjust me this much, in this direction”
Forward passThe model reading a batch of examples and producing predictions
Backward passWorking out, weight by weight, how much each contributed to the error; produces the gradients
Gradient descentThe overall algorithm: nudge weights slightly against their gradients, repeat until loss falls
OptimizerThe algorithm turning gradients into weight updates; AdamW is the standard everywhere
AdamW / adamw_8bitThe optimizer used in every script of this guide; the 8-bit variant stores its bookkeeping in 8-bit to save memory
Gradient checkpointingRecomputing intermediate results instead of storing them; trades roughly 30% speed for large memory savings ("unsloth" is the tuned setting)
CheckpointA saved snapshot of the model mid-training, so a run can resume or roll back
HyperparameterA setting you choose before training (learning rate, batch size, epochs), as opposed to values the model learns
Eval set / held-out setExamples the model is never trained on, used to test whether it learned the pattern or memorized
Loss curveThe plot of loss over steps; should fall, then flatten
PlateauThe flat stretch of a loss curve where further training stops helping
Random seedThe fixed number that makes “random” behavior repeatable between runs (3407 in the demos)
MaskingHiding part of a training example from the loss so the model learns only the assistant’s tokens, not the prompt

Models, data, and formats

Technical wordPlain-English meaning
TransformerThe architecture behind every model in this guide: a stack of identical layers mixing information across tokens
AttentionThe layer mechanism letting each token look at, and pull in information from, other tokens
MLP blockThe per-token “thinking” layer between attentions; where the gate/up/down adapters attach
Base modelThe pre-trained model you start from; frozen during LoRA training
Open weightsModel files published for anyone to download and run (Qwen3, Llama), as opposed to API-only models
Frontier modelThe biggest, smartest API models (GPT-class, Claude, Gemini); rented by the token
TokenizerThe program that chops text into tokens and converts them to and from the numbers the model reads
Context length / max_seq_lengthHow many tokens the model can process at once; doubles as the memory cap per training example
Chat templateThe exact text format (special tokens, role markers) a model expects conversations in; must match between training and serving
Special tokensMarker strings like `<
eos token“End of sequence”; the special token telling the model an answer is finished
System promptThe standing instruction that sets behavior for a whole conversation
Instruction format (Alpaca)Dataset rows shaped instruction, optional input, output
Conversation format (ShareGPT)Dataset rows of alternating human/assistant turns
ChatML / messagesThe modern [{role, content}] structure chat models natively understand
Preference format (RLHF/DPO)Dataset rows of a prompt plus a chosen and a rejected answer
Raw corpus / CPTPlain text without roles, used for continued pre-training (domain absorption)
DistillationUsing a big model to generate training data for a small one; teacher writes, student learns
RAGRetrieval-augmented generation: supplying documents to the model at ask-time instead of training facts into it
HallucinationConfident fabrication; the model producing fluent nonsense
BenchmarkA standard test suite used to compare models numerically
OpenAI-compatible APIA service speaking the same request format as OpenAI’s API, so the same client code works

LoRA, QLoRA, and quantization

Technical wordPlain-English meaning
LoRALow-Rank Adaptation: freeze the base, train small side-matrices (the adapters)
Intrinsic rankThe empirical finding that weight changes during fine-tuning live in far fewer dimensions than the full matrix; the fact LoRA exploits
r (rank)Adapter size dial; 8-16 is the sweet spot; bigger means more capacity and more overfitting risk
lora_alphaScaling factor for adapter outputs, “how loudly the adapter speaks”; usually set equal to r
lora_dropoutA regularizer randomly zeroing adapter values during training; Unsloth’s kernels run best with 0
target_modulesWhich layers get adapters; modern practice is all seven (q/k/v/o attention projections plus gate/up/down MLP projections)
MergeFolding trained adapters into the base weights so the result is one standalone model
QLoRALoRA with the frozen base stored in 4-bit; the small-GPU default
NF4NormalFloat4: the 4-bit number format QLoRA uses, shaped like the bell curve of weight values
Double quantizationCompressing the quantization constants themselves; a further memory saving
4-bit / 16-bitBits per stored weight; 16-bit is full precision, 4-bit is the “travel edition”
GGUFThe single-file quantized format that llama.cpp and Ollama run
q4_k_m / q8_0 / f16GGUF quality rungs: recommended balance, near-full quality at half size, exact
UD- dynamic quantsUnsloth’s pre-quantized GGUFs that protect the most sensitivity-critical layers
safetensorsThe standard safe, fast file format for model weights on the Hub

GPU and systems words

Technical wordPlain-English meaning
GPUThe processor that does math thousands of lanes at a time; where training and fast inference happen
CUDANVIDIA’s GPU software platform; PyTorch’s GPU math runs on it
Driver vs CUDA toolkitYou need only the system driver for QLoRA (PyTorch bundles the rest); the full toolkit is for compiling from source
Compute capabilityNVIDIA’s version number for GPU generations; Unsloth needs 7.0+ (V100/T4 and newer)
KernelOne small block of GPU code performing one operation: a matrix multiply, a normalization, a slice of the loss
TritonOpenAI’s Python-like language for writing custom GPU kernels; what Unsloth is written in
Fused operationSeveral kernels combined into one pass, skipping intermediate writes to VRAM
Automatic differentiationPyTorch’s general machinery that derives backward passes for any computation; universal but slower than hand-derived math
RoPERotary position embeddings; the mechanism telling the model what order tokens come in
RMSNormA numeric stabilizer between layers that keeps values from blowing up
vLLMIndustrial serving engine; batches many users’ requests for high throughput
KV cacheStored attention keys/values so generation doesn’t recompute the past every token
Paged attentionvLLM’s memory trick for managing the KV cache like an operating system manages memory pages
Continuous batchingServing technique feeding new requests into the GPU as others finish
Ollama / llama.cppLocal runtimes that serve GGUF models on your own machines
MLXApple’s array framework for Apple-Silicon GPUs; the Mac-native training path
Unified memoryMacs’ single RAM pool shared by CPU and GPU; budget with headroom for the OS
WSL2Windows’ built-in Linux environment; the smoother Windows path for CUDA work

Reinforcement learning words

Technical wordPlain-English meaning
RLReinforcement learning: learning by trying and being scored, rather than imitating
SFTSupervised fine-tuning: imitation learning from good examples
GRPOGroup Relative Policy Optimization: generate a group of attempts per prompt, reinforce the relative winners
PolicyRL-speak for “the model currently being trained”
RolloutOne candidate completion generated by the model during RL training
Reward functionPlain Python you write that scores each completion (e.g., +2.0 for a correct answer)
Reward shapingStacking small rewards (e.g., +0.5 for format) so the model gets signal before it can be fully correct
Reward modelA trained scorer for fuzzy tasks; unnecessary when rewards are verifiable
RLVRRL with Verifiable Rewards: tasks where correctness is mechanically checkable
KL penalty / betaTether keeping the RL model close to its reference so it doesn’t forget language; 0 disables
PPO / DPOOther RL-family algorithms: PPO is GRPO’s older, heavier parent; DPO learns from preference pairs without rollouts
num_generationsGRPO group size: attempts scored per prompt
TemperatureSampling randomness dial; higher means more varied (and riskier) outputs
top_p / top_kSampling filters restricting choice to the most plausible fraction (or count) of tokens
epsilonClipping range limiting how far one RL update may move the model
mask_truncated_completionsExcluding cut-off generations from the loss; they carry no signal about quality, only length
fast_inferenceUnsloth flag routing rollouts through vLLM
max_prompt_length / max_completion_lengthTruncation bounds for RL inputs and generated attempts

Python functions, objects, and parameters

The exact names that appear in this article’s code blocks and every Part 2 script, one line each.

Technical wordPlain-English meaning
FastModelUnsloth’s front-door object; wraps and accelerates everything below
FastModel.from_pretrained(...)Downloads and loads a base model plus tokenizer; parameters: model_name (which repo), max_seq_length (context cap), load_in_4bit=True (QLoRA mode)
FastModel.get_peft_model(...)Attaches trainable LoRA adapters to the frozen base; parameters: r, lora_alpha, lora_dropout, target_modules, use_gradient_checkpointing, random_state
get_peft_modelThe same attach-adapters operation from the underlying peft library
SFTTrainer / SFTConfigtrl‘s supervised fine-tuning loop, and the settings object it reads
GRPOTrainer / GRPOConfigtrl‘s GRPO loop, and its settings object
.train()The method that runs the loop until the configured steps or epochs finish
tokenizer.apply_chat_template(...)Renders a conversation into the model’s native text; parameters: tokenize (return text or numbers), add_generation_prompt (append the “assistant’s turn starts here” header), enable_thinking (Qwen3’s reasoning-mode switch)
get_chat_template(...)Unsloth helper that stamps a tokenizer with a model family’s template
dataset.map(fn, batched=True)Applies a function to every row (or batch of rows) of a dataset
formatting_prompts_funcThe conventional name for the map function that renders rows to training text
FastModel.for_inference(...)Switches the model into roughly 2x-faster inference mode
model.generate(...)Produces tokens; parameters include max_new_tokens, temperature, top_p, top_k
model.save_pretrained(...)Saves the LoRA adapters alone (tens of MB)
model.save_pretrained_merged(...)Merges and saves full weights; save_method="merged_16bit" is the vLLM/cloud format
model.save_pretrained_gguf(...)Exports llama.cpp/Ollama format; quantization_method="q4_k_m" is the default balance
model.push_to_hub_merged(...) / model.push_to_hub_gguf(...)The same two exports, uploaded straight to your Hub repo
prompts / completions / answerThe arguments your reward function receives: the questions, the model’s attempts, and any extra dataset column by name
**kwargsPython shorthand for “plus any other keyword arguments”; lets a function accept extras it ignores
trainable_parametersThe printed count of trainable versus frozen weights; should read well under 1% for LoRA

Loading