AI for the Curious/Labs/Fine-Tuning in Practice: Building the Fine-Tune

Fine-Tuning in Practice: Building the Fine-Tune

A practical walkthrough of the Python, hardware and decisions behind fine-tuning a small Gemma model to write unbearable LinkedIn posts.

LabsArticle 5 of 12 · ~16 min read

Current

Fine-Tuning in Practice: Building the Fine-Tune

In Part 1, I explained what fine-tuning is, where it fits alongside prompting and RAG, and why I trained a small model to write like the most insufferable person on LinkedIn. This part is about how I did it.

The whole training run lived in a Google Colab notebook. The finished workflow is fewer than 100 lines of Python, but the code conceals quite a few decisions such as which model to start with, how to fit it into GPU memory, which parts of it to train, how aggressively to train them, and what to save at the end.

I am going to walk through those decisions in the same order as the notebook. This is not a general-purpose fine-tuning framework, but a small deliberately silly experiment, and a way to see what the mechanics described in Part 1 look like as working code.

Note

The model ID used in the notebook is google/gemma-3-2B-it. The dataset contained 301 examples.

What I Built

The aim was simple: take an ordinary professional update and return a short LinkedIn post full of recognisable tropes such as humblebrags, dramatic revelations, effusive gratitude, life lessons, emojis and hashtags.

Each training example used a conversational structure with three messages:

{
  "messages": [
    {
      "role": "system",
      "content": "You are a LinkedIn post generator..."
    },
    {
      "role": "user",
      "content": "I passed my AWS certification."
    },
    {
      "role": "assistant",
      "content": "I am incredibly humbled to share..."
    }
  ]
}

The actual examples were stored as JSONL, which means one complete JSON object per line. This format is convenient because Hugging Face Datasets can load it directly, and each row already looks like the conversation format the instruction-tuned Gemma model expects.

The assistant responses were not all copies of one template. They varied in length, opening style and degree of absurdity. This is an important detail because a fine-tune learns repeated patterns in its examples, so a dataset made entirely from one rigid pattern is likely to teach the model to repeat that pattern rather than generalise the style.

The Hardware

I ran the notebook in Google Colab using an NVIDIA A100 GPU. The notebook was written for an A100 runtime and estimated a training time of roughly 20 to 30 minutes.

An A100 is powerful hardware, but the important point is that I did not use it to perform a full fine-tune. I used QLoRA, which loaded a quantised copy of the base model and trained a comparatively small set of adapter parameters. That dramatically reduced the amount of memory required.

The notebook checked the hardware rather than assuming the requested GPU had actually been assigned:

import torch

print(f"GPU: {torch.cuda.get_device_name(0)}")
print(
    "VRAM: "
    f"{torch.cuda.get_device_properties(0).total_memory / 1e9:.1f} GB"
)

This check is important because Colab runtimes can vary, and a notebook that fits comfortably on one GPU may run out of memory on another.

1. Installing the Training Libraries

The notebook used the Hugging Face training ecosystem:

!pip install -q -U \
    transformers \
    peft \
    trl \
    datasets \
    bitsandbytes \
    accelerate

Each library has a distinct job:

LibraryWhat it did
transformersLoaded the Gemma model and tokenizer.
peftAdded and managed the LoRA adapters. PEFT stands for Parameter-Efficient Fine-Tuning.
trlProvided the supervised fine-tuning trainer.
datasetsLoaded the training examples from JSONL.
bitsandbytesLoaded the base model using 4-bit quantisation.
accelerateHelped place and run the model on the available hardware.

The working notebook still contains some evidence of experimentation here, including different installation commands and an interrupted attempt to install transformers directly from GitHub. That is fairly normal while getting a notebook working, but a reusable version should pin known-compatible package versions. Otherwise, running the same notebook months later can produce a different dependency combination.

2. Authenticating and Loading the Dataset

I stored my Hugging Face token as a Colab secret named HF_TOKEN, then authenticated without putting the token directly in the notebook:

from google.colab import userdata
from huggingface_hub import login

hf_token = userdata.get("HF_TOKEN")
login(token=hf_token)

Keeping the token in Colab's secret store matters if a notebook is ever shared. A token pasted into a code cell can easily end up in notebook history or a public repository.

I then uploaded the dataset and loaded it with Hugging Face Datasets:

from google.colab import files
from datasets import load_dataset

uploaded = files.upload()

dataset = load_dataset(
    "json",
    data_files="full_dataset.jsonl",
    split="train",
)

print(f"Dataset: {len(dataset)} examples")
print("Sample input:", dataset[0]["messages"][1]["content"])
print("Sample output:", dataset[0]["messages"][2]["content"][:200])

The final three lines are a sanity check. Before spending time and GPU credit on training, I wanted to confirm that the expected number of rows had loaded, I was loading the right file, and that the user and assistant messages were in the correct positions.

3. Loading Gemma in 4-Bit

The base model was Google's instruction-tuned Gemma model:

MODEL_ID = "google/gemma-3-2B-it"

The it suffix means it is instruction-tuned. I was not starting from a raw base model and teaching it how conversations work. I was starting from a model that already understood system, user and assistant messages, then adapting its response style.

The next block created the quantisation configuration:

from transformers import (
    AutoModelForCausalLM,
    AutoTokenizer,
    BitsAndBytesConfig,
)

bnb_config = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_quant_type="nf4",
    bnb_4bit_compute_dtype=torch.bfloat16,
    bnb_4bit_use_double_quant=True,
)

There are four important choices here:

  • load_in_4bit=True stores the frozen base model in a much smaller 4-bit representation.
  • nf4 selects NormalFloat 4, a 4-bit format designed for normally distributed neural-network weights.
  • bfloat16 uses a wider numerical format for the calculations performed around those quantised weights.
  • Double quantisation also quantises some of the values used to describe the first quantisation, saving a little more memory.

This combination is what makes the training run QLoRA rather than ordinary LoRA. LoRA describes the small trainable adapters. The Q in QLoRA refers to the quantised base model beneath them.

The tokenizer and model were then loaded:

tokenizer = AutoTokenizer.from_pretrained(MODEL_ID)
tokenizer.pad_token = tokenizer.eos_token
tokenizer.padding_side = "right"

model = AutoModelForCausalLM.from_pretrained(
    MODEL_ID,
    quantization_config=bnb_config,
    device_map="auto",
    trust_remote_code=True,
)

model.config.use_cache = False

Setting the padding token avoids errors when examples in a batch have different lengths. Disabling the model cache is also important during training. The cache speeds up token-by-token generation, but my understanding is that it uses additional memory and can conflict with how gradients are calculated.

device_map="auto" lets the Transformers and Accelerate libraries place the model on the available device. In this notebook, that device was the A100.

4. Attaching the LoRA Adapters

Loading a model in 4-bit makes it smaller, but it does not yet make it trainable in the way we need. PEFT first prepares the quantised model for training:

from peft import (
    LoraConfig,
    get_peft_model,
    prepare_model_for_kbit_training,
)

model = prepare_model_for_kbit_training(model)

I also cleared any unused Python and CUDA memory before creating the adapters:

import gc

gc.collect()
torch.cuda.empty_cache()

The adapter configuration was:

lora_config = LoraConfig(
    r=16,
    lora_alpha=32,
    target_modules=["gate_proj", "up_proj", "down_proj"],
    lora_dropout=0.05,
    bias="none",
    task_type="CAUSAL_LM",
)

model = get_peft_model(model, lora_config)
model.print_trainable_parameters()

This is the part of the notebook that decides what the fine-tune can change.

Rank and scaling

r=16 sets the rank of the two smaller matrices described in Part 1. A higher rank gives the adapter more capacity to represent changes, but also adds trainable parameters and consumes more memory.

lora_alpha=32 controls the scale of the adapter's update. In this run it was set to twice the rank, a common starting point rather than a universal rule.

Target modules

I attached adapters to three projection modules:

  • gate_proj
  • up_proj
  • down_proj

These are in the model's feed-forward or MLP blocks. The adapters therefore learned adjustments to a selected part of the model rather than every weight matrix in the network. This is exolained in part 1.

This was a deliberately narrow configuration. Many LoRA recipes also target attention projections such as q_proj, k_proj, v_proj and o_proj. More targets create more capacity for tuning, but that does not automatically mean a better result. It also means more trainable parameters, more memory and another variable to evaluate.

Dropout and bias

lora_dropout=0.05 randomly disables a small proportion of adapter activations during training. This introduces regularisation, which can help reduce overfitting on a small dataset.

bias="none" leaves the model's bias terms untouched. Only the LoRA adapter weights are trained.

Finally, model.print_trainable_parameters() reports how much of the model is actually being updated. This is the practical check that the base model is frozen and the adapters are trainable.

5. Configuring the Training Run

TRL's SFTTrainer handled supervised fine-tuning. The notebook's training cell was configured as follows:

from trl import SFTConfig, SFTTrainer

training_args = SFTConfig(
    output_dir="./linkedin-lora",
    num_train_epochs=5,
    per_device_train_batch_size=1,
    gradient_accumulation_steps=16,
    learning_rate=2e-4,
    lr_scheduler_type="cosine",
    warmup_ratio=0.05,
    bf16=True,
    logging_steps=10,
    save_strategy="epoch",
    packing=False,
    report_to="none",
)

trainer = SFTTrainer(
    model=model,
    args=training_args,
    train_dataset=dataset,
    processing_class=tokenizer,
)

trainer.train()

The settings are easier to understand as a set of trade-offs:

SettingValueWhy it mattered
Epochs5The model saw the complete dataset five times. This is aggressive enough to make overfitting something to watch.
Device batch size1Only one example was processed at a time, keeping memory use low.
Gradient accumulation16Gradients were accumulated across 16 steps before an update, producing an effective batch size of 16.
Learning rate2e-4Controlled the size of each adapter update.
SchedulerCosineGradually reduced the learning rate over the run.
Warm-up5%Started with smaller updates before reaching the main learning rate.
PrecisionBF16Used the A100's support for efficient BF16 calculation.
Save strategyEach epochCreated checkpoints that could be compared or recovered.
PackingOffKept each short conversation as its own training sequence.

On the difference between the device batch size and the effective batch size: The GPU only held one training example for each forward and backward pass. After each pass, the gradients were retained. Once 16 examples had contributed, the optimiser performed one update.

This simulates a larger batch without needing enough VRAM to process all 16 examples at once. It is slower than fitting a batch of 16 directly on the GPU, but much more memory-efficient.

Watching the loss

During training, the loss tells us how poorly the model is predicting the target responses. We expect it to fall as the adapter learns the dataset.

The notebook used a rough expectation that loss might fall from around 2.0 towards 0.5 to 0.8. A very low training loss is not automatically good news. On a dataset of only 301 examples, a loss below roughly 0.3 could mean the adapter is memorising the training responses rather than learning a reusable style.

That is why training loss is a diagnostic, not the final score. A model can become excellent at reproducing its training data while getting worse on new prompts. In Part 3 we will discuss proper evaluation.

6. Testing Before Saving

Before merging or downloading anything, I switched the model back to evaluation mode and restored its generation cache:

model.eval()
model.config.use_cache = True

The notebook used a very direct system prompt:

SYSTEM = (
    "You are a LinkedIn post generator. Make the posts very short. "
    "Transform mundane professional updates. Use some of the following: "
    "emojis, humble-bragging, unsolicited life lessons, gratitude for "
    "irrelevant people, or dramatic emotional revelations. This is "
    "career focussed content so bring everything back to one's job."
)

Fine-tuning did not remove the need for a prompt. The adapter taught the model a behavioural tendency, while the system prompt still described the task and let me steer how strongly that behaviour should appear.

The generation function applied Gemma's chat template before tokenising the prompt:

def generate(prompt, max_new_tokens=100):
    messages = [
        {"role": "system", "content": SYSTEM},
        {"role": "user", "content": prompt},
    ]

    text = tokenizer.apply_chat_template(
        messages,
        tokenize=False,
        add_generation_prompt=True,
    )
    inputs = tokenizer(text, return_tensors="pt").to(model.device)

    with torch.no_grad():
        outputs = model.generate(
            **inputs,
            max_new_tokens=max_new_tokens,
            temperature=0.8,
            top_p=0.9,
            do_sample=True,
            pad_token_id=tokenizer.pad_token_id,
        )

    prompt_length = inputs["input_ids"].shape[1]
    generated_tokens = outputs[0][prompt_length:]
    return tokenizer.decode(
        generated_tokens,
        skip_special_tokens=True,
    )

apply_chat_template() is important because chat models expect roles and control tokens in a model-specific format. Hand-building an approximate prompt can produce different behaviour from the format used during instruction tuning.

The generation settings also introduce variation:

  • temperature=0.8 makes the output less deterministic.
  • top_p=0.9 restricts sampling to a high-probability pool of tokens.
  • do_sample=True samples from that pool rather than always choosing the single most likely next token.

The notebook tested three prompts:

print(generate("I got promoted to Head of Compliance at BoringCorp."))
print(generate("My dog died last week."))
print(generate("I just passed my AWS certification."))

These were quick smoke tests, not a proper evaluation set. They answered a narrow question: did the trained adapter produce coherent text and show the LinkedIn patterns strongly enough to justify saving it?

7. Merging the Adapter

At this point, the trained behaviour existed in the LoRA adapter. It could have been saved and loaded alongside the original Gemma model. For the demo, I instead merged the adapter into a fresh copy of the base model.

The first step was to reload that base model in FP16:

from transformers import AutoModelForCausalLM

MERGED_DIR = "./linkedin-merged"

base_model = AutoModelForCausalLM.from_pretrained(
    MODEL_ID,
    torch_dtype=torch.float16,
    device_map="cpu",
)

The training copy was stored in 4-bit form for QLoRA. The merge used an FP16 copy so the adapter update could be folded into the base weights at higher precision.

It also happened on the CPU. That avoided competing with the training model for GPU memory, although it made the merge slower and required enough ordinary system RAM to hold the FP16 model.

PEFT then loaded the trained adapter and applied its updates:

from peft import PeftModel

merged_model = PeftModel.from_pretrained(
    base_model,
    "./linkedin-lora/checkpoint-last",
)
merged_model = merged_model.merge_and_unload()

merged_model.save_pretrained(MERGED_DIR)
tokenizer.save_pretrained(MERGED_DIR)

merge_and_unload() converted the base-model-plus-adapter arrangement into one model whose weights included the learned update. That made deployment simpler because the inference code no longer needed to load a separate PEFT adapter.

There is a trap in this block: the checkpoint path must match a directory actually produced by the trainer. Depending on the trainer version and how the run was saved, the folder may be named after a training step rather than literally checkpoint-last. A reusable notebook should discover or explicitly record the chosen checkpoint instead of assuming that name.

Merging also gives up one of LoRA's useful properties. Before the merge, the adapter is small and can be swapped for another adapter using the same base model. After the merge, the result is a complete, standalone model. That was convenient for this demo, but it is not always the right production choice.

8. Downloading and Converting the Result

Finally, the notebook zipped the merged model and downloaded it:

import shutil
from google.colab import files

shutil.make_archive(
    "linkedin-merged",
    "zip",
    MERGED_DIR,
)
files.download("linkedin-merged.zip")

The saved directory used the standard Hugging Face format. From there, the notebook suggested two possible deployment routes.

For llama.cpp, it could be converted to GGUF:

python llama.cpp/convert_hf_to_gguf.py \
  linkedin-merged/ \
  --outtype q4_k_m \
  --outfile linkedin-q4.gguf

For a browser-based deployment using Transformers.js, it could instead be exported to ONNX:

pip install "optimum[exporters]"
optimum-cli export onnx \
  --model linkedin-merged/ \
  --task text-generation \
  linkedin-onnx/

Note, I was unable to get the model to deploy using Transformers.js.

These commands do not train the model again. They turn the merged Hugging Face artefact into formats suited to different inference environments.

It is also worth separating the two uses of quantisation in this workflow:

  1. Training quantisation: QLoRA loaded the frozen base model in 4-bit form to reduce training memory.
  2. Deployment quantisation: Converting the merged model to a Q4 GGUF reduced the finished artefact for efficient inference.

They both use lower-precision representations, but they happen at different stages and solve different problems.

The dataset makes the product

The trainer cannot infer what I mean by "LinkedIn-like" unless I show examples demonstrating it consistently. Every trope had to appear often enough to become a learnable pattern, but not so mechanically that the model only learnt one template.

Three hundred examples were enough to produce a visible behavioural change in this narrow experiment. That does not make 300 a generally sufficient number. A broader task, more varied input, a larger model or stricter reliability requirements could need far more data.

Parameter-efficient does not mean decision-free

QLoRA made this experiment practical, but it did not choose the rank, target modules, learning rate, number of epochs or stopping point for me. These settings control how much capacity the adapter has and how aggressively it learns.

The sensible values depend on the model, task and data. The values in this notebook are a working configuration, and are not perfect, so do not copy this blindly.

The test prompt is not the evaluation

Seeing a ridiculous LinkedIn post appear after training is satisfying, and it proves that the pipeline basically worked. It does not tell us whether the fine-tuned model is consistently better than the baseline, whether it handles prompts outside the training distribution, or whether it has lost any useful behaviour.

Those questions need a held-out test set, repeatable criteria and comparison against the unchanged base model. That is the subject of Part 3.

The Complete Flow

The notebook can be reduced to eight stages:

  1. Load about 300 system, user and assistant training examples from JSONL.
  2. Confirm that an A100 GPU is available.
  3. Load google/gemma-3-2B-it in 4-bit NF4 form.
  4. Attach rank-16 LoRA adapters to selected MLP projection modules.
  5. Train those adapters for five epochs with an effective batch size of 16.
  6. Run a few new prompts through the fine-tuned model.
  7. Merge the chosen adapter checkpoint into an FP16 copy of the base model.
  8. Save and convert the result for the target inference environment.

The base model remained frozen during training. The LoRA adapters learnt a compact set of adjustments. The merge then folded those adjustments into a standalone model that could be deployed like any other Hugging Face model.

Conclusion

Fine-tuning a model did not require writing a training algorithm from scratch. Transformers loaded the model, BitsAndBytes quantised it, PEFT attached the adapters, and TRL ran the supervised training loop.

My job was to make the decisions those libraries could not make for me such as what behaviour the examples should demonstrate, which model to adapt, which modules to target, how long to train, and what evidence would count as improvement.

In Part 3, I will look at that evaluation problem: how to compare the fine-tuned and base models fairly, how to turn "this feels more LinkedIn" into repeatable criteria, and how to spot memorisation or regression before calling a fine-tune successful.

← Previous

Fine-Tuning: Training The Actor

Next →

pgvector Setup for RAG