Gemma 4 fine-tune
One QLoRA run on past agent conversations.
What it is
The goal was to build the assistant's working style into a small local model and reduce looping in long conversations. A script turned past agent session logs into instruction pairs, a Colab notebook trained LoRA adapters on Gemma 4 with Unsloth, and the result was merged and exported for Ollama.
By the numbers
What it means for your team
Fine-tuning is rarely the first tool to reach for. Better instructions, retrieval and a strong hosted model usually get there first, and a fine-tune needs an evaluation set before anyone relies on it.
Status
A single run that was never evaluated or deployed, and its outputs were later deleted. The training data is private conversation history, so it cannot be shown.
Dataset
A Python script turned agent session logs into instruction pairs, dropping canned intros, looping replies and very short answers.
Training
A Colab notebook trained QLoRA adapters (rank 16, attention and MLP layers) on a free T4 GPU.
Export
The adapters were merged into the base model and exported as a 4-bit GGUF file for Ollama.