LoRA and QLoRA: Efficient Fine-Tuning on a Single Laptop
LoRA cuts fine-tuning cost for large language models by training only small low-rank adaptation matrices instead of every parameter in the base model. QLoRA adds 4-bit quantization on top, cutting required GPU memory by 65-75%, with quality loss of just 1-3% versus full fine-tuning.