Soup: Fine-Tuning an 8B Model on a 4 GB Laptop GPU, One Layer at a Time

Soup is an open-source tool that fine-tunes large language models from a single config. Its layer streaming mode trains an 8B model at 119.6 tokens per second on a 4 GB laptop GPU, and the project publishes the correctness bugs it found.
artificial-intelligence
software-engineering
Author

Kabui, Charles

Published

2026-09-14

Keywords

layer-streaming, low-vram-fine-tuning, lora, local-ai, model-fine-tuning