RepoDaily · 2026-08-17 · Infrastructure / Runtime

Soup Review: Fine-Tuning an 8B LLM on a 4 GB Laptop GPU From One YAML

#6 Infrastructure / Runtime Python +456 MakazhanAlpamys/Soup Open repository

Soup CLI turns fine-tuning into one YAML and one command. Its opt-in layer streaming trains Llama-3.1-8B in 3.32 GB of VRAM — we checked what the repo documents, benchmarks, and admits.

Repo typeInfrastructure / Runtime
Best forLaptop-GPU owners and small labs fine-tuning 7–8B models within 4 GB, plus shops that want one YAML to drive SFT/DPO/GRPO with ship gates
Risk levelMedium — layer streaming is labeled BETA, the headline tok/s predates the v0.73.0 correctness repair, and only 0.73.x receives security fixes
Time to evaluate~30 minutes: install with the [train] extra, run the chat template, then reproduce notebooks/proof-4gb.ipynb on a free Colab T4

Primary question: Does one YAML-driven CLI with bit-exact 4 GB training hold up on your hardware, at your model sizes?

89/100

RepoDaily adoption score

RepoDaily rates this as 89/100 (strong) for adoption: evidence, installation path, production risk, differentiation, license clarity, and AI/agent fit are scored from the article sources and adoption notes.

Directional score from RepoDaily sources and adoption notes, not a benchmark.Risk: Medium
100Evidence quality

8 source(s) across 4 source category/categories, plus a RepoDaily-specific evidence module when available.

99Installability

6 workflow step(s), 4 next-action step(s), and 2 command/install signal(s) were detected.

63Maintenance confidence

Trending momentum is +456 stars, with maintenance/release/issue signals counted when present.

93Production readiness

Risk is marked medium, with 5 security note(s) and 5 explicit skip condition(s).

100Differentiation

3 opportunity lens item(s), 4 alternative(s), and 4 type-specific section(s) support differentiation.

68License clarity

License source or license wording is present.

78Agent / AI fit

5 AI/agent-related signal(s) were detected in the article text and metadata.

Project overview

Soup is a local-first Python CLI that compresses LLM post-training into one config file and one command: `pip install 'soup-cli[train]'`, `soup init --template chat`, `soup train`. The README badge row pins the basics — Apache-2.0 license, Python 3.10–3.12, PyPI distribution `soup-cli`, CI badge, and a DOI (10.5281/zenodo.21771064) for the layer-streaming paper. The pitch is blunt: 'No SSH, no config hell.'

The headline claim is unusually specific. Layer streaming keeps the frozen base model out of VRAM and feeds it to the GPU one decoder layer at a time; measured on an RTX 3050 Laptop 4 GB, Llama-3.1-8B-Instruct quantized to NF4 ran at 119.6 tok/s with a 3.32 GB peak, asserted bit-exact against a normal resident run and independently reproduced on an H100 at 113.00 tok/s in the same 3.32 GB. The feature is opt-in (`stream_layers: true`), still BETA, and the tok/s figure was measured on v0.72.2 — before the v0.73.0 correctness repair that cost −4.8% at 32B — and has not been re-run on a 4 GB card since.

Beyond the 4 GB trick, the docs index maps a wide surface: SFT through GRPO/PPO/KTO/ORPO and unlearning, a PEFT zoo from DoRA to VeRA, QAT/FP8/NVFP4 quantization, an OpenAI-compatible server, a model registry and `.can` artifacts, HIPAA/SOC2/EU-AI-Act/SR-11-7 compliance templates, MLX and Unsloth backends, and 144 ready-made recipes. This article reads the README, the docs index, the changelog, the security policy, and the contributing guide to separate documented behavior from marketing.

Problem it solves

  • Fine-tuning configs sprawl across multiple files and remote machines; Soup's answer is one `soup.yaml` and one command, with 'No SSH, no config hell' as the tagline.
  • VRAM walls: a resident 8B NF4 base plus optimizer state pushes past 4 GB, pushing laptop users toward cloud rentals they may not want.
  • Ship decisions are usually manual; Soup wires `soup ship` verdicts, eval-gated training, and a `soup advise` step into the same config.
  • Provenance for regulated settings (BOM, attestations, repro receipts) is typically hand-built; Soup ships `soup attest`, `soup card`, and `soup ci init` with named templates for HIPAA/SOC2/EU-AI-Act/SR-11-7.
  • Switching trainers means rewriting configs; `src/soup_cli/migrate/` imports existing LLaMA-Factory, Axolotl, and Unsloth setups.

How it works

  1. `pip install 'soup-cli[train]'` — the `[train]` extra pulls torch, transformers, peft, trl, datasets, bitsandbytes, accelerate; bare `soup-cli` stays a light CLI (extras are opt-in since v0.71.0, not core deps).
  2. `soup init --template chat` writes a `soup.yaml`; alternatively start from one of 144 recipes, such as `qwen3.5-4b-pretrain` (#278) or `deepseek-v4-flash-grpo` (#279).
  3. Choose a method in the YAML: SFT, DPO/GRPO/PPO/KTO/ORPO/SimPO/IPO/BCO, plus pre-training, distillation, classification, PRM, unlearning, and more, each a wrapper under `src/soup_cli/trainer/`.
  4. Run `soup train`. With `stream_layers: true`, the frozen base is pinned in host RAM (a 3.60 GB store across 32 layers) and fed to the GPU one decoder layer at a time through two 113 MB VRAM buffers.
  5. Gate the result: `soup ship` emits a verdict from gate policy, and `eval.ship.noise_floor` is now committable to `soup.yaml` (#406) — bounded to [2, 10], bool-as-int rejected, with CLI > config > default precedence.
  6. Export or serve: merge/export, an OpenAI-compatible server, an Anthropic Messages endpoint, `.can` shareable artifacts, and `soup card` for model-card generation.

Product demo and interface preview

soup train pre-flight for Llama-3.1-8B on a 4 GB card: a 3.60 GB base store pinned in RAM across 32 layers, two 113 MB VRAM buffers, and a measured 3.32 GB peak at 119.6 tok/s that stops short of the 4 GB line
Layer streaming pre-flight — The README's own recording of a 4 GB training run shows where the base weights live (host RAM) versus what touches VRAM — the mechanic behind the headline claim. README.md image

Architecture read: how layer streaming fits in 4 GB

Layer streaming inverts the usual residency model. The frozen base never sits in VRAM; instead a 3.60 GB store is pinned in host RAM and sliced across the model's 32 decoder layers, while two 113 MB VRAM buffers shuttle one layer at a time to the GPU. Trainable adapters and optimizer state stay resident. On an RTX 3050 Laptop 4 GB this yields a measured 3.32 GB peak at 119.6 tok/s for Llama-3.1-8B-Instruct at NF4 — under the 4 GB line — and the streamed run is asserted bit-exact against a normal resident run, independently reproduced on an H100 at 113.00 tok/s in the same 3.32 GB.

The feature is tiered by disk, and the tiers are where the engineering shows. `detect_disk_kind` originally trusted `/sys/block/<dev>/queue/rotational`, which misread paravirtual virtio disks (defaulting to `1`) as HDDs — denying the disk-overflow tier to genuinely NVMe-backed cloud disks measured at 1.5 GB/s read. Fix #365 replaced the flag with a bounded O_DIRECT sequential-read measurement when `rotational=1`: throughput ≥ 1 GB/s earns the tier, a genuinely slow disk is still refused (160 seeks/step, plan P11), and a new `training.stream_disk_kind` setting exposes the decision in config.

Two honesty markers matter for anyone sizing this. First, the feature is opt-in (`stream_layers: true`) and labeled BETA, having gained NF4 in v0.72.2, disk support and wider architectures in v0.72.3, and preference losses in v0.72.4 per the docs anchor. Second, the 119.6 tok/s headline was measured on v0.72.2 — before the v0.73.0 correctness repair that cost −4.8% at 32B — and has not been re-run on a 4 GB card since. The number is real but stale by one correctness release.

Command surface: what `soup` exposes

  • `soup init --template chat` and `soup train` — the two-command happy path from the README quick start.
  • `soup ship` — verdict from gate policy; with #406, `eval.ship.noise_floor` (bounded [2, 10]) joins the other five committable gate flags, excluded from the recipe `config_sha` so setting a floor never invalidates evidence.
  • `soup runs` — SQLite-backed experiment tracking; after #401, a run whose watcher died reconciles on read to `terminated` with an unknown exit code instead of showing `running` forever.
  • `soup card`, `soup ci init`, `soup attest`, `soup adapters sign` — model-card autogeneration, CI gating, and ed25519 provenance; the `[sign]` extra pulls `cryptography` so these tests run in CI.
  • `soup shrink` — depth pruning plus distill-heal, listed in the PEFT and efficiency guide.
  • `soup loop` and `soup advise` — the data flywheel and post-train recommendation steps named in the docs index.
  • `soup train --annex-xi *.pdf` — Annex XI/XII report output, backed by the `[pdf]` extra (`reportlab`).
  • The full command list lives in `docs/commands.md`; `docs/models.md` holds recommended model families, the VRAM size guide, and the pip extras matrix.

Maintenance risk: what the changelog admits

  • #401 (fixed): `ExecutionManager._watch` runs as a `daemon=True` thread, so an MCP-server exit killed it without unwinding and runs stayed `running` forever; the tracker now reconciles on read and never records an unknown outcome as success.
  • #431 (merged, by @Shutaru): MLX SFT dispatch no longer imports the Transformers SFT wrapper; an additive Apple Silicon CI job verifies that `mlx` and `mlx-lm` import, that the PyTorch/TRL stack is absent, and that a one-step real CLI SFT run completes.
  • #394 (open): the changelog states plainly that #431 'does not claim to resolve the still-unpinned torch-present hang' — a known Apple Silicon sharp edge.
  • Version policy: SECURITY.md supports only 0.73.x; the changelog notes 70+ published versions, so older deployments age out of fixes quickly.
  • Release velocity is high — v0.72.0 through v0.73.2 shipped layer streaming, a correctness repair, and new gate flags in quick succession — so pin versions in CI and re-verify after upgrades.

Integration surface: backends, hubs, migration

  • Backends: MLX and Unsloth are first-class (`soup-cli[mlx]` standalone hardened by #431), alongside the PyTorch/TRL stack via `[train]`.
  • Cloud: Modal cloud GPU training, alternative hubs, and HF Hub integration are listed in backends-and-ops.md.
  • Serving: an OpenAI-compatible server, an Anthropic Messages endpoint, batch inference, and speculative decoding — including training your own draft model.
  • Migration: `src/soup_cli/migrate/` converts LLaMA-Factory, Axolotl, and Unsloth configs into Soup's schema.
  • Tracking and ops: `src/soup_cli/experiment/` uses SQLite; env lockfiles and hardware-fit checks round out the backends-and-ops guide.
  • Scale-out: multi-GPU, DeepSpeed, and FSDP appear in performance-and-quantization.md for when 4 GB tricks are not enough.

Who should pay attention?

Good fit if

  • You own a 4–8 GB laptop GPU and want to fine-tune 7–8B models locally instead of renting cloud cards.
  • You want one `soup.yaml` to carry a project from SFT through DPO/GRPO to a gated `soup ship` decision.
  • You sit under HIPAA/SOC2/EU-AI-Act/SR-11-7 constraints and want init templates, BOMs, attestations, and `soup ci init` built in.
  • You are on Apple Silicon and want the `[mlx]` path, which CI verifies runs without the PyTorch/TRL stack.
  • You have existing LLaMA-Factory, Axolotl, or Unsloth configs to port via `src/soup_cli/migrate/`.

Skip for now if

  • You need non-BETA throughput guarantees today — the headline tok/s predates the v0.73.0 correctness repair and has not been re-measured on a 4 GB card.
  • You must run Python 3.13 or newer — `requires-python` is >=3.10,<3.13, and CI enforces exactly 3.10, 3.11, 3.12.
  • You need long-term support for pinned older versions — SECURITY.md supports only 0.73.x and nothing below.
  • You require SLA-backed vendor support — reporting runs through GitHub Security Advisories and two email addresses, with no bounty.
  • Your jobs are large multi-node runs already orchestrated elsewhere; the 4 GB residency trick is not your bottleneck.

Risks and cautions

Medium

The core claim is falsifiable and the docs are candid, but layer streaming is explicitly BETA, the headline tok/s predates the v0.73.0 correctness repair, and only the 0.73.x line receives security fixes.

  • `stream_layers: true` is opt-in and labeled BETA; the 119.6 tok/s figure was measured on v0.72.2, before the v0.73.0 correctness repair that cost −4.8% at 32B, and has not been re-run on a 4 GB card since.
  • SECURITY.md supports only 0.73.x; anything older is out of security support, making upgrades effectively mandatory.
  • The changelog itself flags an unresolved Apple Silicon hang (#394) that the #431 MLX dispatch fix explicitly does not claim to resolve.
  • 70+ published versions plus unreleased items like #406 show fast churn; config and gate behavior can shift between minor releases.
  • Support lanes are GitHub Security Advisories plus two email addresses ([email protected], [email protected]) — a thin, single-maintainer-shaped reporting path.
  • License is Apache-2.0 per the README badge.
  • Supported versions: 0.73.x only; the policy table marks everything below 0.73 as unsupported.
  • In-scope threat classes: path traversal or arbitrary file read/write from config, dataset, or artifact paths; SSRF in synthetic-data providers, the inference server, and hub/endpoint validators; command, Modelfile, Jinja chat-template, and systemd/launchd injection; secret leakage in logs, crash bundles, or generated artifacts; sandbox escape in the RLVR code-execution reward path.
  • Reporting is private-only: GitHub Security Advisories preferred, else [email protected] or the maintainer's personal address; Discord is explicitly rejected as a channel. Acknowledgment target is 5 business days, and there is no bug bounty — credit in release notes is the stated reward.
  • Out of scope by policy: vulnerabilities in third-party weights or datasets you load, already-compromised hosts, and trysoup.dev DNS/email configuration such as DMARC/SPF/DKIM records.

Alternatives to compare

ApproachWhen to useTrade-off
Axolotl
You want a mature YAML-driven multi-GPU trainer with broad method coverage and do not need the 4 GB layer-streaming residency trick; Soup's migrate/ path treats it as a first-class import source.Open source (Apache-2.0)
LLaMA-Factory
You prefer a web UI on top of YAML configs and broad model coverage; also supported as a Soup migration source.Open source
Single-GPU throughput on NVIDIA/AMD cards matters more than one-config portability; usable inside Soup as a backend and as a migration source.Open source (Apache-2.0)
Hugging Face TRL
You want the library Soup itself builds on (the [train] extra pulls trl) with direct control in Python rather than a CLI.Open source (Apache-2.0)

What this trend reveals

Tiny-GPU fine-tuning as a wedge

The 3.32 GB / 119.6 tok/s measurement plus a self-verifying notebook targets the large population of students, consultants, and tinkerers whose only GPU is a 4 GB laptop card — a segment most trainers ignore.

Open notebooks/proof-4gb.ipynb on a free Colab T4; confirm the bit-identical assertion passes and record peak VRAM and tok/s on your own card.

Compliance-grade post-training

compliance.md ships init templates for HIPAA, SOC2, EU-AI-Act, and SR-11-7, plus BOM/attest/repro-receipt provenance, air-gap notes, `soup card`, and `soup ci init` — the pieces a regulated shop would otherwise assemble by hand.

Wire `soup ci init` into one CI job and confirm the gate blocks a recipe whose eval regresses past `eval.ship.noise_floor`.

Config migration as an on-ramp

`src/soup_cli/migrate/` converts LLaMA-Factory, Axolotl, and Unsloth configs, lowering switching cost for anyone already invested in those tools.

Convert one existing YAML, run 200 steps under Soup and under the original trainer, and diff loss curves before committing.

Best next action

Reproduce the 4 GB claim before adopting

Soup's strongest selling point is falsifiable in a single notebook: the process is capped to 4 GB and the streamed model is asserted bit-identical to a resident run. Verify it on hardware you actually own, then re-check after the next release since the published tok/s predates the v0.73.0 correctness repair.

  1. Install with `pip install 'soup-cli[train]'` on a machine with a 4 GB card, or open notebooks/proof-4gb.ipynb on a free Colab T4.
  2. Run `soup init --template chat`, then `soup train` with `stream_layers: true`.
  3. Confirm the notebook's bit-identical assertion between the streamed and resident runs.
  4. Record peak VRAM and tok/s, and pin the version you tested; re-run before any upgrade because only 0.73.x receives security fixes.

RepoDaily verdict

Soup earns its trend slot with a falsifiable claim — an 8B model trained in 3.32 GB, bit-exact and reproducible on a free Colab T4 — backed by candid docs and a changelog that names its own bugs. Treat layer streaming as BETA, pin 0.73.x, and re-measure on your hardware before standardizing.

Sources