Primary question: Does one YAML-driven CLI with bit-exact 4 GB training hold up on your hardware, at your model sizes?
RepoDaily adoption score
RepoDaily rates this as 89/100 (strong) for adoption: evidence, installation path, production risk, differentiation, license clarity, and AI/agent fit are scored from the article sources and adoption notes.
8 source(s) across 4 source category/categories, plus a RepoDaily-specific evidence module when available.
6 workflow step(s), 4 next-action step(s), and 2 command/install signal(s) were detected.
Trending momentum is +456 stars, with maintenance/release/issue signals counted when present.
Risk is marked medium, with 5 security note(s) and 5 explicit skip condition(s).
3 opportunity lens item(s), 4 alternative(s), and 4 type-specific section(s) support differentiation.
License source or license wording is present.
5 AI/agent-related signal(s) were detected in the article text and metadata.
Project overview
Soup is a local-first Python CLI that compresses LLM post-training into one config file and one command: `pip install 'soup-cli[train]'`, `soup init --template chat`, `soup train`. The README badge row pins the basics — Apache-2.0 license, Python 3.10–3.12, PyPI distribution `soup-cli`, CI badge, and a DOI (10.5281/zenodo.21771064) for the layer-streaming paper. The pitch is blunt: 'No SSH, no config hell.'
The headline claim is unusually specific. Layer streaming keeps the frozen base model out of VRAM and feeds it to the GPU one decoder layer at a time; measured on an RTX 3050 Laptop 4 GB, Llama-3.1-8B-Instruct quantized to NF4 ran at 119.6 tok/s with a 3.32 GB peak, asserted bit-exact against a normal resident run and independently reproduced on an H100 at 113.00 tok/s in the same 3.32 GB. The feature is opt-in (`stream_layers: true`), still BETA, and the tok/s figure was measured on v0.72.2 — before the v0.73.0 correctness repair that cost −4.8% at 32B — and has not been re-run on a 4 GB card since.
Beyond the 4 GB trick, the docs index maps a wide surface: SFT through GRPO/PPO/KTO/ORPO and unlearning, a PEFT zoo from DoRA to VeRA, QAT/FP8/NVFP4 quantization, an OpenAI-compatible server, a model registry and `.can` artifacts, HIPAA/SOC2/EU-AI-Act/SR-11-7 compliance templates, MLX and Unsloth backends, and 144 ready-made recipes. This article reads the README, the docs index, the changelog, the security policy, and the contributing guide to separate documented behavior from marketing.
Why it is trending now
- 456 stars in the window and rank 6 on the 2026-08-17 trend list.
- A concrete, falsifiable claim: Llama-3.1-8B-Instruct + NF4 at 119.6 tok/s and 3.32 GB peak on an RTX 3050 Laptop 4 GB, bit-exact against a resident run and reproduced on an H100 at 113.00 tok/s.
- The claim ships with `notebooks/proof-4gb.ipynb`, which caps the process to 4 GB on a free Colab T4 and asserts the streamed model is bit-identical to a normal one.
- Breadth under one CLI: 18 trainer wrappers under `src/soup_cli/trainer/`, 144 recipes, PEFT variants from DoRA and LoRA+ to VeRA and PiSSA, and compliance init templates.
- Candor as a differentiator: a DOI'd paper, a `benchmarks/` directory, and a changelog that names its own unresolved bugs (#394) and the −4.8% cost of its correctness repair.
Problem it solves
- Fine-tuning configs sprawl across multiple files and remote machines; Soup's answer is one `soup.yaml` and one command, with 'No SSH, no config hell' as the tagline.
- VRAM walls: a resident 8B NF4 base plus optimizer state pushes past 4 GB, pushing laptop users toward cloud rentals they may not want.
- Ship decisions are usually manual; Soup wires `soup ship` verdicts, eval-gated training, and a `soup advise` step into the same config.
- Provenance for regulated settings (BOM, attestations, repro receipts) is typically hand-built; Soup ships `soup attest`, `soup card`, and `soup ci init` with named templates for HIPAA/SOC2/EU-AI-Act/SR-11-7.
- Switching trainers means rewriting configs; `src/soup_cli/migrate/` imports existing LLaMA-Factory, Axolotl, and Unsloth setups.
How it works
- `pip install 'soup-cli[train]'` — the `[train]` extra pulls torch, transformers, peft, trl, datasets, bitsandbytes, accelerate; bare `soup-cli` stays a light CLI (extras are opt-in since v0.71.0, not core deps).
- `soup init --template chat` writes a `soup.yaml`; alternatively start from one of 144 recipes, such as `qwen3.5-4b-pretrain` (#278) or `deepseek-v4-flash-grpo` (#279).
- Choose a method in the YAML: SFT, DPO/GRPO/PPO/KTO/ORPO/SimPO/IPO/BCO, plus pre-training, distillation, classification, PRM, unlearning, and more, each a wrapper under `src/soup_cli/trainer/`.
- Run `soup train`. With `stream_layers: true`, the frozen base is pinned in host RAM (a 3.60 GB store across 32 layers) and fed to the GPU one decoder layer at a time through two 113 MB VRAM buffers.
- Gate the result: `soup ship` emits a verdict from gate policy, and `eval.ship.noise_floor` is now committable to `soup.yaml` (#406) — bounded to [2, 10], bool-as-int rejected, with CLI > config > default precedence.
- Export or serve: merge/export, an OpenAI-compatible server, an Anthropic Messages endpoint, `.can` shareable artifacts, and `soup card` for model-card generation.
Product demo and interface preview

Architecture read: how layer streaming fits in 4 GB
Layer streaming inverts the usual residency model. The frozen base never sits in VRAM; instead a 3.60 GB store is pinned in host RAM and sliced across the model's 32 decoder layers, while two 113 MB VRAM buffers shuttle one layer at a time to the GPU. Trainable adapters and optimizer state stay resident. On an RTX 3050 Laptop 4 GB this yields a measured 3.32 GB peak at 119.6 tok/s for Llama-3.1-8B-Instruct at NF4 — under the 4 GB line — and the streamed run is asserted bit-exact against a normal resident run, independently reproduced on an H100 at 113.00 tok/s in the same 3.32 GB.
The feature is tiered by disk, and the tiers are where the engineering shows. `detect_disk_kind` originally trusted `/sys/block/<dev>/queue/rotational`, which misread paravirtual virtio disks (defaulting to `1`) as HDDs — denying the disk-overflow tier to genuinely NVMe-backed cloud disks measured at 1.5 GB/s read. Fix #365 replaced the flag with a bounded O_DIRECT sequential-read measurement when `rotational=1`: throughput ≥ 1 GB/s earns the tier, a genuinely slow disk is still refused (160 seeks/step, plan P11), and a new `training.stream_disk_kind` setting exposes the decision in config.
Two honesty markers matter for anyone sizing this. First, the feature is opt-in (`stream_layers: true`) and labeled BETA, having gained NF4 in v0.72.2, disk support and wider architectures in v0.72.3, and preference losses in v0.72.4 per the docs anchor. Second, the 119.6 tok/s headline was measured on v0.72.2 — before the v0.73.0 correctness repair that cost −4.8% at 32B — and has not been re-run on a 4 GB card since. The number is real but stale by one correctness release.
Command surface: what `soup` exposes
- `soup init --template chat` and `soup train` — the two-command happy path from the README quick start.
- `soup ship` — verdict from gate policy; with #406, `eval.ship.noise_floor` (bounded [2, 10]) joins the other five committable gate flags, excluded from the recipe `config_sha` so setting a floor never invalidates evidence.
- `soup runs` — SQLite-backed experiment tracking; after #401, a run whose watcher died reconciles on read to `terminated` with an unknown exit code instead of showing `running` forever.
- `soup card`, `soup ci init`, `soup attest`, `soup adapters sign` — model-card autogeneration, CI gating, and ed25519 provenance; the `[sign]` extra pulls `cryptography` so these tests run in CI.
- `soup shrink` — depth pruning plus distill-heal, listed in the PEFT and efficiency guide.
- `soup loop` and `soup advise` — the data flywheel and post-train recommendation steps named in the docs index.
- `soup train --annex-xi *.pdf` — Annex XI/XII report output, backed by the `[pdf]` extra (`reportlab`).
- The full command list lives in `docs/commands.md`; `docs/models.md` holds recommended model families, the VRAM size guide, and the pip extras matrix.
Maintenance risk: what the changelog admits
- #401 (fixed): `ExecutionManager._watch` runs as a `daemon=True` thread, so an MCP-server exit killed it without unwinding and runs stayed `running` forever; the tracker now reconciles on read and never records an unknown outcome as success.
- #431 (merged, by @Shutaru): MLX SFT dispatch no longer imports the Transformers SFT wrapper; an additive Apple Silicon CI job verifies that `mlx` and `mlx-lm` import, that the PyTorch/TRL stack is absent, and that a one-step real CLI SFT run completes.
- #394 (open): the changelog states plainly that #431 'does not claim to resolve the still-unpinned torch-present hang' — a known Apple Silicon sharp edge.
- Version policy: SECURITY.md supports only 0.73.x; the changelog notes 70+ published versions, so older deployments age out of fixes quickly.
- Release velocity is high — v0.72.0 through v0.73.2 shipped layer streaming, a correctness repair, and new gate flags in quick succession — so pin versions in CI and re-verify after upgrades.
Integration surface: backends, hubs, migration
- Backends: MLX and Unsloth are first-class (`soup-cli[mlx]` standalone hardened by #431), alongside the PyTorch/TRL stack via `[train]`.
- Cloud: Modal cloud GPU training, alternative hubs, and HF Hub integration are listed in backends-and-ops.md.
- Serving: an OpenAI-compatible server, an Anthropic Messages endpoint, batch inference, and speculative decoding — including training your own draft model.
- Migration: `src/soup_cli/migrate/` converts LLaMA-Factory, Axolotl, and Unsloth configs into Soup's schema.
- Tracking and ops: `src/soup_cli/experiment/` uses SQLite; env lockfiles and hardware-fit checks round out the backends-and-ops guide.
- Scale-out: multi-GPU, DeepSpeed, and FSDP appear in performance-and-quantization.md for when 4 GB tricks are not enough.
Who should pay attention?
Good fit if
- You own a 4–8 GB laptop GPU and want to fine-tune 7–8B models locally instead of renting cloud cards.
- You want one `soup.yaml` to carry a project from SFT through DPO/GRPO to a gated `soup ship` decision.
- You sit under HIPAA/SOC2/EU-AI-Act/SR-11-7 constraints and want init templates, BOMs, attestations, and `soup ci init` built in.
- You are on Apple Silicon and want the `[mlx]` path, which CI verifies runs without the PyTorch/TRL stack.
- You have existing LLaMA-Factory, Axolotl, or Unsloth configs to port via `src/soup_cli/migrate/`.
Skip for now if
- You need non-BETA throughput guarantees today — the headline tok/s predates the v0.73.0 correctness repair and has not been re-measured on a 4 GB card.
- You must run Python 3.13 or newer — `requires-python` is >=3.10,<3.13, and CI enforces exactly 3.10, 3.11, 3.12.
- You need long-term support for pinned older versions — SECURITY.md supports only 0.73.x and nothing below.
- You require SLA-backed vendor support — reporting runs through GitHub Security Advisories and two email addresses, with no bounty.
- Your jobs are large multi-node runs already orchestrated elsewhere; the 4 GB residency trick is not your bottleneck.
Risks and cautions
The core claim is falsifiable and the docs are candid, but layer streaming is explicitly BETA, the headline tok/s predates the v0.73.0 correctness repair, and only the 0.73.x line receives security fixes.
- `stream_layers: true` is opt-in and labeled BETA; the 119.6 tok/s figure was measured on v0.72.2, before the v0.73.0 correctness repair that cost −4.8% at 32B, and has not been re-run on a 4 GB card since.
- SECURITY.md supports only 0.73.x; anything older is out of security support, making upgrades effectively mandatory.
- The changelog itself flags an unresolved Apple Silicon hang (#394) that the #431 MLX dispatch fix explicitly does not claim to resolve.
- 70+ published versions plus unreleased items like #406 show fast churn; config and gate behavior can shift between minor releases.
- Support lanes are GitHub Security Advisories plus two email addresses ([email protected], [email protected]) — a thin, single-maintainer-shaped reporting path.
- License is Apache-2.0 per the README badge.
- Supported versions: 0.73.x only; the policy table marks everything below 0.73 as unsupported.
- In-scope threat classes: path traversal or arbitrary file read/write from config, dataset, or artifact paths; SSRF in synthetic-data providers, the inference server, and hub/endpoint validators; command, Modelfile, Jinja chat-template, and systemd/launchd injection; secret leakage in logs, crash bundles, or generated artifacts; sandbox escape in the RLVR code-execution reward path.
- Reporting is private-only: GitHub Security Advisories preferred, else [email protected] or the maintainer's personal address; Discord is explicitly rejected as a channel. Acknowledgment target is 5 business days, and there is no bug bounty — credit in release notes is the stated reward.
- Out of scope by policy: vulnerabilities in third-party weights or datasets you load, already-compromised hosts, and trysoup.dev DNS/email configuration such as DMARC/SPF/DKIM records.
Alternatives to compare
| Approach | When to use | Trade-off |
|---|---|---|
Axolotl | You want a mature YAML-driven multi-GPU trainer with broad method coverage and do not need the 4 GB layer-streaming residency trick; Soup's migrate/ path treats it as a first-class import source. | Open source (Apache-2.0) |
LLaMA-Factory | You prefer a web UI on top of YAML configs and broad model coverage; also supported as a Soup migration source. | Open source |
| Single-GPU throughput on NVIDIA/AMD cards matters more than one-config portability; usable inside Soup as a backend and as a migration source. | Open source (Apache-2.0) | |
Hugging Face TRL | You want the library Soup itself builds on (the [train] extra pulls trl) with direct control in Python rather than a CLI. | Open source (Apache-2.0) |
What this trend reveals
Tiny-GPU fine-tuning as a wedge
The 3.32 GB / 119.6 tok/s measurement plus a self-verifying notebook targets the large population of students, consultants, and tinkerers whose only GPU is a 4 GB laptop card — a segment most trainers ignore.
Open notebooks/proof-4gb.ipynb on a free Colab T4; confirm the bit-identical assertion passes and record peak VRAM and tok/s on your own card.
Compliance-grade post-training
compliance.md ships init templates for HIPAA, SOC2, EU-AI-Act, and SR-11-7, plus BOM/attest/repro-receipt provenance, air-gap notes, `soup card`, and `soup ci init` — the pieces a regulated shop would otherwise assemble by hand.
Wire `soup ci init` into one CI job and confirm the gate blocks a recipe whose eval regresses past `eval.ship.noise_floor`.
Config migration as an on-ramp
`src/soup_cli/migrate/` converts LLaMA-Factory, Axolotl, and Unsloth configs, lowering switching cost for anyone already invested in those tools.
Convert one existing YAML, run 200 steps under Soup and under the original trainer, and diff loss curves before committing.
RepoDaily verdict
Soup earns its trend slot with a falsifiable claim — an 8B model trained in 3.32 GB, bit-exact and reproducible on a free Colab T4 — backed by candid docs and a changelog that names its own bugs. Treat layer streaming as BETA, pin 0.73.x, and re-measure on your hardware before standardizing.
Sources
- README.md (install, quick start, layer-streaming benchmark)
- docs/README.md (documentation index)
- CHANGELOG.md (unreleased changes)
- SECURITY.md (security policy)
- CONTRIBUTING.md (dev setup and project layout)
- GitHub repository (releases and issue tracker)
- trysoup.dev (official website)
- PyPI: soup-cli (package page)