Skip to content

Add Fizgig LoRA training studio package - #1761

Open
e-nord wants to merge 4 commits into
LykosAI:mainfrom
e-nord:feat-fizgig-package
Open

e-nord wants to merge 4 commits into
LykosAI:mainfrom
e-nord:feat-fizgig-package

Conversation

@e-nord

@e-nord e-nord commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Adds Fizgig as a new package. Fizgig is a "train · fine-tune · repair · explore" workbench for generative image models. Its stated focus is getting LoRA training working on consumer GPUs, fixing broken LoRAs without retraining, and making variations in seconds.

Upstream: https://github.com/shootthesound/Fizgig (Apache-2.0, tagged semver releases)

Why it fits Stability Matrix

  • Focus on newer models: Flux 2 Klein 9B, Krea 2, MiniMax H3 and Qwen Image 2.1 (LoRA training on all of them, plus LoKR and experimental full fine-tuning on some). Minimum VRAM runs from roughly 8 to 16 GB depending on the model.
  • Fast LoRA workflow: adaptive LR, weight averaging, context LoRA, pause/resume, and dataset curation with loss watch.
  • Wizard-style UI: numbered tabs (Start → Image Prep → Captions → Samples → Training) that walk users from dataset to a running training job.
  • More than training: Repair Studio (block-level sliders), LoRA the Explorer, LoRA Royale (epoch comparison and export), Profiler, and Extract (rank reduction).
  • Popularity: about 445 GitHub stars and 42 forks since the repo was created in April 2026.

Scope of changes

  • New Fizgig package with appropriately integrated references into SM. It's a Tkinter desktop app, so it's similar to OneTrainer.
  • Docs: README, supported-packages and hardware-support entries.
  • No changes to shared or framework code.

Notes on the less obvious choices:

  • Visual Studio Build Tools prerequisite (VcBuildTools). This lets triton / torch.compile's inductor backend compile kernels on Windows, which Fizgig's Compile Blocks speedup needs. Among the trainers, only Fizgig requires it (ComfyZluda is the only other package that does), and it makes the first install take a bit longer because the VS bootstrapper runs. Open to making it an option to make this explicitly opt-in.
  • Launches lora_trainer_gui.py directly, not launch.pyw. launch.pyw re-spawns itself under pythonw.exe and exits, which would orphan the GUI and break console output and Stop.
  • DISABLE_CUDA=1 during the requirements install. Without it, hqq's sdist tries to build CUDA kernels. Upstream's requirements.txt warns about this.
  • Pinned torch lines are skipped from requirements.txt, so torch is installed once, from the cu128 index, instead of first pulling a multi-GB PyPI build.
  • cu128 index, because torch 2.10 has no cu130 wheels.
  • Releases are tracked (ShouldIgnoreReleases stays false), unlike the other trainers, because Fizgig publishes real semver releases.
  • Shared folders default to None, like the other trainers. Symlink is opt-in through Shared Model Strategy and junctions only output_loras into the shared Lora folder, so new LoRAs show up in ComfyUI and other packages. Model weights aren't shared, because Fizgig puts every weight type into a single flat models/ directory.

Testing

  • Install tested on Windows 11 with an NVIDIA GPU: the solution builds cleanly, and I installed Fizgig through the SM UI on a dev build. Ran and used it.
  • ROCm is out of scope. Fizgig supports AMD GPUs through ROCm in some form, but I don't have hardware to test it.
  • Currently taking standard Linux support claim at face value from the upstream.

Questions

I'm not sure whether you're accepting PRs for new packages. If not, no worries. I'm happy to adjust anything as well.

Fizgig (shootthesound/Fizgig) is an Apache-2.0 LoRA training studio for
Flux 2 Klein 9B, Krea 2, MiniMax H3 and Qwen Image 2.1. It is a Tkinter
desktop app rather than a web UI, so it follows the same shape as
OneTrainer.

Notes on the less obvious choices:

- LaunchCommand is lora_trainer_gui.py, not launch.pyw. The latter re-spawns
  itself under venv/Scripts/pythonw.exe and exits, which would drop the
  process we track (no console output, no working Stop button) and leave the
  GUI orphaned. lora_trainer_gui.py has a standalone main() and is what
  upstream's run_fizgig.sh invokes directly.

- VcBuildTools is a prerequisite so triton / torch.compile's inductor
  backend can compile kernels on Windows, which Fizgig's Compile Blocks
  speedup needs.

- DISABLE_CUDA=1 is set around the requirements install. hqq ships as an
  sdist whose setup.py kicks off a CUDA kernel build during egg_info without
  it; Fizgig's own requirements.txt warns never to install that line
  otherwise.

- The requirements exclude pattern spells out the version specifiers. The
  pattern is anchored against the whole entry, so the default
  "(torch|torchvision|...)" only catches bare, unpinned names and would let
  Fizgig's pinned torch==2.10.0 through - pulling multi-GB PyPI wheels right
  before the cu128 build force-reinstalls over them.

- Torch pins carry the == operator ("==2.10.0"), since GetTorchPipArgs
  appends the version string directly to the package name.

- CudaIndex is cu128; torch 2.10 has no wheels on SM's default cu130.

Shared folders default to None, matching the other trainers. Symlink is
offered as an opt-in (Package Manager -> Shared Model Strategy) that
junctions output_loras into the shared Lora folder, so a freshly trained
LoRA is immediately visible to ComfyUI and friends. Only output_loras is
mapped: Fizgig flattens every weight it downloads into a single models/
directory, and junctions are directory-level, so DiffusionModels,
TextEncoders and VAE would all collide on the same target path.

ROCm is deliberately left out of this change and will follow separately.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@e-nord
e-nord force-pushed the feat-fizgig-package branch from 64e04e7 to 5e4d417 Compare September 30, 2026 05:38
@NeuralFault

NeuralFault commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

The AMD setup seems pretty standard from the upstream package's documentation.

StabilityMatrix's Windows ROCm helper framework uses the current stable production ROCm 10 repository instead of the old nightlies (no longer getting updated in lieu of new permanent repository)

Plus the custom Win-rocm bitsandbytes wheel from 0xDELUXA.
So pretty much in-line with other packages that have AMD support wired in on Windows for StabilityMatrix.
I can assist in wiring this up if need be, or do it post-merge.

Linux can use upstream pytorch as its rocm 7.14 now. Can be set up with RocmIndex = "rocm7.14" in the PipInstallConfig similar to the ComfyUI package config. Big blocker for it is that the default PyTorch version would have to be set to 2.14 for AMD on Linux for this. So this route may not be the optimal one as Fizgig upstream docs instruct on using AMD's official repository.

The ROCm helper will handle both Windows and Linux soon, so can then wire the package to it and have it pull the Pytorch/ROCm installs from AMD's official repo as per upstream package instructions.

But overall looks to be very doable. Can just be Nvidia/CUDA only in-app at the start and I can work in the AMD support and testing myself afterwards if you're not sure on it yourself.

Copy link
Copy Markdown
Member

Thanks for putting this together! Fizgig looks like a great addition 😁 A couple of installer details to sort out before merging:

  • I’m fine with including VS Build Tools where they help, but can we make it clear before installation that they’re installed system-wide and remain after uninstalling Fizgig?
  • Could we install the intended CUDA Torch build before the remaining requirements, keeping the relevant constraints? Even with the direct Torch requirements filtered out, accelerate can pull it in transitively, which looks like it could cause an extra download/reinstall

Starting with Windows/NVIDIA seems reasonable, with Linux validation and NeuralFault’s offered AMD work tracked separately. Thanks again! :3

e-nord and others added 2 commits October 2, 2026 18:28
…d Tools

Address review feedback on the Fizgig package:

- Install torch 2.10.0 / torchvision 0.25.0 from cu128 via PrePipInstallArgs,
  so it lands before the requirements, and skip the standard torch step.
  accelerate (and others) depend on torch transitively; previously that
  pulled a PyPI torch build during the requirements step which the cu128
  install then had to force-reinstall over.
- Re-state the torch/torchvision pins alongside the requirements. The
  installed +cu128 build satisfies them, and a conflicting requirement
  fails loudly instead of swapping torch.
- Add a Windows-only Disclaimer, shown in the install browser before
  installing, stating that VS Build Tools are installed system-wide and
  remain after Fizgig is uninstalled.

The change is Fizgig-specific; shared install code is untouched.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Fizgig already works on Linux with NVIDIA GPUs; no code change is needed.
The package is CUDA-only and gated on HasNvidiaGpu on every platform, and
the CUDA install carries over unchanged: requirements.txt scopes
triton-windows to win32, pytorch.org publishes Linux cu128 wheels for torch
2.10.0, insightface ships a pure-Python wheel, and the Unix prerequisite
helper skips the Windows-only Tkinter / VS Build Tools prerequisites. The
VS Build Tools disclaimer is already gated on Windows.

Linux ROCm is left out. Upstream labels it highly experimental and installs
it from AMD's TheRock wheels, which do not fit Stability Matrix's existing
Linux ROCm path.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@e-nord

e-nord commented Oct 3, 2026

Copy link
Copy Markdown
Contributor Author

Sure thing, there's now a VS Build tools warning in the pre-defined spot for disclaimers and Torch is installed first via existing pre-pip install args if that suffices.

Also smoke tested an install and started it up in WSL2 if that counts enough for Linux XD

I originally had RoCM in scope and a separate branch for it that 'looks' correct. But since I can't formally test/validate, should I just push it as a draft PR and leave it hanging for NeuralFault to finish out with the Windows and eventual Linux-side RoCM helper support?

@NeuralFault

Copy link
Copy Markdown
Contributor

I originally had RoCM in scope and a separate branch for it that 'looks' correct. But since I can't formally test/validate, should I just push it as a draft PR and leave it hanging for NeuralFault to finish out with the Windows and eventual Linux-side RoCM helper support?

Can just have the CUDA support in place and I can work in the ROCm afterwards post-merge.

@e-nord
e-nord marked this pull request as ready for review October 3, 2026 19:41

@mohnjiles mohnjiles left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the quick turnaround on the last round! I compared the installer against upstream v6.8.3 and found two things worth sorting out, both inline. The rest are small nits, take or leave them :3

new SharedFolderLayoutRule
{
SourceTypes = [SharedFolderType.Lora],
TargetRelativePaths = ["output_loras"],

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Heads up: Fizgig writes more than finished LoRAs into output_loras. Sample images go to output_loras/sample (get_samples_dir() in lora_trainer_gui.py), and with save-state on, the resume state folders (optimizer + weights, often several GB) land there too via save_state_on_epoch_end / save_and_remove_state_stepwise.

With Symlink enabled, all of that ends up in the shared Models/Lora folder, so ComfyUI's LoRA list would show the state-dir weights and the Checkpoint Manager would scan them. Fizgig's own LoRA pickers would also open on the user's whole LoRA library.

I'd lean towards dropping Symlink for now (None only, like the other trainers). Fizgig already has an Output Directory field, so anyone who wants their LoRAs in the shared folder can point it there themselves.

// requirements.txt warns never to install that line without this.
venvRunner.UpdateEnvironmentVariables(env => env.SetItem("DISABLE_CUDA", "1"));

const string torchVersion = "==2.10.0";

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Since releases are tracked and upstream ships often, these hardcoded pins worry me a bit. The exclude pattern strips upstream's own torch== lines, so when upstream bumps torch, an SM update keeps installing 2.10.0 (and re-pins it via ExtraPipArgs) without any error.

The risky part is triton: upstream pins triton-windows>=3.5.1,<3.7 and says to bump it together with torch. After their next torch bump we'd pair a newer triton-windows with the old torch, which is the mismatch their requirements.txt notes as hanging Krea 2 previews inside torch.compile.

Could we read the torch / torchvision lines and the --extra-index-url out of requirements.txt instead (roughly what upstream's uv_install_deps._parse_requirements does), and keep the torch-first ordering? Then the pins follow upstream automatically.

$"torchvision{torchvisionVersion}",
"--extra-index-url",
"https://download.pytorch.org/whl/cu128",
"--force-reinstall",

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: on a fresh install this step runs first, so there's nothing to force over. The flag only kicks in on updates, where it reinstalls torch, torchvision and all the nvidia-* CUDA wheels every time even when 2.10.0+cu128 is already there. Without it, uv keeps a matching build and still replaces a mismatched one.

)
.ConfigureAwait(false);

venvRunner.UpdateEnvironmentVariables(env => env.Remove("DISABLE_CUDA"));

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: this venvRunner is local and disposed by the await using right after, and RunPackage builds a fresh one via SetupVenv, so DISABLE_CUDA can't leak anywhere. This line can just go.

public override string Author => "shootthesound";

public override string Blurb =>
"LoRA training studio for Flux 2 Klein 9B, Krea 2, MiniMax H3 and Qwen Image 2.1 — train, profile, repair and extract";

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: we keep em dashes out of user-facing copy. Maybe ...MiniMax H3 and Qwen Image 2.1. Train, profile, repair and extract?

@@ -0,0 +1,194 @@
using Injectio.Attributes;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: tiny one, the file starts with a UTF-8 BOM. A few older package files have one too, but we're avoiding it in new files.


var packageTitle = args.CurrentPackagePair.BasePackage switch
{
Fizgig => "Fizgig",

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Totally optional and fine as is: this switch just repeats each package's DisplayName, so a _ => args.CurrentPackagePair.BasePackage.DisplayName default would remove the need for this arm (and fix packages like AI-Toolkit showing "Running Stable Diffusion"). Happy to leave that for a follow-up on our side.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants