238 MB
80,108 files
Updated about 10 hours ago
Name
Size
MNIST
mnist_fm
mnist_fm_eval
README.md5.03 kB
xet
run_repro.py20.8 kB
xet
README.md

Linearizer Reproduction — arXiv:2510.08570 ("Who Said Neural Networks Aren't Linear?")

Reduced-scale reproduction of the core claims of the Linearizer paper (Berman, Hallak, Shocher), scope chosen by the user: one-step flow matching on MNIST 32×32. Grounded in the official repo assafshocher/Linearizer @ commit adfca2c, imported verbatim; training/eval logic adapted from the repo's one_step/train_one_step.py / test_one_step.py by the driver run_repro.py.

Configuration (reduced scale)

Item Setting
Model Official Linearizer: InverseUnet g (4 invertible blocks: ActNorm → 2 affine couplings w/ Song-UNet conditioners, model_channels=16 → invertible 1×1 conv), time-conditioned via gx(t=1)/gy(t=0) modes; rank-8 time-dependent LoRA core over dim 1024 — 9,932,800 params
Dataset MNIST train split, 60k images, resized 32×32, [0,1]
Loss Official FM loss: induced-space MSE + noisy reconstruction terms (MSE on x0, LPIPS on x1/x1_pred), noise_level 0.3
Optimizer Adam lr 1e-4, batch 128 (official 256 does not fit 24 GB), 30 epochs (~14,070 steps) vs official 201
Hardware 1× A10G (a10g-small), 0.82 it/s, 4.8 h training

Results

All metrics computed by mnist_fm_eval/results.json on the 30-epoch checkpoint lin_final.pth.

Claim Paper This reproduction Verdict
Linearity under induced ⊕/⊙ ops (Lemmas 1–2) exact by construction f(x1⊕x2) vs f(x1)⊕f(x2): rel. err 2–7×10⁻⁶ (fp32 precision) across 8 addition + 3 scaling checks ✅ CONFIRMED
One-step ≡ multi-step sampling (Euler, T=100) collapse to a single operator B image MSE 4.9×10⁻¹⁰ (PSNR 93.1 dB); latent collapse rel. err 6.6×10⁻⁶; FID between the two sample sets 0.77 ✅ CONFIRMED (exact)
One-step ≡ multi-step (RK, T=100) paper reports 100-vs-1 fidelity PSNR 32.4 / MSE 3.0×10⁻⁴ PSNR 32.0 dB / MSE 6.3×10⁻⁴ ✅ matches paper — the official RK matrix-B is a cruder discretization than the iterative RK sampler (also flagged by third-party reproducers); Euler collapse is exact, RK collapse is approximate
Inversion via Moore–Penrose pseudoinverse (B⁻¹) exact encoding (Lemma 7) roundtrip x1→x0→x̂1: PSNR 76.4 dB, LPIPS 1.2×10⁻⁷ ✅ CONFIRMED
g invertibility required for induced spaces roundtrip max abs err 3.1×10⁻⁴ per pixel ✅ holds to training precision
Generation quality MNIST FID at 201 epochs, multi-step ~127 (CelebA numbers in main table) FID 43.1 (one-step, T=100, 10k samples vs 10k real test, pytorch-fid Inception) partial: generated digits are recognizable (see mnist_fm_eval/samples_one_step.png), but 30-epoch training is far short of the paper's 201 — absolute FID not comparable

Honest deviations / findings

  • T=1000 vs T=100 sampling: one-step T=1000 gives FID 52.1 vs 43.1 at T=100 — the opposite direction from the paper's reported ~8-point improvement (1000→1 vs 100→1). Since the Euler collapse itself is exact (FID one-vs-multi = 0.77), this reflects the sampler trajectory of a model under-trained at 30 epochs, not a failure of the collapse.
  • The RK one-step/multi-step gap (PSNR 32) is an implementation property of the official code (its matrix-B RK omits the nested terms of Appendix F), not an error in our reproduction — and it exactly reproduces the paper's own reported one-vs-100 fidelity numbers.
  • Training used batch 128 instead of 256 (official config OOMs on 24 GB cards) and 30/201 epochs (budget).

Artifacts

Rerun

uv run run_repro.py --epochs 30 --batch-size 128 --fid-samples 10000 \
  --trackio-space <space_id> --out <out_dir>

Training dashboard: https://huggingface.co/spaces/abidlabs/linearizer-mnist-fm-trackio

Provenance

Trained on Hugging Face Jobs (a10g-small) with trackio logging; driver imports the official implementation verbatim from codeload tarball of commit adfca2c. arxiv:2510.08570

Total size
238 MB
Files
80,108
Last updated
Oct 7
Pre-warmed CDN
US EU US EU

Contributors