LuminaV Optimizer

We Were Too Broke for AdamW So We Trapped Gradients in a Hyperbolic Straitjacket and Hired a Traffic Cop to Slap Them

Official Upstream & Standalone Codebase | Current Version: v1.3.3 | Check `Files and Versions`

PyPI Version Changelog Benchmark Software DOI Paper DOI Config JSON License

Official Research Paper

LuminaV Paper Preview

LuminaV Optimizer Theory & Mechanics
Read LuminaV.pdf (Local Mirror)  |  Primary Paper Archive

Click the preview above to read or download the official paper PDF.

Honestly, this is actually the PDF for the first version of LuminaV. So, the major updates we've made since then aren't in here. I might make a new one later. So, for now.. please look => CHANGELOG.md


Notice: Official Upstream Repository

This repository (cloverx-id/LuminaV-Optimizer-Paper) is the official standalone and living development repository for the LuminaV optimizer family.

While LuminaV was originally conceived and validated as the core engine for the XoneLM-1.0 language model series, all subsequent optimizer upgrades, low-precision Triton kernels, PyTorch standards compliance, and bug fixes are actively maintained and released directly in this repository.


What's New in v1.3.3 & v1.3.2

The v1.3.3 and v1.3.2 releases bring substantial compiler compatibility enhancements, distributed pretraining synchronization, zero-allocation memory stabilization, and mathematical precision refinements:

Version 1.3.3 (Compiler Stabilization & Clean Telemetry)

  • Triton AST Compiler Inlining Resolution: Refactored control flow inside @triton.jit helper functions (_store_param and _store_state_buffer) into unified, single-exit branches, completely resolving fatal MLIR/CFG early-return compiler errors during JIT inlining across varying Triton releases.
  • Type-Safe 32-Bit Stochastic Rounding Arithmetic: Enforced strict signed 32-bit integer arithmetic via two's complement constants (-1640531527, -2048173461, -1028477387) and standard bitwise masks (-65536 for BF16, -8192 for FP16), preventing silent 64-bit promotion and ensuring valid bitcast width matching (tl.int32 to tl.float32).
  • Standardized Positional JIT Call Arguments: Converted internal boolean keyword parameters (use_sr=False) into positional parameters across Triton kernel calls to eliminate runtime binding discrepancies.
  • Console De-Noising & Fallback Telemetry: Integrated throttled one-time warning tracking (_warned_triton_failure) with full exception tracebacks routed cleanly to DEBUG, preventing multi-line compiler AST dumps from flooding training consoles.
  • Validated JIT Execution on NVIDIA Architectures: Formally confirmed end-to-end execution on NVIDIA Tesla T4 with zero JIT warnings in both Dual-Pass (fused_single_pass=False) and Fused Single-Pass (fused_single_pass=True) modes in native FP16.

Version 1.3.2 (Distributed Synchronization & Zero-Allocation Caching)

  • Distributed FSDP Collective Synchronization (fsdp_sync): Added the fsdp_sync: bool = False flag with automated runtime guarding (torch.distributed.is_initialized()). Uses cross-GPU collective communications (all_reduce) on global reduction metrics (mask_sum, u_sq_sum, and p_sq_sum), ensuring exact scalar consistency for m_bar, cautious_scale, and bound_scale across distributed parameter shards under PyTorch FSDP and DeepSpeed ZeRO-3.
  • Persistent Contiguous Buffer Caching (Zero-Allocation Execution): Added persistent state caching tensors (p_contig, exp_avg_contig, exp_avg_sq_contig, mp_contig) inside self.state[p], eliminating dynamic memory allocations during iterative training loops, stabilizing the PyTorch VRAM caching allocator, and preserving CUDA Graph Capture address invariance.
  • Dedicated Low-Precision Stochastic Quantization (_sr_quantize): Factored out bitwise stochastic rounding into a modular routine applied directly to momentum states (exp_avg in BF16/FP16) across all PyTorch CPU and CUDA fallback loops, eliminating gradient stagnation caused by LSB truncation.
  • Zero-Copy Direct Tensor Aliasing for FP32 Master Weights: For layers natively in torch.float32, state["master_param"] now directly aliases the parameter tensor pointer (state["master_param"] = p), eliminating auxiliary master weight memory overhead (0 bytes auxiliary footprint).
  • Automatic Checkpoint Sanitizer (state_dict): Overrode state_dict() to automatically purge temporary staging buffers, ensuring on-disk model checkpoints remain compact, clean, and fully backward-compatible.
  • Defensive Hardware Pipeline Flush on Fault (torch.cuda.synchronize): Injected device synchronization within Triton exception recovery handlers prior to PyTorch fallback execution, preventing asynchronous CUDA errors from cascading.
  • High-Entropy Bit-Avalanche PRNG Hashing: Replaced legacy LCG hashing with SplitMix32 / MurmurHash3 integer bit-avalanche hashing in _prng to eliminate low-order bit periodicity in stochastic rounding.
  • Singular Zero-Division & Step-0 Bias Guard: Introduced safe step flooring (safe_base_step = max(base_step, 1) and bc2_base = max(1.0 - (beta2 ** safe_base_step), 1e-15)), eliminating potential zero-division at step 0.

(For the complete patch notes and historical version logs, see CHANGELOG.md.)


Overview

LuminaV is a high-performance adaptive optimizer engineered specifically for deep learning workloads running directly in low precision (FP16 / BF16) without maintaining redundant 4-byte FP32 master weights.

By combining Centered Innovation Variance, Hyperbolic Tangent (tanh) Coordinate Bounding, a Directional Traffic-Cop Mask, and On-Chip Bitwise Stochastic Rounding, LuminaV eliminates the standard 16-byte-per-parameter memory tax imposed by classic optimizers while avoiding weight freezing, gradient shocks, and numerical underflow.


Key Features

  1. Zero Master-Weight Copies (Default): Directly mutates parameter weights in native FP16 or BF16, eliminating the 4-byte FP32 master weight allocation.
  2. On-Chip Bitwise Stochastic Rounding (SR): Implements in-register bitcast hashing in Triton to provide mathematically unbiased stochastic rounding, preventing weight stagnation during fine-grained updates or learning rate decay.
  3. Hyperbolic tanh Bounding Envelope: Maps normalized momentum through a (-1.0, 1.0) transfer function, guaranteeing coordinate updates cannot explode beyond the step learning rate.
  4. The Traffic-Cop Directional Gate: Dynamically eliminates coordinate updates whenever historical momentum conflicts with the incoming mini-batch gradient direction (u_t Β· g_t <= 0).
  5. Centered Innovation Variance: Tracks centered innovation dispersion (g_t - m_t)^2 rather than uncentered raw second moments, suppressing variance inflation during confident descent.
  6. Automatic FP16 Cliff Governor: Built-in asymptotic boundary governor that dampens steps near the IEEE-754 FP16 overflow limit (> 65,504), enabling stable pure FP16 training without external schedulers or clipping.
  7. Direction-Preserving Radial Bounding: Smooth asymptotic parameter squashing (tanh(r)/r) that preserves 100.000% gradient angular fidelity while capping displacement.
  8. Multi-Tier Master Weight Support: Configurable on the fly from 100% master-free up to hybrid ("semi") or full FP32 ("full") modes.
  9. Dual Execution Engine: Fully accelerated custom OpenAI Triton kernels for CUDA and Intel XPU devices, paired with vectorized C++ torch._foreach multi-tensor fallbacks.

Installation

From PyPI (Recommended)

pip install luminav

For GPU acceleration via OpenAI Triton:

pip install luminav[triton]

From Source (Editable Mode)

git clone https://huggingface.co/cloverx-id/LuminaV-Optimizer-Paper
cd LuminaV-Optimizer-Paper
pip install -e .

Direct File Drop-in

Alternatively, you can copy luminav.py directly into your working project directory without packaging overhead:

wget https://huggingface.co/cloverx-id/LuminaV-Optimizer-Paper/raw/main/luminav.py

Quickstart

Standard Instantiation (Master-Free Mode)

import torch
from luminav import LuminaV

# Instantiate your model in native low precision (e.g. BF16 or FP16)
model = YourModel().to(device="cuda", dtype=torch.bfloat16)

# Initialize LuminaV v1.3.3
optimizer = LuminaV(
    model.parameters(),
    lr=8e-4,                    # or 8e-5 / 8e-6 for fine-tuning
    betas=(0.9, 0.999),
    eps=1e-8,
    weight_decay=0.08,
    tau=0.8,
    alpha_ss=0.5,
    cautious=True,
    cautious_clamp_min=0.5,     # Exact power-of-two ceiling (2.00x)
    buffer=2,                   # 2 = Dual-Buffer (Standard), 1 = Single-Buffer (Extreme Low VRAM)
    stochastic_rounding=True,
    bound=True,                 # Smooth asymptotic step bounding
    bound_type="radial",        # "radial" (preserves 100% angular direction) or "coordinate"
    bound_ratio=0.03,
    master_weights="none",      # "none" (Master-Free), "semi" (Hybrid FP32 Master), or "full" (Full FP32)
    fused_single_pass=False,    # True enables ultra-fast 1-pass EMA kernel
    fsdp_sync=False,            # True enables collective metric sync across distributed shards
    execution="auto"
)

# Standard training step
optimizer.zero_grad(set_to_none=True)
loss = model(inputs, targets)
loss.backward()
optimizer.step()

Ultra-Low Memory Training (Single-Buffer Mode)

To cut optimizer memory state by an additional 50% (maintaining only a single momentum buffer and collapsing variance to scalar RMS):

optimizer = LuminaV(
    model.parameters(),
    lr=8e-4,
    buffer=1,                   # LuminaV-1 Single-Buffer Mode
    alpha_ss=0.5,               # Softsign presquashing factor
    master_weights="none"
)

Loading from config.json

import json
import torch
from luminav import LuminaV

with open("config.json", "r") as f:
    config = json.load(f)

# Initialize with verified default configuration
optimizer = LuminaV(model.parameters(), **config["default_params"])

Parameter Reference

Parameter Type Default Description
params iterable Required Iterable of parameters to optimize or dicts defining parameter groups.
lr float 8e-4 Learning rate (Ξ·).
betas Tuple[float, float] (0.9, 0.999) Coefficients (β₁, Ξ²β‚‚) for running momentum and centered innovation variance. First-moment bias correction uses 1.0 - beta1^(step + 1).
eps float 1e-8 Numerical stability term (Ξ΅). Automatically floored to 1e-4 in FP16 to prevent subnormal underflow.
weight_decay float 8e-2 Decoupled weight decay coefficient (Ξ»).
tau float 0.8 Analytical bias correction temperature parameter (Ο„).
alpha_ss float 0.5 Softsign dampening factor (Ξ±_ss) used in single-buffer mode (buffer=1).
cautious bool True If True, enables Traffic-Cop directional verification masking.
cautious_clamp_min float 0.5 Safety floor density clamp (Ξ³_min) enforcing a power-of-two maximum energy scaling ceiling (2.00x, 2ΒΉ) and preventing division by zero.
buffer int 2 Buffer mode: 2 (Dual-buffer tracking m_t and v_t) or 1 (Single-buffer scalar RMS tracking).
stochastic_rounding bool True Enables bitwise stochastic rounding on native FP16/BF16 weights.
bound bool True If True, enables smooth asymptotic parameter bounding to prevent divergence in deep networks.
bound_type str "radial" Asymptotic bounding formulation: "radial" (direction-preserving squashing using tanh(r)/r) or "coordinate" (elementwise squashing).
bound_ratio float 0.03 Maximum allowed step displacement ratio relative to parameter norm or magnitude (R = bound_ratio * max(β€–pβ€–, 1.0)).
master_weights Union[bool, str] "none" Master weight precision mode: "none" / False (Master-Free), "semi" / "half" (FP32 master with 16-bit states), or "full" / "fp32" (Full FP32).
fused_single_pass bool False If True (step > 1), fuses pass 1 and pass 2 into a single unified GPU kernel using running EMA scale estimates.
ema_decay float 0.8 Running scale decay factor (Ξ±_ema) used when fused_single_pass=True.
fsdp_sync bool False If True, enables cross-GPU collective all_reduce synchronization of reduction metrics across sharded parameter ranks under PyTorch FSDP / DeepSpeed ZeRO-3.
execution str "auto" Execution engine: "auto", "triton", "foreach", or "single". Automatically routes to "foreach" if deterministic mode is enabled.
seed int 1337 Base seed for PRNG stochastic rounding and stateless golden-ratio hashing.

Operational Modes

LuminaV-2 (Dual-Buffer Default: buffer=2)

Maintains first moment m_t and centered innovation variance v_t:

mt=Ξ²1mtβˆ’1+(1βˆ’Ξ²1)gt m_t = \beta_1 m_{t-1} + (1 - \beta_1) g_t

vt=Ξ²2vtβˆ’1+(1βˆ’Ξ²2)(gtβˆ’mt)2 v_t = \beta_2 v_{t-1} + (1 - \beta_2)(g_t - m_t)^2

Updates are bounded through the hyperbolic tangent envelope:

ut=tanh⁑(m~tΟƒt),m~t=Ξ²1mt+(1βˆ’Ξ²1)gt u_t = \tanh\left(\frac{\tilde{m}_t}{\sigma_t}\right), \quad \tilde{m}_t = \beta_1 m_t + (1 - \beta_1) g_t

LuminaV-1 (Single-Buffer Extreme-Poverty Mode: buffer=1)

Collapses variance tracking into a scalar Root-Mean-Square (RMS) across the entire tensor, saving 50% optimizer state memory by maintaining only a single state buffer (m_t):

RMS(m~t)=1Nβˆ‘i=1Nm~t,i2+Ο΅ \text{RMS}(\tilde{m}_t) = \sqrt{\frac{1}{N} \sum_{i=1}^N \tilde{m}_{t,i}^2 + \epsilon}

ut=tanh⁑(z1+Ξ±ss∣z∣),z=m~tΟ„β‹…RMS(m~t)+Ο΅(1βˆ’Ξ²1t+1)Ο„ u_t = \tanh\left(\frac{z}{1 + \alpha_{ss}|z|}\right), \quad z = \frac{\tilde{m}_t}{\tau \cdot \text{RMS}(\tilde{m}_t) + \epsilon(1 - \beta_1^{t+1})\tau}


Empirical Benchmarks (Qwen3.5-4B-Base)

LuminaV v1.3.0 was rigorously benchmarked on a 4.0-billion parameter Large Language Model (Qwen/Qwen3.5-4B-Base) initialized from architectural config (AutoModelForCausalLM.from_config) under full pretraining from scratch conditions (all 4B parameters actively optimized, gradient checkpointing disabled) on an NVIDIA 80GB GPU.

LuminaV Qwen3.5-4B Benchmark Charts

Architectural Disambiguation (1P / 2P vs Buffer Count):
The labels 1P and 2P refer strictly to GPU Kernel Execution Passes (1P = Fused Single-Pass Kernel via EMA, 2P = Standard Two-Pass Kernel Reduction). All evaluated configurations below strictly operate under the Dual-Buffer architecture (buffer=2), maintaining both the first moment (m_t) and centered innovation variance (v_t). Single-buffer mode (buffer=1) was not benchmarked in this suite.

Key Benchmark Highlights

Optimizer Mode Kernel Passes Buffer Mode Peak VRAM Static State VRAM Pure Opt Latency Final Loss (250) Weight Cosine Sim (vs 2P)
Master-Free (1P) 1-Pass (EMA) buffer=2 (Dual) 34.32 GB (-47.6%) 24.65 GB (-55.8%) 59.9 ms (1.90x faster) 8.8260 0.9445
Master-Free (2P) 2-Pass (Exact) buffer=2 (Dual) 34.32 GB (-47.6%) 24.65 GB (-55.8%) 68.4 ms 9.1075 1.0000 (Exact)
Semi (1P) 1-Pass (EMA) buffer=2 (Dual) 49.99 GB (-23.6%) 40.32 GB (-27.7%) 70.2 ms (1.62x faster) 8.8192 (Best) 0.9512
Full (2P - Baseline) 2-Pass (Exact) buffer=2 (Dual) 65.47 GB 55.80 GB 113.8 ms 8.9579 1.0000 (Exact)
  • 47.6% Peak VRAM Reduction (31.15 GB Saved): Master-Free mode cuts peak memory from 65.47 GB down to 34.32 GB, allowing full-throughput training of a 4B parameter model on 40GB/48GB GPUs without activation checkpointing.
  • 1.9x Pure Optimizer Speedup via Fused Single-Pass: Slashing redundant VRAM memory passes cuts optimizer latency from 113.8 ms (Full 2P) down to 59.9 ms (Master-Free 1P) while retaining full dual-buffer state tracking (buffer=2).
  • 94.0% Mean Cosine Representation Fidelity: Single-pass running EMA updates preserve 94% angular cosine similarity with exact two-pass trajectories while achieving equal or superior convergence.

Read the Full Empirical Benchmark & Reproducibility Report (BENCHMARKS.md)
Open Interactive Reproduction Notebook in Google Colab


Citation

If you utilize LuminaV in your research or applications, please cite both the foundational paper and this software implementation:

# 1. To cite the official research paper & theoretical mechanics
@misc{luminamoon2026luminav_paper,
  author       = {{Silver Moon (cloverxion)}},
  organization = {Lumina Moon (cloverx-id)},
  title        = {{LuminaV: We Were Too Broke for AdamW So We Trapped Gradients in a Hyperbolic Straitjacket and Hired a Traffic Cop to Slap Them}},
  year         = {2026},
  publisher    = {Hugging Face},
  doi          = {10.57967/hf/10270},
  url          = {https://huggingface.co/cloverx-id/XoneLM-1.0-Paper}
}

# 2. To cite this software implementation & standalone codebase
@software{luminamoon2026luminav_code,
  author       = {{Silver Moon (cloverxion) and Lumina Moon Contributors}},
  organization = {Lumina Moon (cloverx-id)},
  title        = {{LuminaV Optimizer: Official PyTorch Implementation}},
  year         = {2026},
  publisher    = {Hugging Face / PyPI},
  version      = {1.3.3},
  doi          = {10.57967/hf/10365},
  url          = {https://huggingface.co/cloverx-id/LuminaV-Optimizer-Paper}
}

License

All Resources are under Apache License 2.0. See LICENSE for full terms.

Downloads last month
659
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support