NCN-RW
NCN-RW is a 9.72M-parameter recurrent Transformer with a recurrent neuromodulatory controller, trained from scratch on approximately 10 billion tokens. It is a base language model for text completion and continuation scoring.
Its RNN–Transformer hybrid architecture brings together a Transformer trunk, GRU memory, and MLP control and prediction pathways. Five dense experts propose changes to a shared representation, while the controller adjusts how attention, feed-forward networks, and repeated passes process that representation.
Model family
NCN-RW is one of the first members of AxiomicLabs’ Neuromodulatory Control Networks (NCN) family. This expanding family explores biomimicry and experimental brain-inspired architectures, investigating how recurrent memory, predictive feedback, and neuromodulation can shape language-model computation.
Architecture
| Property | Value |
|---|---|
| Unique trainable parameters | 9,720,527 |
| Hidden size | 256 |
| Vocabulary | 4,096-token byte-level BPE |
| Context | 512 tokens |
| Shared core | Six blocks, applied in two passes |
| Additional blocks | One prelude and one coda |
| Main block applications | 14 |
| Attention | Causal grouped-query attention; 8 query heads, 2 KV heads; QK normalization and RoPE |
| Main feed-forward network | SwiGLU, width 896 |
| Experts | Five dense SwiGLU feature adapters, width 256 |
| Neuromodulatory controller width | 64 |
| Dropout | 0 |
Three neural-network components working together
| Component | Role in the model |
|---|---|
| Transformer | Builds the main 256-dimensional language representation using causal attention and SwiGLU FFNs. A prelude produces an anchor, six core blocks form the shared trunk, and a coda produces the final representation. |
| GRU | Maintains a 64-dimensional tonic state across tokens. A separate GRUCell updates controller state across computational depth as the trunk and experts change the representation. |
| MLP pathways | Combine tonic memory with phasic changes, decode neuromodulatory controls, predict internal representation changes, and transform controller state into expert-routing signals. SwiGLU MLPs also implement the expert proposals and shared refiner. |
These components form a feedback loop. The Transformer changes the language representation; an MLP predicts that change; the discrepancy and a compressed observation enter the depth GRUCell; and the resulting controller state changes the controls applied to later computation. Controller feedback uses already-computed causal states. Future tokens are used as training targets for Compass, without entering the inference control path.
A trunk that revisits its own workspace
The main path is:
Token embeddings
→ Prelude and anchor
→ Six-block Transformer trunk
→ Five expert proposals → Weighted consensus → Shared SwiGLU refiner
→ Refresh toward anchor
→ The same six-block trunk, with updated controller state
→ The same experts and refiner, with new proposals and votes
→ Coda → Tied vocabulary readout
The default two passes apply the six core blocks twice, giving 14 main block applications with the prelude and coda. Weight sharing makes additional depth available without duplicating the trunk's parameters. Before the second pass, a learned, per-feature refresh gate blends the current workspace with the prelude anchor. A learned pass embedding distinguishes the rounds for the controller. The second pass reuses the weights with a changed representation and controller state.
There are three forms of recurrence: token recurrence in the tonic GRU, depth recurrence in the feedback GRUCell, and semantic recurrence through the repeated Transformer trunk. Each pass has its own attention-cache slots, even though its Transformer weights are shared.
Neuromodulation throughout the computation
The controller's initial state combines the tonic GRU output, a phasic projection of changes between neighboring token embeddings, and their multiplicative interaction. Learned control heads turn that state into bounded contextual adjustments around learned operating points:
| System | What it changes |
|---|---|
| Attention precision | Per-head scaling of normalized queries, controlling the sharpness of attention scores. |
| Presynaptic gain | Per-head attention output strength before the attention output projection. |
| Postsynaptic gain | Per-feature strength after the attention output projection. |
| Global broadcast and local receptors | A shared 256-dimensional modulation signal acts on each block's SwiGLU gate input through learned, layer-specific receptor sensitivities. |
| FFN output gain | A separate scalar controls the magnitude of each block's FFN contribution. |
| Workspace refresh | Per-feature retention balances the evolving workspace against the prelude anchor before another pass. |
| Workspace update | A scalar gate controls how much expert consensus and refinement enter the residual stream. |
Presynaptic and postsynaptic gains act on opposite sides of a learned projection, so they control different operations. The global signal also has a separate role from FFN output gain: it changes the gate's input, while output gain changes the resulting update's magnitude. The controller therefore influences how features are computed and how strongly they are retained throughout the trunk.
Compass adds three learned forecasts at horizons of 8, 32, and 64 tokens. Their training targets are averages of future token embeddings from an EMA teacher, with masks that respect document boundaries. At inference, predicted Compass features join the current controller state in the routing query. The teacher is a training-only embedding buffer with 1,048,576 values, excluded from the trainable parameter count.
Dense experts and layered SwiGLU refinement
The FFNs have three distinct jobs and widths: 896 in the main Transformer blocks, 256 in each feature expert, and 512 in the shared refiner. Each uses a multiplicative SwiGLU gate/value interaction, with different placement in the computation.
Inside the trunk, the controller's global broadcast modulates the gate input while the value branch receives the normalized language representation. Each of the six core layers has its own gate, value, and down-projection matrices; the entire block, including those matrices, is reused across passes. This is sharing through repeated core computation.
After each trunk pass, all five experts process the same normalized workspace and propose feature updates in a common coordinate system. Each expert contains 196,608 parameters. Their proposals receive a smooth RMS cap, and the controller assigns token-dependent votes using normalized routing queries and expert keys, a learned temperature, and expert biases. Every expert contributes through dense voting; the routing floor reserves 0.001 total weight across the five experts.
The weighted consensus is added to the workspace before a shared nonlinear SwiGLU refiner processes it. The controller then gates the combined consensus and refinement into the residual stream:
consensus = sum(expert_vote × expert_proposal)
refinement = SwiGLU_refiner(normalize(workspace + consensus))
new_workspace = workspace + update_gate × (consensus + refinement)
This sequence lets the experts influence intermediate features that the next trunk pass can use. The final token prediction comes from one shared hidden state and one tied vocabulary head. The unusual feature is this coordinated loop: recurrent memory, predictive feedback, attention control, modulated FFNs, expert consensus, and repeated Transformer computation all work on the same evolving workspace.
Training and data
The training corpus used this data mix for the entirety of the model's training. TinyStories was included in an attempt to help beginner grammar, but whether or not this was successful is unknown.
| Source | Share |
|---|---|
| FineWeb-Edu | 35% |
| Cosmopedia | 25% |
| DCLM | 22% |
| FineMath 4+ | 14% |
| TinyStories | 4% |
The selected checkpoint is step 38,000, with 9,961,472,000 training tokens consumed, as recorded in the checkpoint's token cursor. This is the training total for the selected model.
Training used AdamW with BF16 CUDA autocast and a 512-token context. A microbatch of 128 sequences and four accumulation steps gave 262,144 tokens per optimizer update. Peak learning rates were 0.003 for the language-model components and 0.001 for the controller and router, with betas (0.9, 0.95), weight decay 0.01 on eligible weights, gradient clipping at 1.0, and seed 111. Warmup ended at step 500; the selected checkpoint is in the stable phase. A controller-recovery continuation began at step 4,020.
At step 38,000, validation loss was 2.780727 and perplexity was 16.130741, measured on a fixed 131,072-token prefix of the recorded validation binary.
Zero-shot evaluation
The checkpoint was selected because it achieved the highest Intelligence Index among the completed evaluations in this run. Public benchmark results informed this selection.
Author-run results for the selected step-38,000 checkpoint are shown below. Accuracy is reported in percent. HellaSwag, ARC-Easy, ARC-Challenge, and PIQA use character-normalized continuation likelihood; ArithMark-3 uses token-normalized continuation likelihood.
| Benchmark | Score |
|---|---|
| HellaSwag | 27.73 |
| ARC-Easy | 36.95 |
| ARC-Challenge | 22.70 |
| PIQA | 58.00 |
| ArithMark-3 | 30.20 |
| Chance-normalized Intelligence Index | 8.378012 |
These are recorded local evaluator results. They have not been independently verified against an external harness. Zero-shot describes the evaluation prompts; a complete training-data overlap audit has not been established. Dataset fingerprints and original per-example candidate scores are retained in the run artifacts.
Use
NCN-RW uses custom PyTorch code. Run this example from a downloaded model repository containing model.py, inference_state.py, config.json, tokenizer.json, and model.safetensors:
import json
from contextlib import nullcontext
import torch
from safetensors.torch import load_file
from tokenizers import Tokenizer
from model import RecurrentWorkspace, WorkspaceConfig
device = "cuda" if torch.cuda.is_available() else "cpu"
with open("config.json", encoding="utf-8") as f:
model = RecurrentWorkspace(WorkspaceConfig(**json.load(f)))
result = model.load_state_dict(load_file("model.safetensors"), strict=False)
if result.missing_keys != ["ema_embedding"] or result.unexpected_keys:
raise RuntimeError(f"Unexpected checkpoint mismatch: {result}")
model = model.to(device).eval()
tokenizer = Tokenizer.from_file("tokenizer.json")
ids = torch.tensor(
[tokenizer.encode("The color of the sky is").ids], device=device
)
autocast = (
torch.autocast("cuda", dtype=torch.bfloat16)
if device == "cuda" else nullcontext()
)
with torch.inference_mode(), autocast:
for _ in range(min(32, model.cfg.context - ids.shape[1])):
token = model(ids, auxiliary=False)["logits"][:, -1].argmax(-1)
ids = torch.cat([ids, token[:, None]], dim=1)
if token.item() == model.cfg.eos_id:
break
print(tokenizer.decode(ids[0].tolist(), skip_special_tokens=True))
The inference checkpoint omits the training-only EMA teacher buffer. It contains all 9,720,527 learned parameters and the required prediction-scale buffer, with input/output embedding weights tied by the architecture.
Runtime dependencies are PyTorch >=2.5, Tokenizers >=0.20, and Safetensors >=0.4.5. Stored weights are FP32; BF16 CUDA autocast matches the recorded evaluation precision. The tokenizer's EOS token is <|endoftext|> (ID 0).
The model also supports use_cache, cache, and cache.select for independent answer branches. Preserve question-to-answer recurrent state, use unpadded inputs, and batch equal token counts. The context limit is 512 tokens. config.json contains WorkspaceConfig arguments; the supplied architecture is not a Transformers AutoModel configuration.
Limitations
NCN-RW is a small base model and can produce incorrect or incoherent completions. The documented training recipe does not include instruction tuning or safety alignment. Evaluation covers the five tasks above, and benchmark-based checkpoint selection can make these scores optimistic estimates of generalization. The two-pass configuration is the evaluated setting; additional passes and longer contexts are not established by these results.
As far as throughput is concerned, no attempt has been made to optimize NCN-RW. It is, at the moment, meant purely as a research model.
Creator and license
Creator: Mmorgan-ML.
The model package is released under the Apache License 2.0. Upstream dataset licenses apply to their respective data.
- Downloads last month
- 205
.png)
