NVIDIA Nemotron 3 Embed Ranks #1 Overall on RTEB, Advancing Agentic Retrieval nvidia • 4 days ago • 50
Kimi K3 Model Overview: 2.8T Parameters, MXFP4 Quantization, and What the Open Weights Mean for the Community ResterChed • 3 days ago • 39
Aether-7B-5Attn: A 100% Open-Source Sovereign Foundation Model — and a Controlled Experiment in Heterogeneous Attention FINAL-Bench • 1 day ago • 17
One Adapter, Both Modalities: Field Notes from Building and Serving a Multimodal Reranker lightonai • 4 days ago • 15
VKUE: No GPU? Runs Anyway — a 34.7B Reasoner on a Laptop and on Bare CPU FINAL-Bench • 8 days ago • 17
makeMoE: Implement a Sparse Mixture of Experts Language Model from Scratch AviSoori1x • May 7, 2024 • 124
How I Trained Action Chunking Transformer (ACT) on SO-101: My Journey, Gotchas, and Lessons sherryxychen • Sep 30, 2025 • 76
NVIDIA Nemotron 3 Embed Ranks #1 Overall on RTEB, Advancing Agentic Retrieval nvidia • 4 days ago • 50
Kimi K3 Model Overview: 2.8T Parameters, MXFP4 Quantization, and What the Open Weights Mean for the Community ResterChed • 3 days ago • 39
Aether-7B-5Attn: A 100% Open-Source Sovereign Foundation Model — and a Controlled Experiment in Heterogeneous Attention FINAL-Bench • 1 day ago • 17
One Adapter, Both Modalities: Field Notes from Building and Serving a Multimodal Reranker lightonai • 4 days ago • 15
VKUE: No GPU? Runs Anyway — a 34.7B Reasoner on a Laptop and on Bare CPU FINAL-Bench • 8 days ago • 17
makeMoE: Implement a Sparse Mixture of Experts Language Model from Scratch AviSoori1x • May 7, 2024 • 124
How I Trained Action Chunking Transformer (ACT) on SO-101: My Journey, Gotchas, and Lessons sherryxychen • Sep 30, 2025 • 76