Shrinking Full-Duplex Speech: Training a 600M Moshi-Style Model on a Weekend Budget edwixx • about 21 hours ago
Compute:Arena / Measuring Local Inference Across Models, Quants, Chips, and Runtimes basecompute • 1 day ago • 5
Summarization Bias: A Pre-Registered Test for a Directional Failure in LLM Judges leventbulut • 1 day ago • 1
Jev AI API Tutorial: Build Your First Structured Decision with Choice, Score, and Noul sora-2 • 2 days ago
Jev AI vs LLMs: When Should You Use a Decision Model Instead of a Chat Model? sora-2 • 2 days ago • 2
Your model already knows it's wrong. Asking costs 0.06 seconds and zero tokens. FINAL-Bench • 2 days ago • 11
Five Raters, One Rule, Five Different Answers: What Happened When We Measured LLM Annotation Agreement leventbulut • 3 days ago • 1