Proto_AGI's picture

Proto_AGI

mayafree

AI & ML interests

None yet

Recent Activity

reacted to SeaWolf-AI's post with šŸ˜Ž about 10 hours ago
🧠 We just released Darwin-27B-ZTC, a judgment engine that reaches a verdict without generating anything. Most LLMs answer by generating, decoding one token at a time. Darwin-27B-ZTC takes a different route. āš™ļø How it works šŸ”¹ It makes its call in a single forward pass. šŸ”¹ Zero generated tokens, and no decoding loop. šŸ”¹ That keeps latency and cost far below what a generative model needs. šŸŽÆ What it judges šŸ”¹ It handles several question types: free-form correctness (noul), multiple choice (choice), and scoring (score). šŸ”¹ For each one it hands back a calibrated confidence, not just an answer. šŸ“Š How well calibrated (measured) šŸ”¹ KL 0.204, Brier 0.097, so the confidence it reports lines up with what actually happens. šŸ”¹ 0.743 accuracy (zero-shot, general split), across 2,000 judgments with zero errors. šŸ”¹ By type: noul 0.847, choice 0.723, score 0.675. šŸ”¹ None of the benchmark's train split went into it. It is pure zero-shot. šŸš€ Where it fits šŸ”¹ Grading at scale, model routing, safety gating, anywhere you want a fast decision without paying for generation. šŸ† It currently sits at #1 on the official typed-decisions leaderboard on Hugging Face (0.743 accuracy, zero-shot). šŸ”— Links Model: https://huggingface.co/FINAL-Bench/Darwin-27B-ZTC Leaderboard: https://huggingface.co/datasets/LocalLLaMA/typed-decisions Curious to hear what you make of the single-pass, no-generation approach. šŸ™Œ
reacted to SeaWolf-AI's post with šŸ‘€ about 10 hours ago
🧠 We just released Darwin-27B-ZTC, a judgment engine that reaches a verdict without generating anything. Most LLMs answer by generating, decoding one token at a time. Darwin-27B-ZTC takes a different route. āš™ļø How it works šŸ”¹ It makes its call in a single forward pass. šŸ”¹ Zero generated tokens, and no decoding loop. šŸ”¹ That keeps latency and cost far below what a generative model needs. šŸŽÆ What it judges šŸ”¹ It handles several question types: free-form correctness (noul), multiple choice (choice), and scoring (score). šŸ”¹ For each one it hands back a calibrated confidence, not just an answer. šŸ“Š How well calibrated (measured) šŸ”¹ KL 0.204, Brier 0.097, so the confidence it reports lines up with what actually happens. šŸ”¹ 0.743 accuracy (zero-shot, general split), across 2,000 judgments with zero errors. šŸ”¹ By type: noul 0.847, choice 0.723, score 0.675. šŸ”¹ None of the benchmark's train split went into it. It is pure zero-shot. šŸš€ Where it fits šŸ”¹ Grading at scale, model routing, safety gating, anywhere you want a fast decision without paying for generation. šŸ† It currently sits at #1 on the official typed-decisions leaderboard on Hugging Face (0.743 accuracy, zero-shot). šŸ”— Links Model: https://huggingface.co/FINAL-Bench/Darwin-27B-ZTC Leaderboard: https://huggingface.co/datasets/LocalLLaMA/typed-decisions Curious to hear what you make of the single-pass, no-generation approach. šŸ™Œ
reacted to SeaWolf-AI's post with šŸš€ about 10 hours ago
🧠 We just released Darwin-27B-ZTC, a judgment engine that reaches a verdict without generating anything. Most LLMs answer by generating, decoding one token at a time. Darwin-27B-ZTC takes a different route. āš™ļø How it works šŸ”¹ It makes its call in a single forward pass. šŸ”¹ Zero generated tokens, and no decoding loop. šŸ”¹ That keeps latency and cost far below what a generative model needs. šŸŽÆ What it judges šŸ”¹ It handles several question types: free-form correctness (noul), multiple choice (choice), and scoring (score). šŸ”¹ For each one it hands back a calibrated confidence, not just an answer. šŸ“Š How well calibrated (measured) šŸ”¹ KL 0.204, Brier 0.097, so the confidence it reports lines up with what actually happens. šŸ”¹ 0.743 accuracy (zero-shot, general split), across 2,000 judgments with zero errors. šŸ”¹ By type: noul 0.847, choice 0.723, score 0.675. šŸ”¹ None of the benchmark's train split went into it. It is pure zero-shot. šŸš€ Where it fits šŸ”¹ Grading at scale, model routing, safety gating, anywhere you want a fast decision without paying for generation. šŸ† It currently sits at #1 on the official typed-decisions leaderboard on Hugging Face (0.743 accuracy, zero-shot). šŸ”— Links Model: https://huggingface.co/FINAL-Bench/Darwin-27B-ZTC Leaderboard: https://huggingface.co/datasets/LocalLLaMA/typed-decisions Curious to hear what you make of the single-pass, no-generation approach. šŸ™Œ
View all activity

Organizations

mayafree_ai's profile picture Gemma Challenge's profile picture