Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
Mikhail Gribov
PRO
mihailgribov
1
2
4
Follow
tegridydev's profile picture
quantid's profile picture
Hari5115's profile picture
16 followers
·
10 following
https://subsemantic.com
MihailGGribov
mihail-gribov
mihail-gribov-rs
AI & ML interests
Understanding LLMs from the inside - probing internals, and testing what survives when the model becomes an agent
Recent Activity
updated
a model
1 day ago
mihailgribov/typecastlm-qwen3.5-3.8b
posted
an
update
1 day ago
How good are LLMs as decision models? A benchmark of 22 systems A decision model answers a question about a text with a probability for each predefined option. We measured how well LLMs do this on the 3 358 English questions of Decision Questions: 21 LLMs and TypeSafe's Jev. Each question was asked on its own, with reasoning off, so the comparison is about each model's direct judgement, read from its first answer token, not about how long it thinks. All systems are placed on an item response theory scale tied to fixed anchor questions. What we found Size is not the ranking. Qwen3.8-27B, an open-weights 27B model, comes first (+5.39), statistically indistinguishable from GPT-5.4 (+5.26), Kimi K2.6 (+5.24) and Qwen3.5-397B (+5.23). Accuracy alone misleads. Claude Sonnet 5.5 has the highest accuracy (0.948) but leaves 5.7% of questions undecided, and ranks 7th by IRT ability, which penalises undecided answers. "The text doesn't say" is the hardest answer. When the answer to a true/false/unknown question is "unknown", systems recognise it from 16% to 92% of the time — a 5.75× gap. How to run it on your models The LLMs were asked through llm2decision 0.1.1 (pip install llm2decision), which turns any LLM into a decision model behind one interface. Works with hosted APIs (Nebius, OpenAI, Anthropic, OpenRouter, Mistral), your own vLLM / SGLang / llama.cpp / Ollama, or Jev; switching is one string. Questions can be yes/no, true/false/unknown, choice or score, and probabilities come from the model's token probabilities. Benchmark your own: llm2decision bench <model> --data my_questions.jsonl. No dependencies, Python 3.10+. Where to look All answers, item parameters and scoring script are in the repo — a new model lands on the same scale with one command. Dataset: https://huggingface.co/datasets/mihailgribov/decision-questions Code: https://github.com/mihail-gribov/llm2decision Package: https://pypi.org/project/llm2decision
updated
a dataset
2 days ago
mihailgribov/decision-questions
View all activity
Organizations
mihailgribov
's models
1
Sort: Recently updated
mihailgribov/typecastlm-qwen3.5-3.8b
Text Classification
•
4B
•
Updated
1 day ago
•
1.95k
•
2