Mikhail Gribov's picture

Mikhail Gribov PRO

mihailgribov

AI & ML interests

Understanding LLMs from the inside - probing internals, and testing what survives when the model becomes an agent

Recent Activity

updated a model 1 day ago
mihailgribov/typecastlm-qwen3.5-3.8b
posted an update 1 day ago
How good are LLMs as decision models? A benchmark of 22 systems A decision model answers a question about a text with a probability for each predefined option. We measured how well LLMs do this on the 3 358 English questions of Decision Questions: 21 LLMs and TypeSafe's Jev. Each question was asked on its own, with reasoning off, so the comparison is about each model's direct judgement, read from its first answer token, not about how long it thinks. All systems are placed on an item response theory scale tied to fixed anchor questions. What we found Size is not the ranking. Qwen3.8-27B, an open-weights 27B model, comes first (+5.39), statistically indistinguishable from GPT-5.4 (+5.26), Kimi K2.6 (+5.24) and Qwen3.5-397B (+5.23). Accuracy alone misleads. Claude Sonnet 5.5 has the highest accuracy (0.948) but leaves 5.7% of questions undecided, and ranks 7th by IRT ability, which penalises undecided answers. "The text doesn't say" is the hardest answer. When the answer to a true/false/unknown question is "unknown", systems recognise it from 16% to 92% of the time — a 5.75× gap. How to run it on your models The LLMs were asked through llm2decision 0.1.1 (pip install llm2decision), which turns any LLM into a decision model behind one interface. Works with hosted APIs (Nebius, OpenAI, Anthropic, OpenRouter, Mistral), your own vLLM / SGLang / llama.cpp / Ollama, or Jev; switching is one string. Questions can be yes/no, true/false/unknown, choice or score, and probabilities come from the model's token probabilities. Benchmark your own: llm2decision bench <model> --data my_questions.jsonl. No dependencies, Python 3.10+. Where to look All answers, item parameters and scoring script are in the repo — a new model lands on the same scale with one command. Dataset: https://huggingface.co/datasets/mihailgribov/decision-questions Code: https://github.com/mihail-gribov/llm2decision Package: https://pypi.org/project/llm2decision
updated a dataset 2 days ago
mihailgribov/decision-questions
View all activity

Organizations

AI Cordon's profile picture