Mikhail Gribov PRO
mihailgribov
AI & ML interests
Understanding LLMs from the inside - probing internals, and testing what survives when the model becomes an agent
Recent Activity
updated a model about 14 hours ago
mihailgribov/typecastlm-qwen3.5-3.8b repliedto their post about 14 hours ago
TypeCastLM: Jev-class decision models on frozen LLMs
TypeCastLM is a decision model built around a frozen LLM. Given a text and a question, it returns calibrated probabilities over the allowed answers in one forward pass. The LLM is not fine-tuned: not one of its weights changes. TypeCastLM swaps its LM head for a mini one: a linear matrix whose rows are mostly the model's own output rows, with the three verdict rows (true, false, unsure) fitted. A fitted row catches an answer the model spreads over many tokens, such as "yes", "true" or "correct", where a single vocabulary row sees one word. A head that small leaves little room to overfit. It speaks the Jev API, so a Jev client switches by changing the base URL.
The first model, [typecastlm-qwen3.5-3.8b](https://huggingface.co/mihailgribov/typecastlm-qwen3.5-3.8b) on Qwen3.5-4B, is among the top open frozen-4B models on JevBench v1.6. Besides Jev's yes/no, choice and scale modes it has `tfu`: yes/no with a third answer, `unsure`, for when the text lacks what the decision needs.
Specs:
• Size: 3.76B params; 7.5 GB bf16, 4.0 GB Q8_0
• VRAM: 12 GB is enough, 16 GB comfortable
• Latency p50 (batch 1, bf16, RTX 5060 Ti): ≤200 tok 48 ms · 1k 93 ms · 4k 566 ms
• Calibration (ECE): BoolQ 0.052 · RTE 0.012 · FEVER 0.019
• Context: 32k tokens
• Runtime: transformers or llama.cpp, fully offline
• JevBench v1.6: 19.7
`pip install typecastlm` posted an update 1 day ago
TypeCastLM: Jev-class decision models on frozen LLMs
TypeCastLM is a decision model built around a frozen LLM. Given a text and a question, it returns calibrated probabilities over the allowed answers in one forward pass. The LLM is not fine-tuned: not one of its weights changes. TypeCastLM swaps its LM head for a mini one: a linear matrix whose rows are mostly the model's own output rows, with the three verdict rows (true, false, unsure) fitted. A fitted row catches an answer the model spreads over many tokens, such as "yes", "true" or "correct", where a single vocabulary row sees one word. A head that small leaves little room to overfit. It speaks the Jev API, so a Jev client switches by changing the base URL.
The first model, [typecastlm-qwen3.5-3.8b](https://huggingface.co/mihailgribov/typecastlm-qwen3.5-3.8b) on Qwen3.5-4B, is among the top open frozen-4B models on JevBench v1.6. Besides Jev's yes/no, choice and scale modes it has `tfu`: yes/no with a third answer, `unsure`, for when the text lacks what the decision needs.
Specs:
• Size: 3.76B params; 7.5 GB bf16, 4.0 GB Q8_0
• VRAM: 12 GB is enough, 16 GB comfortable
• Latency p50 (batch 1, bf16, RTX 5060 Ti): ≤200 tok 48 ms · 1k 93 ms · 4k 566 ms
• Calibration (ECE): BoolQ 0.052 · RTE 0.012 · FEVER 0.019
• Context: 32k tokens
• Runtime: transformers or llama.cpp, fully offline
• JevBench v1.6: 19.7
`pip install typecastlm`