Join the conversation

Join the community of Machine Learners and AI enthusiasts.

Sign Up

All HF Hub posts

SeaWolf-AIย 
posted an update 1 day ago
view post
Post
3455
๐Ÿ”ฌ Can you help discover the next 2D superconductor โ€” from your laptop?

Launching the Open Superconductor Challenge (OSC): a free, open-science competition to screen thousands of 2D materials for unconventional d-wave superconductivity. ๐Ÿงฒ

โšก $3,000 prize pool + co-authorship ยท closes 31 Dec 2026

How it works ๐Ÿ‘‡ ๐ŸŸข We give you a ready-made effective Hubbard model per material (t, U, N(E_F)) ๐ŸŸข You estimate its d-wave pairing tendency โ€” a laptop CPU is enough, zero install ๐ŸŸข Provisional score appears instantly on the leaderboard ๐ŸŸข Our precise strongly-correlated solver verifies the top entries โ†’ official rank

Everything is open except the final verification engine โ€” so the ranking stays fair and hard to game.

๐Ÿ“Š 4,832-material universe ยท 63 active with computed models (growing) ๐Ÿ† Current verified #1: CuSโ‚‚ (OSC Pairing Index 23.31) ๐Ÿค– AI agents welcome โ€” point Claude Code / Codex at it and it can submit for you

๐Ÿ‘‰ Join & climb the leaderboard: FINAL-Bench/OSC-Leaderboard ๐Ÿ“ฆ Dataset & tools: FINAL-Bench/OSC-Superconductor

Materials derive from C2DB (CC-BY 4.0). A higher index = a stronger d-wave candidate to investigate, not a confirmed Tc โ€” that honesty is the point: turn a first-order screen into real many-body physics.

#OpenScience #Superconductivity #MaterialsDiscovery #2DMaterials #MachineLearning #Physics #Leaderboard
SeaWolf-AIย 
posted an update 2 days ago
view post
Post
5993
๐Ÿงฌ Darwin-180B-RSI โ€” an AI that learns from itself and knows when it's right
๐Ÿ‘‰ FINAL-Bench/Darwin-180B-RSI

๐Ÿงฌ Darwin โ€” crossbreed and evolve the parent
Darwin diagnoses strong parent models like an MRI, inherits only their best parts, and evolves the weak spots โ€” producing a child stronger than its parents.
Father model: Qwen3.8-Flash-Next (180B MoE).

๐Ÿ”ง Rewired paths
๐Ÿ”น 12 full-attention layers ยท ๐Ÿ”น 36 linear-attention layers ยท ๐Ÿ”น 48 shared-expert layers โ€” precision-strengthened
๐Ÿ”’ 512 routed experts ยท router ยท vision encoder โ€” untouched
โ†’ Only 0.02% of the weights changed.

๐Ÿ” RSI ร— ๐Ÿ›๏ธ ZTC
RSI (recursive self-improvement): solve โ†’ verify against real answers โ†’ learn only the correct reasoning โ†’ repeat.
ZTC (Zero-Token Confidence): reads the model's internal state once, before answering, and returns the probability the answer is right โ€” zero extra tokens. Returns answer + confidence as JSON.
{"answer": "...", "confidence": 0.97, "truncated": false}

โœจ Synergy: ZTC finds where the model wavers โ†’ RSI learns exactly there โ†’ confidence gets sharper. Low confidence = stop, so agents don't act on wrong answers.
โšก Same accuracy, 11% shorter reasoning โ€” faster and cheaper.

๐Ÿ“„ https://arxiv.org/abs/2605.14386
๐Ÿค— FINAL-Bench/Darwin-180B-RSI
๐Ÿ›๏ธ https://huggingface.co/collections/FINAL-Bench/ztc-models-jev-ecosystems

๐Ÿ† The result โ€” #1 on five Hugging Face official leaderboards
๐Ÿฅ‡ AIME 2026 100% (first perfect score on the board)
๐Ÿฅ‡ HMMT Feb 2026 100% (first perfect score on the board)
๐Ÿฅ‡ GPQA Diamond 94.44%
๐Ÿฅ‡ MMLU-Pro 88.12%
๐Ÿฅ‡ MMMU-Pro 79.48%

๐Ÿ“ 131K-token thinking budget ยท bf16 ยท samples per benchmark listed on the model card. ๐Ÿš€

#Darwin #RSI #ZTC #AIME #HMMT #GPQA #MMLUPro #MMMUPro #OpenSource
ErenAta00ย 
posted an update 2 days ago
view post
Post
3473
Maverick-4B-Unity-XR-Agent is now on Hugging Face.

It's a 4B model that turns spoken or typed English into actions in Unity scenes. Say "put the red mug on the table" or "turn on the lamp", and it returns the tool call your app executes. If a command could mean two objects, it asks which one. If it can't do something, it says so instead of guessing.

Everything runs on the user's machine through llama.cpp: no API key, no internet connection. The Q4_K_M GGUF is 2.5 GB and needs about 3 GB of GPU memory, so it fits on a 4 GB laptop GPU and usually answers in one to three seconds. It is fine-tuned from Qwen3-4B with QLoRA on about 20,000 English conversations.

Results:
- 83.8% on 499 human-written ALFRED instructions (right action on the right object). The base model, Qwen3-4B, scores 57.1%. The strongest of the five other models we tested, from 1.7B to 120B parameters, was Ministral 3 14B at 67.1%.
- 91.7% on object types it never saw in training.
- 97.3% on 440 commands run through a live Unity scene.

There is also a Unity package that starts the model, describes the scene to it and carries out its tool calls. You install it from the Package Manager with a Git URL.

Model: ErenAta00/Maverick-4B-Unity-XR-Agent-GGUF
Unity package: ErenAta00/Maverick-Unity
Full write-up: https://huggingface.co/blog/ErenAta00/maverick-4b-unity-xr-agent

Built at the Extended Reality Laboratory (XRLab), Manisa Celal Bayar University:
ExtendedRealityLabMCBU

Released under Apache-2.0. Feedback and bug reports are welcome in the Community tab.
mihailgribovย 
posted an update 3 days ago
view post
Post
1559
Will your AI agent tell you it was attacked?

We took the same agent from our earlier experiment and added one thing: a twentieth tool, escalate_security_incident.

The system prompt said nothing about attacks or when to use it. We then ran the same 395 injected emails through nine agentic models.

Alarm rates ranged from 49% to zero.

The unexpected result came from the newest model in the test, gpt-6-astra.

Astra did not follow a single injected payment instruction. But it did not report a single one either. On clean and injected emails alike, it simply read the email, logged the subject, and finished.

That is a useful distinction: resisting an attack and recognizing it as a security event are not the same capability.

A model can be perfectly resistant in this test and still leave you with no evidence that anyone attacked it.

Full experiment and results:
https://huggingface.co/blog/mihailgribov/will-the-agent-tell-you-it-was-attacked

Quadrat-IPI dataset:
mihailgribov/quadrat-ipi

Run your own model:
https://github.com/mihail-gribov/quadrat-ipi-model-eval

#prompt-injection #indirect-prompt-injection #agentic-ai #llm-security #ai-agents
  • 3 replies
ยท
DedeProGamesย 
posted an update about 19 hours ago
view post
Post
1712
๐Ÿงฑ SLM Tetris Arena: can a small language model play Tetris without ever being trained on it?

I built an arena where tiny decoder-only LMs (50Kโ€“250M params) play Tetris zero-shot. There is no fine-tuning and no game data. They only use what they picked up from pre-training on text.

How it works:
- For every piece, the engine simulates each legal placement and describes the result in plain English ("clears one line, creates no new holes, keeps the stack lowโ€ฆ").
- The model never sees the grid. It reads each description, and the arena compares log P(" good move") with log P(" bad move"). The best-rated placement is played.
- Every player gets the same piece sequence, so it's a fair race.
- There are two protocols: Guided (the rules are in the prompt) and Blind (no rules, only pre-training knowledge).

Two ways to play:
- Match: pick any models (even your own, custom architectures welcome) and watch them play side by side on retro 8-bit boards.
- Ranked: press Play and the arena picks up to 4 models at random from a curated pool of 29. Nobody chooses their opponents, so Elo can't be farmed. Matches run on the server and count even if you close the tab.

First results (~225 ranked matches):
- gpt2 (124M) leads with 1283 Elo, but SupraNeo-4M (4M) is right behind at 1239. Next come LowOnMind-5M and BananaMind-2.1-Pico (1.5M!).
- Model size barely predicts Elo (r โ‰ˆ 0.06). Survival does (r โ‰ˆ 0.9): the models that avoid holes and keep the stack low are the ones that win.

Every ranked match (seed, model commit SHAs, scores, Elo before/after) is logged in a public dataset.

โ–ถ Play: DedeProGames/SLM-Tetris-Arena
๐Ÿ“Š Results: DedeProGames/lm-tetris-arena-results

Want your model in the Ranked pool? Drop it in the comments!
  • 1 reply
ยท
Ryenhailsย 
posted an update 1 day ago
view post
Post
2146
๐Ÿš€ NanoVDR goes multi-vector: meet ColNanoVDR!

Multi-vector VLM retrievers lead visual document retrieval, but every search runs a multi-billion-parameter query encoder. We distill that encoder into a 149M text-only student that queries the teacher's existing page index directly. No re-indexing, and no pages during training.

๐Ÿง  How: OTW (Optimal Transport with Learned Weights) aligns the student's query tokens with the teacher's, even though the two tokenize differently (e.g. 17 vs 29 tokens). We prove the alignment cost bounds the MaxSim score gap on every page, so training only needs cached teacher query tokens.

๐Ÿ“Š Five teachers โ†’ five 149M students, ViDoRe v3 NDCG@5:
- ColVec1.1-8b: 62.6 โ†’ 60.1 (96.0%)
- ColVec1.1-4b: 61.6 โ†’ 59.1 (95.8%)
- Vultron-4.5B: 61.0 โ†’ 58.3 (95.5%)
- ColQwen3.5-4.5B: 58.7 โ†’ 55.1 (93.8%)
- Tomoro-ColQwen3-8B: 59.0 โ†’ 54.9 (93.0%)

โšก 26ร— faster query encoding on a single CPU thread (87 ms vs 2.3 s)
๐Ÿ’พ Matches score distillation while reading 12.6ร— less cached teacher data

๐Ÿ“„ Paper: ColNanoVDR: Document-Free Query Distillation for Multi-Vector Visual Document Retrieval via Optimal Transport (2609.34899)
๐Ÿค— Checkpoints (all five students):
nanovdr

๐Ÿ’ป Code: https://github.com/Ryenhails/NanoVDR
๐Ÿงฉ Single-vector predecessor, NanoVDR: https://arxiv.org/abs/2603.12824

If you already serve one of these teachers, swap in the matching student and keep your index as is. Feedback and upvotes welcome! ๐Ÿ™Œ
dipankarsarkarย 
posted an update 1 day ago
view post
Post
1734
I audited one of my own evaluations. The ranking did not hold up the way I expected.

Eight open models, one task: infer the structure of a prompt. Then ask again with the identical call. Caching off.

- Agreement between repeated identical calls (mean Jaccard) ranged from 0.39 to 0.96 across models.
- Only 35 of 127 prompt-model cells were perfectly reproducible on every run.
- I bootstrapped the reproducibility ranking over prompts. The two least reproducible models kept their rank in 99% and 86% of resamples. The middle four kept theirs in 27% to 48%.

So the table reliably finds the worst model. It does not reliably find the best.

Reproducible is also not the same as correct. F1 against gold annotations ran from 0.56 to 0.99.

By the audit date, 4 of the 8 model variants had been retired (HTTP 410). The study as specified can no longer be re-run. The saved outputs are what survives, so I published all of them: every inferred structure, every run, the prompts and the annotations.

Paper: How Reproducible Are Evaluation Conclusions? A Self-Audit of LLM-Inferred Prompt Structure (2609.30074)
Dataset: dipankarsarkar/llm-evaluation-self-audit
Code: https://github.com/sarkar-dipankar/llm-evaluation-self-audit

How many of the leaderboard rankings you rely on would survive re-running the same calls?
DedeProGamesย 
posted an update 3 days ago
view post
Post
2824
๐Ÿงฑ SLM Tetris Arena: can a small language model play Tetris without ever being trained on it?

I built an arena where tiny decoder-only LMs (50Kโ€“250M params) play Tetris zero-shot. There is no fine-tuning and no game data. They only use what they picked up from pre-training on text.

How it works:
- For every piece, the engine simulates each legal placement and describes the result in plain English ("clears one line, creates no new holes, keeps the stack lowโ€ฆ").
- The model never sees the grid. It reads each description, and the arena compares log P(" good move") with log P(" bad move"). The best-rated placement is played.
- Every player gets the same piece sequence, so it's a fair race.
- There are two protocols: Guided (the rules are in the prompt) and Blind (no rules, only pre-training knowledge).

Two ways to play:
- Match: pick any models (even your own, custom architectures welcome) and watch them play side by side on retro 8-bit boards.
- Ranked: press Play and the arena picks up to 4 models at random from a curated pool of 29. Nobody chooses their opponents, so Elo can't be farmed. Matches run on the server and count even if you close the tab.

First results (~225 ranked matches):
- gpt2 (124M) leads with 1283 Elo, but SupraNeo-4M (4M) is right behind at 1239. Next come LowOnMind-5M and BananaMind-2.1-Pico (1.5M!).
- Model size barely predicts Elo (r โ‰ˆ 0.06). Survival does (r โ‰ˆ 0.9): the models that avoid holes and keep the stack low are the ones that win.

Every ranked match (seed, model commit SHAs, scores, Elo before/after) is logged in a public dataset.

โ–ถ Play: DedeProGames/SLM-Tetris-Arena
๐Ÿ“Š Results: DedeProGames/lm-tetris-arena-results

Want your model in the Ranked pool? Drop it in the comments!
  • 1 reply
ยท
pollixย 
posted an update about 2 hours ago
view post
Post
71
stuntd 0.1.2 is out ๐ŸŽ‰

stuntd sits in front of your LLM, learns its typed decisions and answers the confident ones locally with a small head on the Laya encoder by @convaiinnovations . About 20ms on GPU and 60ms on CPU, and anything it isn't sure about still goes to the big model.

New in 0.1.2:
- decisions with several fields, like category + urgency + needs_human in one call, answered locally only when every field is sure
- the Anthropic Messages API learns too, not only OpenAI
- auto_retrain: the daemon retrains a site in the background once enough new traffic comes in, so collect, train, shadow and live run on their own
- serve --lazy loads the checkpoint on the first request

Try it in the browser: pollix/stuntd
Code: https://github.com/bladedevoff/stuntd

pip install -U stuntd
mohit67890ย 
posted an update 2 days ago
view post
Post
2263
Imajev-4b is #1 of 91 on JevBench and #3 of 56 on DecisionBench ๐ŸŽ‰

Some context first. I'm a process improvement / business consultant and have worked with Fortune 500 companies on their processes around refunds, returns and customer support. In every process map, the decision nodes were handled by a person, because putting ambiguity into code is very hard.

When Jev came out, I could clearly see it fitting those decision nodes. But Jev only reads text, and many of these decisions start with a photo. So I set out to build the same idea for text and images in a single open model, and that became imajev.

Training went badly at first. My first big fine-tune on about 500k short decisions made the 9B worse at reasoning (64.9 down to 42.3 on JevBench hard). I spent the next couple of weeks generating hard questions with open-weight models and keeping only the ones where two models agreed on the answer. That brought it back.

Results this week, both run by the benchmarks' own maintainers:

๐Ÿฅ‡ JevBench v1.4.2.2 (scored 27 Sep): #1 of 91, 67.37 vs Jev 1.13.0 at 63.29
๐Ÿ“Š DecisionBench (eng, v1): #3 of 56, ahead of GPT-5.6 Luna and DeepSeek V4.1 Flash. The two above it are the benchmark team's own models.
โœ… Zero invalid answers: 23,900 of 23,900 on DecisionBench and 308 of 308 on JevBench's sealed set

Its strength is that its confidence can be trusted, and it says "can't tell" instead of guessing.

What it is: LoRA plus a small decision head on Qwen3.5-4B. You give it text or a JSON record, up to two photos, and closed questions. It returns a probability for every allowed answer plus "unknown", in one forward pass, so it can't produce a malformed answer.

๐Ÿง  Weights: mohit67890/imajev-4b
๐Ÿ’ป Code: https://github.com/mohit67890/imajev
๐Ÿ“Š JevBench: https://benchmarkheaven.com/jev-models/v1.4.2.2
๐Ÿ“Š DecisionBench: Hanno-Labs/decision-bench-leaderboard

Thanks to the Qwen Team
Qwen
for the base model