In a Training Loop 🔄

Stefan Schweter PRO

stefan-it

AI & ML interests

Flair Library 💕, NER & PoS Tagging, LM Pretraining (mostly encoder-only & encoder-decoder), Historical Language Models, German Language Models, Bavarian NLP 🥨

Recent Activity

liked a Space about 11 hours ago

FINAL-Bench/all-bench-leaderboard

liked a model 1 day ago

allenai/olmOCR-2-7B-1025-FP8

liked a dataset 1 day ago

unicamp-dl/mmarco

View all activity

Organizations

liked a Space about 11 hours ago

All Bench Leaderboard

🔥

benchmarking metrics to 31 leading LLM models

liked a model 1 day ago

allenai/olmOCR-2-7B-1025-FP8

Image-Text-to-Text • 8B • Updated 14 days ago • 277k • 209

liked a dataset 1 day ago

unicamp-dl/mmarco

Updated Mar 6, 2024 • 2.14k • 90

upvoted a collection 3 days ago

🤏 Smol-Data

Collection

Tried and tested mixes for strong pretraining. Inspired by https://huggingface.co/blog/codelion/optimal-dataset-mixing • 14 items • Updated 3 days ago • 11

reacted to hannayukhymenko's post with 🔥❤️ 4 days ago

Post

1862

Do you translate your benchmarks from English correctly? 🤔
Turns out, for many languages it is much harder than you can imagine!

Introducing Recovered in Translation 🌍 together with @aalexandrov
ritranslation.insait.ai

Translating benchmarks is a painful process, requiring a lot of manual inspection and adjustments. You start from setting up the whole pipeline and adapting to every format type, including task specifics. There already exist some massive benchmarks, but they still have some simple (and sometimes silly) bugs, which can hurt the evaluations :( We present a novel automated translation framework to help with that!

Eastern and Southern European languages introduce richer linguistic structures compared to English and for benchmarks which heavily rely on grammatical coherence machine translation presents a risk of harming evaluations. We discover potential answer leakage or misleading through grammatical structure of the questions. Some benchmarks are also just outdated and need to be retranslated with newer and better models.

We present a framework with novel test-time scaling methods which allow to control time and cost investments, while at the same time mitigate the need for human-in-the-loop verification. While working on Ukrainian-focused MamayLM models, we had to translate 10+ benchmarks in a short span of time. Finding human evaluators is costly and time-consuming, same goes for using professional translators. With our pipeline we were able to do it in 3 days🏎️

We hope our findings will help enable stronger multilingual evaluations and developments. We release all produced benchmarks on Hugging Face together with the source code and Arxiv paper 🤗

Paper: Recovered in Translation: Efficient Pipeline for Automated Translation of Benchmarks and Datasets (2602.22207)
Code: https://github.com/insait-institute/ritranslation
Benchmarks: https://huggingface.co/collections/INSAIT-Institute/multilingual-benchmarks

1 reply

upvoted a paper 4 days ago

Recovered in Translation: Efficient Pipeline for Automated Translation of Benchmarks and Datasets

Paper • 2602.22207 • Published 8 days ago • 39

liked a dataset 9 days ago

windprak/steuerllm_instruct_dataset

Preview • Updated 21 days ago • 41 • 1

reacted to SeaWolf-AI's post with 🔥 9 days ago

Post

4085

Do Bubbles Form When Tens of Thousands of AIs Simulate Capitalism?

We gave LLMs autonomous trading over 30 real tickers at 100x leverage. All went bankrupt in 30 minutes from hallucination. This spawned FINAL Bench (first metacognition benchmark) and AI NPC Trading Arena — tens of thousands of metacognition-equipped AI agents competing under capitalist rules. Humans can only watch.

Live Demo: Heartsync/Prompt-Dump
Article: https://huggingface.co/blog/FINAL-Bench/pumpdump

NPCs form a society: 3-tier memory, self-modifying parameters, mutual criticism, strategy propagation, and a virtual SEC enforcing fines every 20 minutes. Every trade passes 4-stage verification including Brave Search fact-check. FINAL Bench confirmed across 9 SOTA models that AI can say "I might be wrong" (MA 0.694) but cannot actually fix errors (ER 0.302).

Six findings: Bubbles form naturally through knowledge transfer and swarm herding. Identical NPCs diverge irreversibly from their first three trades. Metacognition blocks individual hallucination but not collective herding — this is the key finding. Information asymmetry solidifies hierarchy. Fraud and regulation co-evolve. Criticism improves returns.

Individual intelligence does not guarantee collective intelligence.

Dataset & Paper:
FINAL-Bench/Metacognitive

1 reply

liked a dataset 9 days ago

castorini/NanoKnow-Fineweb-Edu-Index

Updated 8 days ago • 1.38k • 2

upvoted a paper 9 days ago

NanoKnow: How to Know What Your Language Model Knows

Paper • 2602.20122 • Published 10 days ago • 6

upvoted an article 9 days ago

Article

Do Bubbles Form When Tens of Thousands of AIs Simulate Capitalism?

10 days ago

•

liked a dataset 9 days ago

BabyLM-community/babylm-deu

Viewer • Updated Oct 15, 2025 • 36.6k • 56 • 2

upvoted a paper 10 days ago

The Million-Label NER: Breaking Scale Barriers with GLiNER bi-encoder

Paper • 2602.18487 • Published 23 days ago • 5

reacted to umarbutler's post with ❤️ 13 days ago

Post

2176

@abdurrahmanbutler and I just dropped Legal RAG Bench, the first benchmark for legal RAG systems to simultaneously evaluate hallucinations, retrieval failures, and reasoning errors.

Our key takeaways are:
1. Embedding models, not generative models, are the primary driver of RAG accuracy. Switching from a general-purpose embedder like OpenAI's Text Embedding 3 Large to a legal domain embedder like Isaacus' Kanon 2 Embedder can raise accuracy by ~19 points.
2. Hallucinations are often triggered by retrieval failures. Fix your retrieval stack, and, in most cases, you end up fixing hallucinations.
3. Once you have a solid legal retrieval engine like Kanon 2 Embedder, it doesn’t matter as much what generative model you use; GPT-5.2 and Gemini 3.1 Pro perform relatively similarly, with Gemini 3.1 Pro achieving slightly better accuracy at the cost of more hallucinations.
4. Google's latest LLM, Gemini 3.1 Pro, is actually a bit worse than its predecessor at legal RAG, achieving 79.3% accuracy instead of 80.3%.

These findings confirm what we already knew at Isaacus: that information retrieval sets the ceiling on the accuracy of legal RAG systems. It doesn’t matter how smart you are; you aren’t going to magically know what the penalty is for speeding in California without access to an up-to-date copy of the California Vehicle Code.

Even still, to our knowledge, we’re the first to actually show this empirically.

Unfortunately, as we highlight in our write-up, high-quality open legal benchmarks like Legal RAG Bench and our earlier MLEB are few and far between.

In the interests of transparency, we have not only detailed exactly how we built Legal RAG Bench, but we’ve also released all of our data openly on Hugging Face. You can read our write up [here](https://isaacus.com/blog/legal-rag-bench), noting that we’ll soon be publishing it as a paper.

Kudos to my brother @abdurrahmanbutler for serving as the lead author on this monumental release.

2 replies

liked a dataset 13 days ago

Eurolingua/HPLT3_DE_0.9_Quantile_Adult_Filtered

Viewer • Updated 13 days ago • 9.99M • 30 • 1

reacted to qgallouedec's post with 🔥 14 days ago

Post

2659

@CohereLabs just released 🌿 Tiny Aya: a fully open-source 3B parameter model that speaks 70+ languages 🌍! But there’s a catch:

Tiny Aya is just a language model. It doesn’t support tool calling, the key capability that turns frontier models into powerful *agents*.
So the real question is:

How hard is it to turn Tiny Aya into an agent?

Turns out… it’s simple, thanks to Hugging Face TRL.
We’re sharing a hands-on example showing how to train Tiny Aya to turn it into a tool-calling agent using TRL, unlocking what could become the first *massively multilingual open agent*.

Small model. Global reach. Agent capabilities.

👉 https://github.com/huggingface/trl/blob/main/examples/notebooks/sft_tool_calling.ipynb