Instructions to use nasrellahkharroubi/DarijaDZ-DialectID-Classifier with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Scikit-learn
How to use nasrellahkharroubi/DarijaDZ-DialectID-Classifier with Scikit-learn:
from huggingface_hub import hf_hub_download import joblib model = joblib.load( hf_hub_download("nasrellahkharroubi/DarijaDZ-DialectID-Classifier", "sklearn_model.joblib") ) # only load pickle files from sources you trust # read more about it here https://skops.readthedocs.io/en/stable/persistence.html - Notebooks
- Google Colab
- Kaggle
DarijaDZ Dialect Identification Classifier
The problem. Algerian online text mixes several dialects and languages -- Algerian Darija in Arabic script, Modern Standard Arabic, Arabizi (Darija in Latin script), French, and English -- often switching between them inside a single message. Telling these apart is the first step for almost any downstream Algerian-NLP task. This model does that classification, trained and selected on DarijaDZ-DialectID. It is part of DarijaDZ, an attempt to build an NLP ecosystem for Algerian Darija.
The model. Character n-gram TF-IDF features + an RBF-kernel SVM, with one model per script group. It was the best of twelve approaches compared on this task (word-cluster baselines, a transformer classifier head, an HMM, a fastText-style model, KNN over sentence embeddings, a from-scratch language-model-weighted vote, and more) -- a simple architecture beating the heavier ones.
The byte-vs-codepoint, n-gram-order comparison it was chosen from is based on Baldwin & Lui, "Language Identification: The Long and the Short of the Matter" (NAACL 2010) -- aclanthology.org/N10-1027.
How it works
Same script-gated design as the dataset's own labeling pipeline:
- A deterministic script-detection rule decides
code_switch(both Arabic and Latin script present) directly . - Arabic-script-only text -> the arabic model, restricted to
{msa, darija}. - Latin-script-only text -> the latin model, restricted to
{arabize, french, english}.
| Model | Classes | Features |
|---|---|---|
arabic_* |
msa, darija |
character trigrams, TF-IDF, 3,000-term vocab |
latin_* |
arabize, french, english |
character bigrams, TF-IDF, 3,000-term vocab |
Training data
Trained on DarijaDZ-DialectID
after its class-balancing pass -- 59,997 rows, ~10,000 per class.
The dataset is hybrid: darija/arabize/code_switch are real
Algerian text, while msa/french/english are mostly external
(Arabic Wikipedia / film reviews / news-site comments), so those three
classes partly reflect their source register rather than Algerian
usage.
Performance
Measured on a fresh, stratified 80/20 held-out split of the balanced 59,997-row dataset (train-only fit; the shipped weights are then refit on all 60k).
| Group | Accuracy | Per-class recall |
|---|---|---|
arabic (msa/darija) |
0.8912 | msa 0.848, darija 0.935 |
latin (arabize/french/english) |
0.9562 | arabize 0.972, french 0.938, english 0.959 |
How to Load
pip install skops scikit-learn
import re
from huggingface_hub import hf_hub_download
import skops.io as sio
REPO_ID = "nasrellahkharroubi/DarijaDZ-DialectID-Classifier"
def load_group(group: str):
vec_path = hf_hub_download(REPO_ID, f"{group}_vectorizer.skops")
svm_path = hf_hub_download(REPO_ID, f"{group}_svm.skops")
# Both files only contain standard scikit-learn/scipy objects --
# inspect with sio.get_untrusted_types(file=...) yourself if you
# want to verify before trusting.
vec = sio.load(vec_path, trusted=sio.get_untrusted_types(file=vec_path))
svm = sio.load(svm_path, trusted=sio.get_untrusted_types(file=svm_path))
return vec, svm
arabic_vec, arabic_svm = load_group("arabic")
latin_vec, latin_svm = load_group("latin")
ARABIC_CLASSES = ["msa", "darija"]
LATIN_CLASSES = ["arabize", "french", "english"]
_ARABIC_RE = re.compile(r"[-ۿݐ-ݿࢠ-ࣿﭐ-﷿ﹰ-]")
_LATIN_RE = re.compile(r"[a-zA-ZÀ-ɏ]")
def classify(text: str) -> str:
has_ar, has_lat = bool(_ARABIC_RE.search(text)), bool(_LATIN_RE.search(text))
if has_ar and has_lat:
return "code_switch" # decided here, no model call
if has_ar:
idx = arabic_svm.predict(arabic_vec.transform([text]))[0]
return ARABIC_CLASSES[idx]
if has_lat:
idx = latin_svm.predict(latin_vec.transform([text]))[0]
return LATIN_CLASSES[idx]
return "other" # no real script content
print(classify("wach rak khouya")) # arabize
print(classify("Bonjour tout le monde")) # french
Part of DarijaDZ
This model is part of DarijaDZ, an attempt to build an NLP ecosystem for Algerian Darija -- see the DarijaDZ corpus, the DarijaDZ-DialectID dataset this model was trained on, and the Arabizi transliterator for the other pieces of that effort.
Citation
Kharroubi Nasrellah.
DarijaDZ Dialect Identification Classifier.
2026.
- Downloads last month
- -