DarijaDZ Dialect Identification Classifier

The problem. Algerian online text mixes several dialects and languages -- Algerian Darija in Arabic script, Modern Standard Arabic, Arabizi (Darija in Latin script), French, and English -- often switching between them inside a single message. Telling these apart is the first step for almost any downstream Algerian-NLP task. This model does that classification, trained and selected on DarijaDZ-DialectID. It is part of DarijaDZ, an attempt to build an NLP ecosystem for Algerian Darija.

The model. Character n-gram TF-IDF features + an RBF-kernel SVM, with one model per script group. It was the best of twelve approaches compared on this task (word-cluster baselines, a transformer classifier head, an HMM, a fastText-style model, KNN over sentence embeddings, a from-scratch language-model-weighted vote, and more) -- a simple architecture beating the heavier ones.

The byte-vs-codepoint, n-gram-order comparison it was chosen from is based on Baldwin & Lui, "Language Identification: The Long and the Short of the Matter" (NAACL 2010) -- aclanthology.org/N10-1027.


How it works

Same script-gated design as the dataset's own labeling pipeline:

  1. A deterministic script-detection rule decides code_switch (both Arabic and Latin script present) directly .
  2. Arabic-script-only text -> the arabic model, restricted to {msa, darija}.
  3. Latin-script-only text -> the latin model, restricted to {arabize, french, english}.
Model Classes Features
arabic_* msa, darija character trigrams, TF-IDF, 3,000-term vocab
latin_* arabize, french, english character bigrams, TF-IDF, 3,000-term vocab

Training data

Trained on DarijaDZ-DialectID after its class-balancing pass -- 59,997 rows, ~10,000 per class. The dataset is hybrid: darija/arabize/code_switch are real Algerian text, while msa/french/english are mostly external (Arabic Wikipedia / film reviews / news-site comments), so those three classes partly reflect their source register rather than Algerian usage.

Performance

Measured on a fresh, stratified 80/20 held-out split of the balanced 59,997-row dataset (train-only fit; the shipped weights are then refit on all 60k).

Group Accuracy Per-class recall
arabic (msa/darija) 0.8912 msa 0.848, darija 0.935
latin (arabize/french/english) 0.9562 arabize 0.972, french 0.938, english 0.959

How to Load

pip install skops scikit-learn
import re
from huggingface_hub import hf_hub_download
import skops.io as sio

REPO_ID = "nasrellahkharroubi/DarijaDZ-DialectID-Classifier"

def load_group(group: str):
    vec_path = hf_hub_download(REPO_ID, f"{group}_vectorizer.skops")
    svm_path = hf_hub_download(REPO_ID, f"{group}_svm.skops")
    # Both files only contain standard scikit-learn/scipy objects --
    # inspect with sio.get_untrusted_types(file=...) yourself if you
    # want to verify before trusting.
    vec = sio.load(vec_path, trusted=sio.get_untrusted_types(file=vec_path))
    svm = sio.load(svm_path, trusted=sio.get_untrusted_types(file=svm_path))
    return vec, svm

arabic_vec, arabic_svm = load_group("arabic")
latin_vec, latin_svm = load_group("latin")

ARABIC_CLASSES = ["msa", "darija"]
LATIN_CLASSES = ["arabize", "french", "english"]

_ARABIC_RE = re.compile(r"[؀-ۿݐ-ݿࢠ-ࣿﭐ-﷿ﹰ-]")
_LATIN_RE = re.compile(r"[a-zA-ZÀ-ɏ]")

def classify(text: str) -> str:
    has_ar, has_lat = bool(_ARABIC_RE.search(text)), bool(_LATIN_RE.search(text))
    if has_ar and has_lat:
        return "code_switch"          # decided here, no model call
    if has_ar:
        idx = arabic_svm.predict(arabic_vec.transform([text]))[0]
        return ARABIC_CLASSES[idx]
    if has_lat:
        idx = latin_svm.predict(latin_vec.transform([text]))[0]
        return LATIN_CLASSES[idx]
    return "other"                    # no real script content

print(classify("wach rak khouya"))          # arabize
print(classify("Bonjour tout le monde"))    # french

Part of DarijaDZ

This model is part of DarijaDZ, an attempt to build an NLP ecosystem for Algerian Darija -- see the DarijaDZ corpus, the DarijaDZ-DialectID dataset this model was trained on, and the Arabizi transliterator for the other pieces of that effort.


Citation

Kharroubi Nasrellah.
DarijaDZ Dialect Identification Classifier.
2026.
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train nasrellahkharroubi/DarijaDZ-DialectID-Classifier