A comprehensive and trustworthy benchmark of AI methods for change detection in Earth observation

Tadej Tomanič1, 2 · Alice Baudhuin1 · Jan Sotošek1 · Jure Brence1, 3 · Panče Panov1, 3 · Nikola Simidjievski1, 4 · Dragi Kocev1, 3


1Bias Variance Labs, d.o.o.  2University of Ljubljana, Faculty of Mathematics and Physics
3Department of Knowledge Technologies, Jožef Stefan Institute  4Télécom Paris, Institut Polytechnique de Paris




Summary

Companion artifacts for the paper, A comprehensive and trustworthy benchmark of AI methods for change detection in Earth observation (currently under review).

This study presents a standardized benchmark for change detection in Earth observation (EO), developed as part of the OSCARS-funded FAIR-EO project. The framework is integrated into the AiTLAS toolbox.

Despite rapid advancements in machine learning for remote sensing, accurately evaluating and comparing change detection models remains a significant challenge. Variations in dataset preprocessing, data splits, and evaluation metrics often lead to inconsistent results, making it difficult to determine whether a new method genuinely outperforms existing ones. To address this, we introduce a comprehensive, transparent, and trustworthy benchmarking framework designed to eliminate these ambiguities. Aligned with the core principles of FAIR (Findable, Accessible, Interoperable, and Reusable) and open science, our framework provides a unified pipeline for the Earth Observation community.

GitHub repo for the benchmark: FAIR-EO-CD-benchmark

GitHub repo for AiTLAS: AiTLAS Toolbox


Benchmark Grid

Checkpoints and training curves for a change detection benchmark across 10 datasets, 10 model architectures, and 2 initialisations/training types — 180 runs.

(Note: ChangeFormerV6 and CSSM are released scratch-only; model weights pre-trained on ImageNet-1K were not available, so the grid is 180 runs rather than a full 200).

Axis Values
Datasets (10) BANDON, CLCD, DSIFN, EGY-BCD, LEVIR-CD+, MSBC, MSOSCD, OMBRIA, Season-varying CDD, SYSU-CD
Models (10) BIT, CGNet, CSSM, ChangeFormerV6, ChangeViT, HRNet SiamConc, SiamCRNN, STANet, TinyCD, U-Net SiamConc
Initialisation (2) pretrained_imagenet1K, scratch

Logged metrics per run: Loss/train, Loss/val, Accuracy/val, IOU/val, IOU mean/val, and micro/macro/weighted plus per-class (no change, change) Precision, Recall, F1_score.


Repository Layout

checkpoints/
  best_checkpoint_<dataset>_<init>_<model>.pth.tar     # 180 files
runs/
  <dataset>_<init>_<model>/
    tensorboard_<dataset>_<init>_<model>.tfevents.0    # 180 files
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support