A comprehensive and trustworthy benchmark of AI methods for change detection in Earth observation
Tadej Tomanič1, 2 · Alice Baudhuin1 · Jan Sotošek1 · Jure Brence1, 3 · Panče Panov1, 3 · Nikola Simidjievski1, 4 · Dragi Kocev1, 3
1Bias Variance Labs, d.o.o. 2University of Ljubljana, Faculty of Mathematics and Physics
3Department of Knowledge Technologies, Jožef Stefan Institute 4Télécom Paris, Institut Polytechnique de Paris
Summary
Companion artifacts for the paper, A comprehensive and trustworthy benchmark of AI methods for change detection in Earth observation (currently under review).
This study presents a standardized benchmark for change detection in Earth observation (EO), developed as part of the OSCARS-funded FAIR-EO project. The framework is integrated into the AiTLAS toolbox.
Despite rapid advancements in machine learning for remote sensing, accurately evaluating and comparing change detection models remains a significant challenge. Variations in dataset preprocessing, data splits, and evaluation metrics often lead to inconsistent results, making it difficult to determine whether a new method genuinely outperforms existing ones. To address this, we introduce a comprehensive, transparent, and trustworthy benchmarking framework designed to eliminate these ambiguities. Aligned with the core principles of FAIR (Findable, Accessible, Interoperable, and Reusable) and open science, our framework provides a unified pipeline for the Earth Observation community.
GitHub repo for the benchmark: FAIR-EO-CD-benchmark
GitHub repo for AiTLAS: AiTLAS Toolbox
Benchmark Grid
Checkpoints and training curves for a change detection benchmark across 10 datasets, 10 model architectures, and 2 initialisations/training types — 180 runs.
(Note: ChangeFormerV6 and CSSM are released scratch-only; model weights pre-trained on ImageNet-1K were not available, so the grid is 180 runs rather than a full 200).
| Axis | Values |
|---|---|
| Datasets (10) | BANDON, CLCD, DSIFN, EGY-BCD, LEVIR-CD+, MSBC, MSOSCD, OMBRIA, Season-varying CDD, SYSU-CD |
| Models (10) | BIT, CGNet, CSSM, ChangeFormerV6, ChangeViT, HRNet SiamConc, SiamCRNN, STANet, TinyCD, U-Net SiamConc |
| Initialisation (2) | pretrained_imagenet1K, scratch |
Logged metrics per run: Loss/train, Loss/val, Accuracy/val, IOU/val, IOU mean/val, and micro/macro/weighted plus per-class (no change, change) Precision, Recall, F1_score.
Repository Layout
checkpoints/
best_checkpoint_<dataset>_<init>_<model>.pth.tar # 180 files
runs/
<dataset>_<init>_<model>/
tensorboard_<dataset>_<init>_<model>.tfevents.0 # 180 files