Geometrically grounded change detection
Argos adapts implicit 3D priors through co-visibility-aware inter-session comparison for joint change detection and reconstruction from a pair of RGB sequences.
Preprint · 2026
1MIT 2Gwangju Institute of Science and Technology
Robots operating in dynamic environments require reliable detection of how their surroundings change over time. Existing learning-based methods largely rely on pairwise 2D image features, which struggle under large viewpoint changes and occlusions, are sensitive to noise, and show limited generalization across domains, while explicit 3D approaches typically require costly offline optimization. We show that the implicit 3D knowledge of Geometric Foundation Models (GFMs) provides a strong basis for addressing these limitations. We introduce Argos, which adapts GFM features for joint scene change detection and 3D reconstruction. To address data scarcity and take a step toward a foundation model for scene change detection, we introduce a large-scale benchmark comprising two synthetic datasets and one real-world dataset, and train jointly across diverse datasets to improve cross-domain generalization. We further introduce Argos-SLAM, a real-time system designed for robotics, which performs online change detection and change-aware 4D mapping. Across benchmarks, our framework substantially outperforms existing baselines, with gains of up to 42.01% in change IoU and 27.91% in F1, while supporting scalable deployment in changing real-world environments.
Argos adapts implicit 3D priors through co-visibility-aware inter-session comparison for joint change detection and reconstruction from a pair of RGB sequences.
The Argos-CD dataset, combined with pretrained GFM features and multi-dataset training, enables strong benchmark performance and synthetic-to-real transfer.
Argos-SLAM connects learned change predictions to online reconstruction of a change-aware point-cloud map.
Our key insight is that the geometric knowledge a GFM learns to associate scene content across viewpoints also provides a prior for recognizing inconsistent content across sessions. Crucially, inconsistency must be interpreted through co-visibility: observing that the space previously occupied by an object is now empty is evidence of removal, whereas simply failing to see an occluded object is not.
| Method | SceneDiff | Argos-CD-real | ||||||
|---|---|---|---|---|---|---|---|---|
| Static | Change | mIoU | F1 | Static | Change | mIoU | F1 | |
| CSCDNet-finetune | 93.33 | 15.25 | 54.29 | 20.23 | 94.75 | 8.88 | 51.81 | 14.20 |
| C-3PO*-finetune | 94.42 | 17.88 | 56.15 | 22.47 | 95.49 | 8.71 | 52.10 | 12.41 |
| GeSCF† | 69.69 | 5.04 | 37.37 | 15.12 | 63.64 | 2.95 | 33.30 | 11.13 |
| RobustSCD-finetune | 94.41 | 18.97 | 56.69 | 21.25 | 94.46 | 9.06 | 51.76 | 15.92 |
| ZSSCD† | 88.63 | 6.74 | 47.69 | 13.22 | 79.97 | 3.79 | 41.88 | 11.47 |
| TERDNet-finetune | 94.30 | 20.32 | 57.30 | 25.14 | 93.90 | 8.56 | 51.23 | 16.75 |
| VSCDNet | 91.40 | 10.48 | 50.94 | 15.46 | 91.06 | 7.26 | 49.16 | 12.91 |
| VSCDNet-finetune | 95.72 | 20.70 | 58.21 | 23.84 | 96.76 | 11.72 | 54.24 | 17.88 |
| VSCDNet-foundation | 95.85 | 19.27 | 57.56 | 24.26 | 93.75 | 9.54 | 51.65 | 19.77 |
| SceneDiff† | 96.48 | 30.31 | 63.40 | 37.58 | 95.50 | 19.49 | 57.50 | 36.90 |
| Ours-synthetic-only | 97.13 | 39.84 | 68.48 | 37.42 | 95.88 | 20.87 | 58.37 | 34.17 |
| Ours-finetune | 98.62 | 59.91 | 79.27 | 44.73 | 97.69 | 36.29 | 66.99 | 39.38 |
| Ours-foundation | 98.78 | 62.54 | 80.66 | 45.98 | 97.36 | 29.50 | 63.43 | 39.90 |
| Method | Argos-CD-HSSD | Argos-CD-SceneSmith | VSCD** | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Static | Change | mIoU | F1 | Static | Change | mIoU | F1 | Static | Change | mIoU | F1 | |
| CSCDNet | 92.91 | 38.10 | 65.51 | 36.53 | 94.32 | 31.11 | 62.72 | 34.83 | 94.03 | 31.95 | 62.99 | 32.44 |
| C-3PO* | 91.85 | 37.65 | 64.75 | 37.68 | 94.46 | 38.68 | 66.57 | 40.77 | 93.88 | 32.98 | 63.43 | 32.90 |
| GeSCF† | 76.73 | 13.81 | 45.27 | 28.41 | 81.56 | 13.72 | 47.64 | 29.80 | 77.74 | 12.26 | 45.00 | 23.98 |
| RobustSCD | 88.86 | 33.18 | 61.02 | 34.76 | 92.70 | 33.92 | 63.30 | 36.85 | 89.54 | 27.18 | 58.36 | 30.59 |
| ZSSCD† | 90.03 | 16.58 | 53.31 | 16.22 | 91.63 | 15.15 | 53.39 | 15.49 | 93.00 | 15.72 | 54.36 | 16.02 |
| TERDNet | 93.05 | 42.14 | 67.59 | 40.62 | 95.19 | 38.58 | 66.88 | 40.31 | 90.90 | 27.92 | 59.41 | 30.92 |
| VSCDNet | 93.50 | 42.71 | 68.11 | 41.77 | 94.62 | 38.94 | 66.78 | 41.17 | 94.49 | 34.55 | 64.52 | 36.31 |
| SceneDiff† | 93.80 | 33.44 | 63.62 | 32.44 | 96.38 | 40.42 | 68.40 | 36.58 | 96.71 | 48.21 | 72.46 | 38.70 |
| Ours | 98.38 | 78.64 | 88.51 | 62.16 | 99.08 | 82.43 | 90.75 | 69.08 | 98.24 | 68.69 | 83.46 | 53.14 |
| Ours-foundation | 98.07 | 75.51 | 86.79 | 60.69 | 98.91 | 80.41 | 89.66 | 68.93 | 98.05 | 68.52 | 83.28 | 53.33 |
Argos-SLAM integrates Argos with VGGT-SLAM for online reconstruction and change updates. We evaluate it on offline human-captured iPhone videos and online deployments on an Agilex mobile manipulator with an Intel RealSense D455, across small, medium and multi-room indoor scenes with 4–10 object changes. During online deployment the SLAM loop produces a new submap every 8.2 s on average, while the asynchronous change-detection job finishes in 6.9 s, so change updates keep pace with reconstruction.
@article{xu2026argos,
title = {Argos: Adapt Rich Geometric Priors for Generalizable
Online Scene-Change-Detection},
author = {Xu, Ruihan and Yoon, Jiae and Zhou, Kaichen and
Kim, Ue-Hwan and Carlone, Luca},
journal = {arXiv preprint},
year = {2026}
}