Preprint · 2026

Argos: Adapt Rich Geometric Priors for Generalizable Online Scene-Change-Detection

Ruihan Xu1 Jiae Yoon2 Kaichen Zhou1 Ue-Hwan Kim2 Luca Carlone1

1MIT    2Gwangju Institute of Science and Technology

Argos takes two unposed RGB visits of an office and produces a joint 3D reconstruction, highlighting removed scissors, diving fins replaced by a bag, an added guitar and a removed can.
We propose Argos, a feedforward model for sub-second change detection and 3D reconstruction, and integrate it into Argos-SLAM for online 4D reconstruction of change-aware point-cloud maps. Argos captures scene changes such as the added guitar (3) and the removed can (4), while keeping a consistent spatio-temporal representation across visits.

Abstract

Robots operating in dynamic environments require reliable detection of how their surroundings change over time. Existing learning-based methods largely rely on pairwise 2D image features, which struggle under large viewpoint changes and occlusions, are sensitive to noise, and show limited generalization across domains, while explicit 3D approaches typically require costly offline optimization. We show that the implicit 3D knowledge of Geometric Foundation Models (GFMs) provides a strong basis for addressing these limitations. We introduce Argos, which adapts GFM features for joint scene change detection and 3D reconstruction. To address data scarcity and take a step toward a foundation model for scene change detection, we introduce a large-scale benchmark comprising two synthetic datasets and one real-world dataset, and train jointly across diverse datasets to improve cross-domain generalization. We further introduce Argos-SLAM, a real-time system designed for robotics, which performs online change detection and change-aware 4D mapping. Across benchmarks, our framework substantially outperforms existing baselines, with gains of up to 42.01% in change IoU and 27.91% in F1, while supporting scalable deployment in changing real-world environments.

Contributions

Geometrically grounded change detection

Argos adapts implicit 3D priors through co-visibility-aware inter-session comparison for joint change detection and reconstruction from a pair of RGB sequences.

Large-scale dataset and generalization

The Argos-CD dataset, combined with pretrained GFM features and multi-dataset training, enables strong benchmark performance and synthetic-to-real transfer.

Online change-aware mapping

Argos-SLAM connects learned change predictions to online reconstruction of a change-aware point-cloud map.

Method

Our key insight is that the geometric knowledge a GFM learns to associate scene content across viewpoints also provides a prior for recognizing inconsistent content across sessions. Crucially, inconsistency must be interpreted through co-visibility: observing that the space previously occupied by an object is now empty is evidence of removal, whereas simply failing to see an occluded object is not.

Argos architecture: a frozen VGGT-Omega backbone and heads, plus a trainable co-visibility-aware spatio-temporal comparator that predicts per-visit change masks and a change-aware 4D map.
Building on VGGT-Ω, we introduce a Co-visibility-aware Spatio-Temporal Comparator with three modules: a Camera Reader that injects final camera tokens into intermediate patch features, Inter-session Cross-Attention that aggregates information across visits, and a lightweight DPT Fusion Head that combines multi-layer features to predict per-visit change masks. These masks then disentangle the joint reconstruction into a change-aware map.
Attention scores for a query token on an added bottle: high response in v1 frames where the bottle is present, low response in v0 frames.
GFM features already encode cross-visit correspondence. The token on the added green bottle in v1 attends strongly to the bottle region in other views of the same visit, and finds no high-attention match in v0, where the bottle is absent.

Results

+42.01change IoU over the best baseline on Argos-CD-SceneSmith (82.43 vs. 40.42)
62.54change IoU on real-world SceneDiff (Ours-foundation), vs. 30.31 for the best baseline
0.59 sper change-detection forward pass with up to 64 input images

Real-world generalization

IoU and F1 (%). Bold: best, underline: second best. † zero-shot. * C-3PO with the “remove” branch disabled. Even trained only on synthetic data, Argos beats learned baselines fine-tuned on real data.
MethodSceneDiffArgos-CD-real
StaticChangemIoUF1StaticChangemIoUF1
CSCDNet-finetune93.3315.2554.2920.2394.758.8851.8114.20
C-3PO*-finetune94.4217.8856.1522.4795.498.7152.1012.41
GeSCF†69.695.0437.3715.1263.642.9533.3011.13
RobustSCD-finetune94.4118.9756.6921.2594.469.0651.7615.92
ZSSCD†88.636.7447.6913.2279.973.7941.8811.47
TERDNet-finetune94.3020.3257.3025.1493.908.5651.2316.75
VSCDNet91.4010.4850.9415.4691.067.2649.1612.91
VSCDNet-finetune95.7220.7058.2123.8496.7611.7254.2417.88
VSCDNet-foundation95.8519.2757.5624.2693.759.5451.6519.77
SceneDiff†96.4830.3163.4037.5895.5019.4957.5036.90
Ours-synthetic-only97.1339.8468.4837.4295.8820.8758.3734.17
Ours-finetune98.6259.9179.2744.7397.6936.2966.9939.38
Ours-foundation98.7862.5480.6645.9897.3629.5063.4339.90

Video scene change detection

IoU and F1 (%). Bold: best, underline: second best. † zero-shot. * C-3PO with the “remove” branch disabled. ** VSCD relabeled with “appeared” labels in both reference and query frames.
MethodArgos-CD-HSSDArgos-CD-SceneSmithVSCD**
StaticChangemIoUF1StaticChangemIoUF1StaticChangemIoUF1
CSCDNet92.9138.1065.5136.5394.3231.1162.7234.8394.0331.9562.9932.44
C-3PO*91.8537.6564.7537.6894.4638.6866.5740.7793.8832.9863.4332.90
GeSCF†76.7313.8145.2728.4181.5613.7247.6429.8077.7412.2645.0023.98
RobustSCD88.8633.1861.0234.7692.7033.9263.3036.8589.5427.1858.3630.59
ZSSCD†90.0316.5853.3116.2291.6315.1553.3915.4993.0015.7254.3616.02
TERDNet93.0542.1467.5940.6295.1938.5866.8840.3190.9027.9259.4130.92
VSCDNet93.5042.7168.1141.7794.6238.9466.7841.1794.4934.5564.5236.31
SceneDiff†93.8033.4463.6232.4496.3840.4268.4036.5896.7148.2172.4638.70
Ours98.3878.6488.5162.1699.0882.4390.7569.0898.2468.6983.4653.14
Ours-foundation98.0775.5186.7960.6998.9180.4189.6668.9398.0568.5283.2853.33
Color legend for the qualitative results.
Qualitative change masks on a SceneDiff table scene. Qualitative change masks on a SceneDiff kitchen scene, comparing VSCDNet, SceneDiff and Argos with ground truth.
Qualitative results on real-world SceneDiff for the best-performing models. Argos produces cleaner masks with fewer false positives and false negatives.

Argos-SLAM on real robots

Argos-SLAM integrates Argos with VGGT-SLAM for online reconstruction and change updates. We evaluate it on offline human-captured iPhone videos and online deployments on an Agilex mobile manipulator with an Intel RealSense D455, across small, medium and multi-room indoor scenes with 4–10 object changes. During online deployment the SLAM loop produces a new submap every 8.2 s on average, while the asynchronous change-detection job finishes in 6.9 s, so change updates keep pace with reconstruction.

Change-aware 4D map of a multi-room office at visits v0 and v1, with nine numbered changes including a removed chair, an added honey bottle and a relocated Husky robot.
Argos-SLAM in a multi-room scene localizes additions and removals across object scales, such as a removed chair, an added honey bottle and a relocated Husky robot, consistently in both 2D observations and the global 4D map.

BibTeX

@article{xu2026argos,
  title   = {Argos: Adapt Rich Geometric Priors for Generalizable
             Online Scene-Change-Detection},
  author  = {Xu, Ruihan and Yoon, Jiae and Zhou, Kaichen and
             Kim, Ue-Hwan and Carlone, Luca},
  journal = {arXiv preprint},
  year    = {2026}
}