SkyScraper: A Multi-Agent Feedback System for Detecting and Describing News Events in Satellite Imagery

1MIT 2Planet Labs
*Correspondence to mloui [at] mit.edu.
SkyScraper leverages multi-modal LLMs to detect multi-temporal Earth-observation imagery corresponding with text derived from news articles, enabling detailed captioning of diverse remote sensing events.

Abstract

Changes in satellite imagery often occur over multiple time steps. Despite the emergence of bi-temporal change captioning datasets, there is a lack of multi-temporal event captioning datasets (at least two images per sequence) in remote sensing. This gap exists because (1) searching for visible events in satellite imagery and (2) labeling multi-temporal sequences require significant time and labor. To address these challenges, we present SkyScraper, an iterative multi-agent workflow that geocodes news articles and synthesizes captions for corresponding satellite image sequences. Our experiments show that SkyScraper successfully finds 5x more events than traditional geocoding methods, demonstrating that agentic feedback is an effective strategy for surfacing new multi-temporal events in satellite imagery. We apply our framework to a large database of global news articles, curating a new multi-temporal captioning dataset with 5,000 sequences. By automatically identifying imagery related to news events, our work also supports journalism and reporting efforts.

Motivation

Challenges:

  • Identifying and labeling events in multi-temporal satellite imagery require significant time and labor.
  • Prior methods typically convert existing change labels into captions, which limits their scalability, sequence length, event diversity, and geographic coverage [1] [2].

Goals:

  • Develop an automated data curation framework for geocoding news articles and captioning corresponding multi-temporal satellite imagery.
  • Evaluate agent-based geocoding against traditional methods.
  • Curate a new global, multi-temporal satellite event dataset grounded in diverse news sources.

Approach

We implement agentic iterative feedback with the following steps:

  1. Extract location name and event timeline from the article text.
  2. Geocode the location name into latitude-longitude coordinates.
  3. Fetch satellite imagery at the coordinates over the timeline dates.
  4. Verify the event visibility by cross-referencing the article.
  5. Caption the image sequence using the article as context.

Experiments

We compare our method against two geocoding baselines: (1) weighted centroid and (2) GIPSY [3].

Weighted Centroid

GIPSY

We applied SkyScraper to an initial set of 1,000 news articles. After manual validation, our method outperforms both baselines, improving weighted centroid by nearly 5 times. Here, yield represents the percentage of correct detections (true positives) out of the initial set of articles.

Dataset

We applied SkyScraper to news articles sampled from the Global Database of Events, Language, and Tone (GDELT) [4] from 2022–2024 using PlanetScope and Sentinel-2 imagery. Annotators verified captions and event dates to produce the final datasets. Each dataset includes around 5,000 total image sequences with about 3,000 captioned visible events, with the remaining as negative examples.

Below are several examples of captioned multi-temporal sequences corresponding with news articles from our dataset.

Conclusion

We introduce SkyScraper, a novel multi-agent feedback system that locates and captions events in multi-temporal satellite imagery using news articles. Compared to traditional geocoding methods, our approach increases event detections by nearly 5x. We apply SkyScraper to the GDELT database to produce new PlanetScope and Sentinel-2 multi-temporal captioning datasets with about 5,000 sequences. These results demonstrate the effectiveness of agentic feedback for facilitating event geocoding, multi-image captioning, and benchmark dataset curation in the remote sensing domain.

References

  1. Karaca, Ali Can, et al. "Robust change captioning in remote sensing: Second-cc dataset and mmodalcc framework." IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing (2025).
  2. Irvin, Jeremy, et al. "Teochat: A large vision-language assistant for temporal earth observation data." International Conference on Learning Representations. Vol. 2025. 2025.
  3. Woodruff, Allison Gyle, and Christian Plaunt. "Gipsy: Georeferenced information processing system,"." Journal of the American Society for Information Science 45.9 (1994): 645-655.
  4. https://www.gdeltproject.org/

BibTeX

If you find this work useful, please cite the following:

@article{anderson2026multi,
  title={A Multi-Agent Feedback System for Detecting and Describing News Events in Satellite Imagery},
  author={Anderson, Madeline and Klassen, Mikhail and Hoover, Ash and Cahoy, Kerri},
  journal={arXiv preprint arXiv:2604.12772},
  year={2026}
}