Layer42
All projects

EO-VLM: Benchmarking Vision-Language Models for Earth Observation

A benchmarking framework for evaluating vision-language models on Earth observation tasks using satellite and UAV imagery.

Type
Research
Status
Ongoing
Developed by
Bouamer Tarek
Period
Since 2026
EO-VLM overview figure: a satellite and a drone capture imagery of the Earth, which vision-language models are evaluated on across tasks including visual question answering, classification, counting, spatial reasoning, temporal analysis, and change detection.

Objective

Measure how well current vision-language models handle Earth observation imagery from satellites and UAVs, using one unified, reproducible evaluation pipeline.

Contribution

EO-VLM provides command-line tools for single-image evaluation, temporal (multi-frame) evaluation, and inspection of a model's predictions on individual images. It covers visual question answering, classification, counting, spatial reasoning, temporal analysis, and change detection, and evaluates models on GEOBench-VLM, a third-party benchmark of satellite and aerial imagery.

Current progress

First released in 2026. EO-VLM currently supports one benchmark dataset, GEOBench-VLM, and the project page reports results for 12 models.

Evaluation

EO-VLM reports results on GEOBench-VLM for single-image and temporal evaluation, as overall accuracy: the percentage of the benchmark's questions answered correctly in each mode. On the project page, LLaVA OneVision 7B has the highest single-image accuracy, 50.20%, followed by Qwen2-VL 7B Instruct at 49.11%. In temporal evaluation, Qwen2-VL 7B Instruct has the highest accuracy, 53.12%, followed by InternVL2 8B at 47.75%.

The two modes use different question sets (16,055 single-image and 8,565 temporal questions), so their scores are not directly comparable. Single-image results cover all 12 models and temporal results cover 10. The project page provides model comparisons and task-level scores for these evaluation modes.

Resources