EO-VLM: Benchmarking Vision-Language Models for Earth Observation
A benchmarking framework for evaluating vision-language models on Earth observation tasks using satellite and UAV imagery.
- Type
- Research
- Status
- Ongoing
- Developed by
- Bouamer Tarek
- Period
- Since 2026

Objective
Measure how well current vision-language models handle Earth observation imagery from satellites and UAVs, using one unified, reproducible evaluation pipeline.
Contribution
EO-VLM provides command-line tools for single-image evaluation, temporal (multi-frame) evaluation, and inspection of a model's predictions on individual images. It covers visual question answering, classification, counting, spatial reasoning, temporal analysis, and change detection, and evaluates models on GEOBench-VLM, a third-party benchmark of satellite and aerial imagery.
Current progress
First released in 2026. EO-VLM currently supports one benchmark dataset, GEOBench-VLM, and the project page reports results for 12 models.
Evaluation
EO-VLM reports results on GEOBench-VLM for single-image and temporal evaluation, as overall accuracy: the percentage of the benchmark's questions answered correctly in each mode. On the project page, LLaVA OneVision 7B has the highest single-image accuracy, 50.20%, followed by Qwen2-VL 7B Instruct at 49.11%. In temporal evaluation, Qwen2-VL 7B Instruct has the highest accuracy, 53.12%, followed by InternVL2 8B at 47.75%.
The two modes use different question sets (16,055 single-image and 8,565 temporal questions), so their scores are not directly comparable. Single-image results cover all 12 models and temporal results cover 10. The project page provides model comparisons and task-level scores for these evaluation modes.