VisLoc: Image Retrieval, Matching, and Localization
A collection of models and algorithms for image retrieval, image matching, and image-to-3D localization in XR and robotics.
- Type
- Research
- Status
- Ongoing
- Developed by
- Bouamer Tarek
- Period
- Since 2022

Objective
Estimate where an image was taken within a 3D map by chaining three steps: find similar mapped images, match local features, and solve for the camera pose. VisLoc targets extended reality (XR) and robotics.
Contribution
Three libraries, one per step:
- Retrieval: global-descriptor models (GeM and HOW) for image retrieval and place recognition, with ASMK aggregation of local features.
- Matching (ImMatch): a shared interface for feature extraction, image matching, and geometric estimation, including methods such as SuperPoint, DISK, XFeat, LightGlue, SuperGlue, and LoFTR.
- Localization: image-to-3D localization and absolute pose estimation, combining retrieval and matching with PyColmap or PoseLib.

Current progress
The retrieval and localization libraries were started in 2022, and ImMatch in 2024. Open items in the repositories include knowledge distillation into smaller retrieval models and a localization visualization.
Evaluation
The localization repository reports results on the Aachen Day-Night benchmark. See the repository for the evaluated configurations and reported results.
The retrieval repository reports results on ROxford5k and RParis6k as mean average precision (mAP, in %), for both single-scale and multi-scale descriptors. With multi-scale global descriptors (images at scales 0.7071, 1.0, and 1.4142), a ResNet-50 GeM model trained on Google Landmarks 2018 reaches 68.05 (Medium) and 43.42 (Hard) on ROxford5k, and 79.75 (Medium) and 61.14 (Hard) on RParis6k. Medium counts both easy and hard ground-truth matches as relevant; Hard counts only the hard ones.
The localization table notes that one comparison is still to be re-run.