Layer42
All projects

VisLoc: Image Retrieval, Matching, and Localization

A collection of models and algorithms for image retrieval, image matching, and image-to-3D localization in XR and robotics.

Type
Research
Status
Ongoing
Developed by
Bouamer Tarek
Period
Since 2022
VisLoc logo: a gradient loop in blue and teal above the word VISLOC.

Objective

Estimate where an image was taken within a 3D map by chaining three steps: find similar mapped images, match local features, and solve for the camera pose. VisLoc targets extended reality (XR) and robotics.

Contribution

Three libraries, one per step:

  • Retrieval: global-descriptor models (GeM and HOW) for image retrieval and place recognition, with ASMK aggregation of local features.
  • Matching (ImMatch): a shared interface for feature extraction, image matching, and geometric estimation, including methods such as SuperPoint, DISK, XFeat, LightGlue, SuperGlue, and LoFTR.
  • Localization: image-to-3D localization and absolute pose estimation, combining retrieval and matching with PyColmap or PoseLib.
A 3D point-cloud map of the city of Aachen, seen from above at an angle, with camera positions marked in red.
Aachen map from the visloc_localization repository.

Current progress

The retrieval and localization libraries were started in 2022, and ImMatch in 2024. Open items in the repositories include knowledge distillation into smaller retrieval models and a localization visualization.

Evaluation

The localization repository reports results on the Aachen Day-Night benchmark. See the repository for the evaluated configurations and reported results.

The retrieval repository reports results on ROxford5k and RParis6k as mean average precision (mAP, in %), for both single-scale and multi-scale descriptors. With multi-scale global descriptors (images at scales 0.7071, 1.0, and 1.4142), a ResNet-50 GeM model trained on Google Landmarks 2018 reaches 68.05 (Medium) and 43.42 (Hard) on ROxford5k, and 79.75 (Medium) and 61.14 (Hard) on RParis6k. Medium counts both easy and hard ground-truth matches as relevant; Hard counts only the hard ones.

The localization table notes that one comparison is still to be re-run.

Resources