Research
How can intelligent systems understand and interact with the world?
Computer Vision & Spatial Intelligence
We study how systems recognize places, estimate where an image was taken, and reconstruct three-dimensional scenes from images and video.
- Place recognition: image retrieval across changes in viewpoint, season, and lighting.
- Feature matching and localization: learned local features and camera pose estimation.
- Reconstruction and mapping: structure-from-motion, SLAM, and feed-forward 3D reconstruction.
- Scene representations: Gaussian Splatting and compact 3D models.
Research question: How can spatial representations built from images stay accurate and compact as viewpoints and conditions change?
Related work
- VisLoc(Research · Ongoing)

Multimodal Learning & Reasoning
We study how models connect images, video, language, and 3D information, and how reliably they reason about what they observe.
- Model adaptation: adapting language and vision-language models to new domains.
- Evaluation: measuring what these models understand, and where they fail.
- Cross-modal representations: linking images, video, language, and 3D data.
- Spatial and temporal reasoning: reasoning about layout, geometry, and change over time.
Research question: When does a model genuinely ground language in what it observes, and how can that be measured?
Related work
- EO-VLM(Research · Ongoing)

Robotics & Autonomous Systems
We study embodied intelligence, sometimes called physical AI: how robots learn, plan, and act in the physical world.
- Robot learning: imitation and reinforcement learning, and vision-language-action models.
- Planning and control: motion planning and model predictive control for navigation and manipulation.
- World models: predicting how the environment responds to a robot's actions.
- Simulation to reality: testing whether behavior learned in simulation holds on real robots.
Research question: How can learned policies and world models work with planning and control to act reliably beyond their training conditions?

Earth Observation & Remote Sensing
An applied direction that brings representation learning, multimodal models, and efficient inference to satellite and aerial observations.
- Representation learning: self-supervised and foundation models for satellite and aerial imagery.
- Remote-sensing vision-language models: describing and answering questions about Earth observation imagery.
- Change analysis: detecting change across repeated observations of the same area.
- Generalization: robustness across regions, sensors, and acquisition conditions.
Research question: How can Earth observation models generalize across regions and sensors while running efficiently near where data is acquired?
Related work
- EO-VLM(Research · Ongoing)

Have a related research question?
We welcome conversations about shared research interests and potential collaborations.