Projects / Ongoing

Spatial Intelligence for Robotic Autonomy: From World Reasoning to Precision Control

Robotics

We develop spatial intelligence for robotic autonomy, focusing on how robots can reason over large-scale environments and execute robust actions in the physical world. Our research combines structured world representations, language-based reasoning, and geometry-aware robot learning to connect high-level task understanding with generalizable robot behavior.

We aim to connect reliable world reasoning with generalizable robot behavior through structured spatial representations and geometry-aware learning.

Symbolic World Models for Spatial Reasoning and Task Planning

Counterfactual feedback compares current hand pose with an expert pose.

Large-scale robotic tasks require more than recognizing individual objects. Robots need a structured understanding of objects, regions, states, and their spatial and semantic relationships to reason about the environment and plan sequences of actions.

We develop ontology-grounded symbolic world models that represent this information as knowledge graphs and continuously update them from robot observations and actions. An LLM-based agent interacts with the world model through explicit tools, retrieving relevant information through structured queries rather than directly processing the entire environment representation. The resulting world state can then be connected to symbolic planners to transform natural-language instructions into consistent and executable long-horizon task plans.

Geometry-Aware and Data-Efficient Robot Policy Learning

LLM copilot agent retrieves documents through RAG and queries tools through MCP.

Robot policies must remain reliable even when objects appear at new positions and orientations or when the robot encounters configurations not covered by its training data. We study geometry-aware Vision-Language-Action (VLA) models that explicitly incorporate three-dimensional spatial information and geometric structure.

In particular, we investigate SE(3)-equivariant robot learning to make policies respond consistently to spatial transformations. By combining 3D information with pretrained VLA models and exploiting geometric symmetries during learning, we aim to improve generalization while reducing the amount of task-specific robot data required. We further explore offline reinforcement learning to adapt these policies efficiently using previously collected demonstrations and interaction data.

Research Timeline

Year 1

We establish the foundations for symbolic world modeling and geometry-aware robot learning, including ontology-based world representations, 3D-aware VLA models, and SE(3)-equivariant learning methods.

Year 2

We develop LLM–world model interfaces for spatial reasoning and task planning and advance geometry-aware VLA policies for robust and data-efficient adaptation.

Year 3

We integrate and evaluate the developed world reasoning and robot policy learning methods on complex, long-horizon robotic tasks.

What Students May Work On

  • Ontology and knowledge-graph world models for robotics
  • LLM agents for spatial reasoning and task planning
  • Symbolic planning for long-horizon robotic tasks
  • 3D-aware Vision-Language-Action models
  • SE(3)-equivariant robot learning
  • Data-efficient robot policy learning and adaptation

Interested in joining this project? See how to apply →