MVDGC: Joint 3D and 2D Multi-view Pedestrian Detection via Dual Geometric Constraints
Dual-State Slot Attention: Decoupling Appearance and Identity for Video Object-Centric Learning
Gripper-aware Vision Language Action Models
CodeGraphVLP: Code-as-Planner Meets Semantic-Graph State for Non-Markovian Vision-Language-Action Models
SemLT3D: Semantic-Guided Expert Distillation for Camera-only Long-Tailed 3D Object Detection
SlotVLA: Towards Modeling of Object-Relation Representations in Robotic Manipulation
Clutter-Robust Vision-Language-Action Models through Object-Centric and Geometry Grounding
Rethinking Progression of Memory State in Robotic Manipulation: An Object-Centric Perspective
Dr. Ngan Le has received a prestigious 2025 NSF Faculty Early Career Development (CAREER) Award
CT-ScanGaze: A Dataset and Baselines for 3D Volumetric Scanpath Modeling