Paper acceptedLIBERO-Mem, our work on stress-testing robots to remember object-level interaction histories for manipulation, was accepted to AAAI 2026.
Jan 2026
Paper acceptedSlotVLA, which gives robots compact object-relation representations for manipulation, was accepted to ICRA 2026. UNO, which unifies box- and pixel-level video scene graph generation in one object-centric model, was accepted to WACV 2026.
Nhat develops semantic representations that help robots connect perception, memory, and action.
SlotVLA
models objects and their relations for manipulation, while
LIBERO-Mem
maintains object-centric task states over long horizons, giving embodied agents structured knowledge for deciding what to do next.
Trustworthy Systems
Nhat studies whether semantic representations remain dependable against ambiguities and adversaries.
OBEYED-VLA
grounds actions in task-relevant objects and geometry, while
DepthVanish,
MAGIC, and
typographic attacks
expose how physical and linguistic perturbations can distort understanding and downstream decisions.
Visual Perception
Nhat's visual perception research turns complex scenes into semantic concepts that embodied agents can reason about.
Open-vocabulary camouflage segmentation
identifies concealed instances,
BgSubNet
separates foreground dynamics, and
UNO
represents interactions as scene relations when objects, motion, and context are difficult to interpret.
Nhat's research now asks how semantic representations can support decisions that remain meaningful, grounded, and robust in the open world. The long-term goal is to build contextually safe robots that understand what matters in a scene, recognize when their understanding may be unreliable, and act accordingly around people.