Home

Nhat's work spans embodied AI, robotic manipulation, video understanding, autonomous driving perception, and trustworthy computer vision.

Research Papers

  1. Camouflaged object feature visualizations
    2026

    Catch Me If You Can Describe Me: Open-Vocabulary Camouflaged Instance Segmentation with Diffusion

    TL;DR: Uses open-vocabulary diffusion features to find and segment camouflaged objects, including object categories unseen during training.

    Tuan-Anh Vu, Duc Thanh Nguyen, Qing Guo, Nhat Chung, Binh-Son Hua, Ivor W. Tsang, Sai-Kit Yeung

    International Journal of Computer Vision (IJCV), 2026

  2. MAGIC collaborative-agent framework for generating physical adversarial patches
    2026

    MAGIC: Mastering Physical Adversarial Generation in Context Through Collaborative LLM Agents

    TL;DR: Coordinates three multimodal LLM agents to design, place, and refine context-aware physical patches that fool driving object detectors.

    Yun Xing, Nhat Chung, Jie Zhang, Yue Cao, Ivor Tsang, Yang Liu, Lei Ma, Qing Guo

    AAAI 2026

  3. 2026

    Rethinking Progression of Memory State in Robotic Manipulation: An Object-Centric Perspective

    TL;DR: Introduces LIBERO-Mem and a slot-based VLA that tracks object histories efficiently for object-level non-Markovian manipulation.

    Nhat Chung*, Taisei Hanyu*, Toan Nguyen, Huy Le, Frederick Bumgarner, Duy Minh Ho Nguyen, Khoa Vo, Kashu Yamazaki, Chase Rainwater, Tung Kieu, Anh Nguyen, Ngan Le

    AAAI 2026

  4. 2026

    SlotVLA: Towards Modeling of Object-Relation Representations in Robotic Manipulation

    TL;DR: Represents objects and their relations as compact slots, enabling interpretable robot control with far fewer visual tokens.

    Taisei Hanyu*, Nhat Chung*, Huy Le, Toan Nguyen, Yuki Ikebe, Anthony Gunderman, Duy Minh Ho Nguyen, Khoa Vo, Tung Kieu, Kashu Yamazaki, Chase Rainwater, Anh Nguyen, Ngan Le

    ICRA 2026

  5. 2026

    UNO: Unifying One-stage Video Scene Graph Generation via Object-Centric Visual Representation Learning

    TL;DR: Unifies box- and pixel-level video scene graph generation in one efficient, object-centric model without explicit tracking.

    Huy Le, Nhat Chung, Tung Kieu, Jingkang Yang, Ngan Le

    WACV 2026

  6. OBEYED-VLA object-centric and geometry-grounded manipulation framework
    2025

    Clutter-Resistant Vision-Language-Action Models through Object-Centric and Geometry Grounding

    TL;DR: Grounds robot observations in task-relevant objects and 3D geometry so VLA policies resist clutter, distractors, and background changes.

    Khoa Vo, Taisei Hanyu, Yuki Ikebe, Trong Thang Pham, Nhat Chung, Minh Nhat Vu, Duy Nguyen Ho Minh, Anh Nguyen, Anthony Gunderman, Chase Rainwater, Ngan Le

    arXiv Preprint, 2025

  7. DepthVanish physical adversarial patch framework for stereo depth estimation
    2025

    DepthVanish: Optimizing Adversarial Interval Structures for Stereo-Depth-Invisible Patches

    TL;DR: Jointly optimizes a patch's texture and grid spacing to expose real-world vulnerabilities in stereo depth systems.

    Yun Xing, Yue Cao, Nhat Chung, Jie M. Zhang, Ivor Tsang, Ming-Ming Cheng, Yang Liu, Lei Ma, Qing Guo

    NeurIPS 2025

  8. 2025

    BiMa: Towards Biases Mitigation for Text-Video Retrieval via Scene Element Guidance

    TL;DR: Uses scene elements to reduce visual and linguistic bias, improving fine-grained text-video retrieval and out-of-distribution robustness.

    Huy Le, Nhat Chung, Tung Kieu, Anh Nguyen, Ngan Le

    ACM MM 2025

  9. Examples of typographic attacks against driving vision-language models
    2024

    Towards Transferable Attacks Against Vision-LLMs in Autonomous Driving with Typography

    TL;DR: Shows that realistic roadside text can transfer across vision-language models and mislead autonomous-driving reasoning tasks.

    Nhat Chung, Sensen Gao, Tuan-Anh Vu, Jie Zhang, Aishan Liu, Yun Lin, Jin Song Dong, Qing Guo

    arXiv Preprint, 2024

  10. BgSubNet training and inference framework
    2024

    BgSubNet: Robust Semi-Supervised Background Subtraction in Realistic Scenes

    TL;DR: Combines synthetic data, self-supervised pretraining, and mixed-domain learning for background subtraction in unseen scenes.

    Nhat Minh Chung, Synh Viet-Uyen Ha

    IEEE Sensors Journal, 2024

  11. 2023

    Tracked-Vehicle Retrieval by Natural Language Descriptions With Multi-Contextual Adaptive Knowledge

    TL;DR: Combines CLIP, semi-supervised domain adaptation, and contextual pruning to retrieve vehicles from natural-language descriptions.

    Huy Le, Quang Qui-Vinh Nguyen, Duc Trung Luu, Truc Chau, Nhat Chung, Synh Ha

    CVPR Workshop 2023 - AI City Challenge Track 2 Winner

  12. 2023

    Multi-camera People Tracking With Mixture of Realistic and Synthetic Knowledge

    TL;DR: Tracks people across cameras by correcting ID switches and transferring synthetic-data knowledge to real indoor scenes.

    Quang Qui-Vinh Nguyen, Huy Le, Truc Chau, Duc Trung Luu, Nhat Chung, Synh Ha

    CVPR Workshop 2023 - AI City Challenge Track 1 Runner-up

  13. 2022

    Multi-Camera Multi-Vehicle Tracking with Domain Generalization and Contextual Constraints

    TL;DR: Improves multi-camera vehicle tracking with motion cues, synthetic-data domain generalization, and contextual matching constraints.

    Nhat Chung, Huy Le, Vuong Nguyen, Quang Qui-Vinh Nguyen, Thong Nguyen, Tin Thai, Synh Ha

    CVPR Workshop 2022

  14. 2022

    Tracked-vehicle Retrieval by Natural Language Descriptions with Domain Adaptive Knowledge

    TL;DR: Adapts CLIP to unseen camera domains with pseudo-labeling and pruning for language-based vehicle retrieval.

    Huy Le*, Quang Qui-Vinh Nguyen*, Vuong Nguyen, Thong Nguyen, Nhat Chung, Tin Thai, Synh Ha

    CVPR Workshop 2022

  15. 2021

    Tiny-PIRATE: A Tiny Model With Parallelized Intelligence for Real-Time Analysis as a Traffic countEr

    TL;DR: Delivers accurate real-time vehicle counting on resource-limited IoT devices with a lightweight detect-track-count pipeline.

    Synh Ha, Nhat Chung, Tien-Cuong Nguyen, Hung Ngoc Phan

    CVPR Workshop 2021 - AI City Challenge Track 1 Runner-up

* denotes equal contribution.