360° Image Perception with MLLMs: A Comprehensive Benchmark and a Training-Free Method
A comprehensive benchmark and a training-free method for 360° image perception using MLLMs.
Hi there! I’m an AI researcher, doing research on Computer Vision, Multimodal AI for vision-language understanding, and their AI applications for domain-specific problems. I am passionate about bridging the gap between foundational AI research and impactful technology with research papers published in international venues like CVPR, ECCV, EACL, and IJCAI.
I’m currently working at the University of Wollongong, Australia. Previously, I had a 3-year postdoctoral fellowship at RIKEN AIP and a visiting research position at Tohoku University, Japan.
Feel free to reach out for collaborations!
A comprehensive benchmark and a training-free method for 360° image perception using MLLMs.
Improving multimodal table understanding with code-driven reasoning.
Large language models deliver expert-level assessments for landslide image analysis and hazard response.
Studies multimodal perception models for anticipating hazardous events in autonomous driving scenarios.
Presents GRIT, a dual-feature transformer that improves both speed and accuracy for image captioning.
Enhances interactive instruction following agents with wide-context perception and iterative reasoning.
Introduces an efficient attention design capturing full interactions in visual dialog systems.
Applies capsule networks to the challenging task of recognizing subtle micro-expressions.