🔭 I’m currently working on video understanding and large multimodal models
📫 How to reach me: shehan.munasinghe@mbzuai.ac.ae
🔭 I’m currently working on video understanding and large multimodal models
📫 How to reach me: shehan.munasinghe@mbzuai.ac.ae
[CVPR 2025 🔥]A Large Multimodal Model for Pixel-Level Visual Grounding in Videos
PG-Video-LLaVA: Pixel Grounding in Large Multimodal Video Models
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.