Finetuning a T2I diffusion model's text encoder for strong attention localization of part-level concepts (e.g., 'left-front leg'); this leads to accurate part-level instance segmentation and generative control.
Hi! I am Vaibhav, a visiting student with Adam Kortylewski. My research is centered around models that can imagine the physical world (e.g., generative models) and how this capability can be used for visual understanding and embodied intelligence. In the past, I have worked on object-level 3D control in text-to-image generative models.
Previously, I spent beautiful years at CVIT, IIIT Hyderabad, where I was fortunate to be advised by Ravi Kiran S and co-advised by Venkatesh Babu R (from IISc Bengaluru), working closely with Rishubh Parihar.
I spend most of my time either thinking about research or doing experiments; some of that work can be found on my publications and blogs. When I am not working, I am usually listening to music or playing the piano. I love receiving emails! Hence, feel free to reach out to discuss about anything.
* denotes equal contribution, † denotes equal advising.
Finetuning a T2I diffusion model's text encoder for strong attention localization of part-level concepts (e.g., 'left-front leg'); this leads to accurate part-level instance segmentation and generative control.
A 3D bounding box layout consisting of translucent boxes effectively models occluding scene regions; this representation is used to spatially condition a T2I model to enable occlusion aware 3D control.
A continuous textual token is learnt to model generalized object orientation, and multiple such tokens can be combined to enable disentangled multi-object control, using attention based disentanglement.
A line-based parametrization over pixel-wise segmentation enables highly accurate text-line prediction. Further, a context-adaptive patching scheme enables generalization to arbitrary documents.
A multi-modal (visual + depth) re-identification improves SLAM performance by effectively leveraging visual cues.
Coming soon.
Coming soon.