ViSTA: Visual Storytelling using Multi-modal Adapters for Text-to-Image Diffusion Models

Sibo Dong, Ismail Shaheen, Maggie Shen, Rupayan Mallick, and Sarah Adel Bargal

IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2026 ยท Oral