ViTCap-R: Unifying Image Captioning and Cross-Modal Retrieval with Vision Transformers and Patch-Level Attention
Abstract: Generating accurate image captions and retrieving semantically matched image-text pairs are typically treated as independent problems, limiting the depth of visual-semantic understanding a model can develop. This talk presents ViTCap-R, a unified framework that integrates a Vision Transformer encoder, a patch-level attention LSTM decoder, and a dual-encoder contrastive retrieval module into a single cohesive pipeline. By attending over 196 spatial patch embeddings during decoding, the model produces interpretable captions grounded in specific image regions. A shared 256-dimensional embedding space, trained with symmetric InfoNCE loss and hard negative mining, enables effective cross-modal retrieval and caption reranking at inference time. We evaluate ViTCap-R on Flickr8k and MS COCO 2014, reporting BLEU, METEOR, and Recall@K metrics alongside qualitative attention heatmaps, PCA projections, and t-SNE visualizations of the learned embedding space. Results demonstrate that unified training improves semantic alignment between modalities, while patch-level attention offers meaningful interpretability gains over standard global-feature baselines. We also discuss the challenges of scaling to high-diversity datasets and the trade-offs introduced by hard negative mining, pointing toward future directions involving transformer-based decoders and larger-scale pretraining. Speaker(s): Prayash Das Agenda: Hybrid event, in-person or online: 2:30 Introduction Technical Talks Question and Answer Networking 4:30 Conclusion Room: Meeting Rooms 2,3, 2 Civic Center Drive, East Brunswick, New Jersey, United States, 08816, Virtual: https://events.vtools.ieee.org/m/558867