Contrastive Video Question Answering via Video Graph Transformer

We propose to perform video question answering (VideoQA) in a Contrastive manner via a Video Graph Transformer model (CoVGT). CoVGT's uniqueness and superiority are three-fold: 1) It proposes a dynamic graph transformer module which encodes video by explicitly capturing the visual objects, thei...

Full description

Saved in:

Bibliographic Details
Main Authors	Xiao, Junbin, Zhou, Pan, Yao, Angela, Li, Yicong, Hong, Richang, Yan, Shuicheng, Chua, Tat-Seng
Format	Journal Article
Language	English
Published	27.02.2023
Subjects	Computer Science - Computer Vision and Pattern Recognition Computer Science - Multimedia
Online Access	Get full text

Cover

Loading…

Be the first to leave a comment!