Human-Centric Spatio-Temporal Video Grounding With Visual Transformers

In this work, we introduce a novel task - Human-centric Spatio-Temporal Video Grounding (HC-STVG). Unlike the existing referring expression tasks in images or videos, by focusing on humans, HC-STVG aims to localize a spatio-temporal tube of the target person from an untrimmed video based on a given...

Full description

Saved in:
Bibliographic Details
Published inIEEE transactions on circuits and systems for video technology Vol. 32; no. 12; pp. 8238 - 8249
Main Authors Tang, Zongheng, Liao, Yue, Liu, Si, Li, Guanbin, Jin, Xiaojie, Jiang, Hongxu, Yu, Qian, Xu, Dong
Format Journal Article
LanguageEnglish
Published New York IEEE 01.12.2022
The Institute of Electrical and Electronics Engineers, Inc. (IEEE)
Subjects
Online AccessGet full text

Cover

Loading…