P2ANet: A Large-Scale Benchmark for Dense Action Detection from Table Tennis Match Broadcasting Videos

While deep learning has been widely used for video analytics, such as video classification and action detection, dense action detection with fast-moving subjects from sports videos is still challenging. In this work, we release yet another sports video benchmark P2ANet for Ping Pong-Action detection...

Full description

Saved in:

Bibliographic Details
Published in	ACM transactions on multimedia computing communications and applications Vol. 20; no. 4; pp. 1 - 23
Main Authors	Bian, Jiang, Li, Xuhong, Wang, Tao, Wang, Qingzhong, Huang, Jun, Liu, Chen, Zhao, Jun, Lu, Feixiang, Dou, Dejing, Xiong, Haoyi
Format	Journal Article
Language	English
Published	New York, NY ACM 01.04.2024
Subjects	Activity recognition and understanding Computing methodologies Neural networks Datasets action recognition and localization table tennis annotation toolbox video analysis
Online Access	Get full text

Cover

Loading…

More Information
Summary:	While deep learning has been widely used for video analytics, such as video classification and action detection, dense action detection with fast-moving subjects from sports videos is still challenging. In this work, we release yet another sports video benchmark P2ANet for Ping Pong-Action detection, which consists of 2,721 video clips collected from the broadcasting videos of professional table tennis matches in World Table Tennis Championships and Olympiads. We work with a crew of table tennis professionals and referees on a specially designed annotation toolbox to obtain fine-grained action labels (in 14 classes) for every ping-pong action that appeared in the dataset, and formulate two sets of action detection problems—action localization and action recognition. We evaluate a number of commonly seen action recognition (e.g., TSM, TSN, Video SwinTransformer, and Slowfast) and action localization models (e.g., BSN, BSN++, BMN, TCANet), using P2ANet for both problems, under various settings. These models can only achieve 48% area under the AR-AN curve for localization and 82% top-one accuracy for recognition since the ping-pong actions are dense with fast-moving subjects but broadcasting videos are with only 25 FPS. The results confirm that P2ANet is still a challenging task and can be used as a special benchmark for dense action detection from videos. We invite readers to examine our dataset by visiting the following link: https://github.com/Fred1991/P2ANET.
ISSN:	1551-6857 1551-6865
DOI:	10.1145/3633516