Predicting Goal-Directed Human Attention Using Inverse Reinforcement Learning

Human gaze behavior prediction is important for behavioral vision and for computer vision applications. Most models mainly focus on predicting free-viewing behavior using saliency maps, but do not generalize to goal-directed behavior, such as when a person searches for a visual target object. We pro...

Full description

Saved in:

Bibliographic Details
Published in	Proceedings (IEEE Computer Society Conference on Computer Vision and Pattern Recognition. Online) Vol. 2020; pp. 190 - 199
Main Authors	Yang, Zhibo, Huang, Lihan, Chen, Yupei, Wei, Zijun, Ahn, Seoyoung, Zelinsky, Gregory, Samaras, Dimitris, Hoai, Minh
Format	Conference Proceeding Journal Article
Language	English
Published	United States IEEE 01.06.2020
Subjects	Computational modeling Context modeling Learning (artificial intelligence) Predictive models Search problems Task analysis Visualization
Online Access	Get full text

Cover

Loading…

More Information
Summary:	Human gaze behavior prediction is important for behavioral vision and for computer vision applications. Most models mainly focus on predicting free-viewing behavior using saliency maps, but do not generalize to goal-directed behavior, such as when a person searches for a visual target object. We propose the first inverse reinforcement learning (IRL) model to learn the internal reward function and policy used by humans during visual search. We modeled the viewer's internal belief states as dynamic contextual belief maps of object locations. These maps were learned and then used to predict behavioral scanpaths for multiple target categories. To train and evaluate our IRL model we created COCO-Search18, which is now the largest dataset of high-quality search fixations in existence. COCO-Search18 has 10 participants searching for each of 18 target-object categories in 6202 images, making about 300,000 goal-directed fixations. When trained and evaluated on COCO-Search18, the IRL model outperformed baseline models in predicting search fixation scanpaths, both in terms of similarity to human search behavior and search efficiency. Finally, reward maps recovered by the IRL model reveal distinctive target-dependent patterns of object prioritization, which we interpret as a learned object context.
Bibliography:	ObjectType-Article-1 SourceType-Scholarly Journals-1 ObjectType-Feature-2 content type line 23
ISSN:	1063-6919 1063-6919
DOI:	10.1109/CVPR42600.2020.00027