Weakly Supervised PatchNets: Describing and Aggregating Local Patches for Scene Recognition

Traditional feature encoding scheme (e.g., Fisher vector) with local descriptors (e.g., SIFT) and recent convolutional neural networks (CNNs) are two classes of successful methods for image recognition. In this paper, we propose a hybrid representation, which leverages the discriminative capacity of...

Full description

Saved in:

Bibliographic Details
Published in	IEEE transactions on image processing Vol. 26; no. 4; pp. 2028 - 2041
Main Authors	Wang, Zhe, Wang, Limin, Wang, Yali, Zhang, Bowen, Qiao, Yu
Format	Journal Article
Language	English
Published	United States IEEE 01.04.2017 The Institute of Electrical and Electronics Engineers, Inc. (IEEE)
Subjects	Artificial neural networks Coding Customization Dictionaries Feature extraction Feature recognition Image coding Image recognition Image representation Neural networks Object recognition Patches (structures) PatchNet Representations scene recognition semantic codebook Semantics Training Visualization VSAD
Online Access	Get full text

Cover

Loading…

More Information
Summary:	Traditional feature encoding scheme (e.g., Fisher vector) with local descriptors (e.g., SIFT) and recent convolutional neural networks (CNNs) are two classes of successful methods for image recognition. In this paper, we propose a hybrid representation, which leverages the discriminative capacity of CNNs and the simplicity of descriptor encoding schema for image recognition, with a focus on scene recognition. To this end, we make three main contributions from the following aspects. First, we propose a patch-level and end-to-end architecture to model the appearance of local patches, called PatchNet. PatchNet is essentially a customized network trained in a weakly supervised manner, which uses the image-level supervision to guide the patch-level feature extraction. Second, we present a hybrid visual representation, called VSAD, by utilizing the robust feature representations of PatchNet to describe local patches and exploiting the semantic probabilities of PatchNet to aggregate these local patches into a global representation. Third, based on the proposed VSAD representation, we propose a new state-of-the-art scene recognition approach, which achieves an excellent performance on two standard benchmarks: MIT Indoor67 (86.2%) and SUN397 (73.0%).
Bibliography:	ObjectType-Article-1 SourceType-Scholarly Journals-1 ObjectType-Feature-2 content type line 14 content type line 23
ISSN:	1057-7149 1941-0042
DOI:	10.1109/TIP.2017.2666739