Zero-Shot In-Distribution Detection in Multi-Object Settings Using Vision-Language Foundation Models

Extracting in-distribution (ID) images from noisy images scraped from the Internet is an important preprocessing for constructing datasets, which has traditionally been done manually. Automating this preprocessing with deep learning techniques presents two key challenges. First, images should be col...

Full description

Saved in:

Bibliographic Details
Main Authors	Miyai, Atsuyuki, Yu, Qing, Irie, Go, Aizawa, Kiyoharu
Format	Journal Article
Language	English
Published	10.04.2023
Subjects	Computer Science - Computer Vision and Pattern Recognition
Online Access	Get full text

Cover

Loading…

Abstract	Extracting in-distribution (ID) images from noisy images scraped from the Internet is an important preprocessing for constructing datasets, which has traditionally been done manually. Automating this preprocessing with deep learning techniques presents two key challenges. First, images should be collected using only the name of the ID class without training on the ID data. Second, as we can see why COCO was created, it is crucial to identify images containing not only ID objects but also both ID and out-of-distribution (OOD) objects as ID images to create robust recognizers. In this paper, we propose a novel problem setting called zero-shot in-distribution (ID) detection, where we identify images containing ID objects as ID images (even if they contain OOD objects), and images lacking ID objects as OOD images without any training. To solve this problem, we leverage the powerful zero-shot capability of CLIP and present a simple and effective approach, Global-Local Maximum Concept Matching (GL-MCM), based on both global and local visual-text alignments of CLIP features. Extensive experiments demonstrate that GL-MCM outperforms comparison methods on both multi-object datasets and single-object ImageNet benchmarks. The code will be available via https://github.com/AtsuMiyai/GL-MCM.
AbstractList	Extracting in-distribution (ID) images from noisy images scraped from the Internet is an important preprocessing for constructing datasets, which has traditionally been done manually. Automating this preprocessing with deep learning techniques presents two key challenges. First, images should be collected using only the name of the ID class without training on the ID data. Second, as we can see why COCO was created, it is crucial to identify images containing not only ID objects but also both ID and out-of-distribution (OOD) objects as ID images to create robust recognizers. In this paper, we propose a novel problem setting called zero-shot in-distribution (ID) detection, where we identify images containing ID objects as ID images (even if they contain OOD objects), and images lacking ID objects as OOD images without any training. To solve this problem, we leverage the powerful zero-shot capability of CLIP and present a simple and effective approach, Global-Local Maximum Concept Matching (GL-MCM), based on both global and local visual-text alignments of CLIP features. Extensive experiments demonstrate that GL-MCM outperforms comparison methods on both multi-object datasets and single-object ImageNet benchmarks. The code will be available via https://github.com/AtsuMiyai/GL-MCM.
Author	Yu, Qing Aizawa, Kiyoharu Miyai, Atsuyuki Irie, Go
Author_xml	– sequence: 1 givenname: Atsuyuki surname: Miyai fullname: Miyai, Atsuyuki – sequence: 2 givenname: Qing surname: Yu fullname: Yu, Qing – sequence: 3 givenname: Go surname: Irie fullname: Irie, Go – sequence: 4 givenname: Kiyoharu surname: Aizawa fullname: Aizawa, Kiyoharu
BackLink	https://doi.org/10.48550/arXiv.2304.04521$$DView paper in arXiv
BookMark	eNotj71OwzAAhD3AAIUHYMIv4ODfNBlRS6FSqg4tDCyRf4NRsFHsIHj7hsB0p9PdSd8lOAsxWABuCC54JQS-k8O3_yoow7zAXFByAcyrHSI6vMUMtwGtfcqDV2P2McC1zVbPzge4G_vs0V69TxE82Jx96BJ8TpPAF5-mFmpk6EbZWbiJYzByXu6isX26AudO9sle_-sCHDcPx9UTavaP29V9g2S5JIhLakVdS-4Y07pSQjLDa7vUZa2xcJow46ipuSkZxdhUiggqtFSKUO2c1WwBbv9uZ8z2c_Afcvhpf3HbGZedAFcLU4c
ContentType	Journal Article
Copyright	http://arxiv.org/licenses/nonexclusive-distrib/1.0
Copyright_xml	– notice: http://arxiv.org/licenses/nonexclusive-distrib/1.0
DBID	AKY GOX
DOI	10.48550/arxiv.2304.04521
DatabaseName	arXiv Computer Science arXiv.org
DatabaseTitleList
Database_xml	– sequence: 1 dbid: GOX name: arXiv.org url: http://arxiv.org/find sourceTypes: Open Access Repository
DeliveryMethod	fulltext_linktorsrc
ExternalDocumentID	2304_04521
GroupedDBID	AKY GOX
ID	FETCH-LOGICAL-a671-4a2e599a4f33cc8b5a3d49e7c69c05fc13df2d94d63200d8b1525cabb12cffec3
IEDL.DBID	GOX
IngestDate	Mon Jan 08 05:37:09 EST 2024
IsDoiOpenAccess	true
IsOpenAccess	true
IsPeerReviewed	false
IsScholarly	false
Language	English
LinkModel	DirectLink
MergedId	FETCHMERGED-LOGICAL-a671-4a2e599a4f33cc8b5a3d49e7c69c05fc13df2d94d63200d8b1525cabb12cffec3
OpenAccessLink	https://arxiv.org/abs/2304.04521
ParticipantIDs	arxiv_primary_2304_04521
PublicationCentury	2000
PublicationDate	2023-04-10
PublicationDateYYYYMMDD	2023-04-10
PublicationDate_xml	– month: 04 year: 2023 text: 2023-04-10 day: 10
PublicationDecade	2020
PublicationYear	2023
Score	1.8776387
SecondaryResourceType	preprint
Snippet	Extracting in-distribution (ID) images from noisy images scraped from the Internet is an important preprocessing for constructing datasets, which has...
SourceID	arxiv
SourceType	Open Access Repository
SubjectTerms	Computer Science - Computer Vision and Pattern Recognition
Title	Zero-Shot In-Distribution Detection in Multi-Object Settings Using Vision-Language Foundation Models
URI	https://arxiv.org/abs/2304.04521
hasFullText	1
inHoldings	1
isFullTextHit
isPrint
link	http://utb.summon.serialssolutions.com/2.0.0/link/0/eLvHCXMwdV09T8MwELVKJxYEAlQ-5YHVEMeOE4-IUgoCOrSgiKXy2Y6ohJIqCYifj-0ElYXR9k3Plu58H-8hdAFci7Sgimjnawk3RUxAWUGYBBW7AFnKMOH99CymL_whT_IBwr-zMKr-Xn11_MDQXPmM5aUn_Xb_m6049i1bd7O8K04GKq7efmPnYsyw9cdJTHbRTh_d4evuOvbQwJb7yLzZuiLz96rF9yUZe6baXmQKj20bWqFKvCpxmIUlM_CZETy3oSG5waGmj1_DCDh57LOLeKOGhL2a2UdzgBaT28XNlPTiBkSJ1KGiYptIqXjBmNYZJIoZLm2qhdRRUmjKHGxGciOYe8cmA69TpBUAjbVv9GCHaFhWpR0hDMpBnKWFpYZziARQKrR2K8tdsMLUERoFSJbrjr9i6dFaBrSO_z86QdteWd0XTmh0ioZt_WnPnP9t4Txcwg9aDYfP
link.rule.ids	228,230,783,888
linkProvider	Cornell University
openUrl	ctx_ver=Z39.88-2004&ctx_enc=info%3Aofi%2Fenc%3AUTF-8&rfr_id=info%3Asid%2Fsummon.serialssolutions.com&rft_val_fmt=info%3Aofi%2Ffmt%3Akev%3Amtx%3Ajournal&rft.genre=article&rft.atitle=Zero-Shot+In-Distribution+Detection+in+Multi-Object+Settings+Using+Vision-Language+Foundation+Models&rft.au=Miyai%2C+Atsuyuki&rft.au=Yu%2C+Qing&rft.au=Irie%2C+Go&rft.au=Aizawa%2C+Kiyoharu&rft.date=2023-04-10&rft_id=info:doi/10.48550%2Farxiv.2304.04521&rft.externalDocID=2304_04521