CLIP-Llama: A New Approach for Scene Text Recognition with a Pre-Trained Vision-Language Model and a Pre-Trained Language Model
This study focuses on Scene Text Recognition (STR), which plays a crucial role in various applications of artificial intelligence such as image retrieval, office automation, and intelligent transportation systems. Currently, pre-trained vision-language models have become the foundation for various d...
Saved in:
Published in | Sensors (Basel, Switzerland) Vol. 24; no. 22; p. 7371 |
---|---|
Main Authors | , , , |
Format | Journal Article |
Language | English |
Published |
Switzerland
MDPI AG
19.11.2024
MDPI |
Subjects | |
Online Access | Get full text |
Cover
Loading…
Be the first to leave a comment!