21597@AAAI

Total: 1

#1 Bridging the Gap between Expression and Scene Text for Referring Expression Comprehension (Student Abstract) [PDF] [Copy] [Kimi]

Authors: Yuqi Bu ; Jiayuan Xie ; Liuwu Li ; Qiong Liu ; Yi Cai

Referring expression comprehension aims at grounding the object in an image referred to by the expression. Scene text that serves as an identifier has a natural advantage in referring to objects. However, existing methods only consider the text in the expression, but ignore the text in the image, leading to a mismatch. In this paper, we propose a novel model that can recognize the scene text. We assign the extracted scene text to its corresponding visual region and ground the target object guided by expression. Experimental results on two benchmarks demonstrate the effectiveness of our model.