vix.ing · top · new · best · stats · spec

Zeng, Gangyan

  1. Track the Answer: Extending TextVQA from Image to Video with Spatio-Temporal Clues
    2024/12/17 by Yan Zhang, Gangyan Zeng, Zhang, Yan +9 · 4 citations
    Computer Science · #Natural Language Processing Techniques #Topic Modeling #Multimodal Machine Learning Applications
  2. TextBlockV2: Towards Precise-Detection-Free Scene Text Spotting with Pre-trained Language Model
    2024/03/15 by Jiahao Lyu, Jin Wei, Lyu, Jiahao +11 · 2 citations
    Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Handwritten Text Recognition Techniques #Multimodal Machine Learning Applications #Video Analysis and Summarization
  3. Focus, Distinguish, and Prompt: Unleashing CLIP for Efficient and Flexible Scene Text Retrieval
    2024/08/01 by Gangyan Zeng, Zeng, Gangyan, Yuan Zhang +13 · 2 citations
    Computer Science · #Image Retrieval and Classification Techniques #Topic Modeling #Text and Document Classification Technologies
  4. When Semantics Mislead Vision: Mitigating Large Multimodal Models Hallucinations in Scene Text Spotting and Understanding
    2025/06/05 by Shu, Yan, Lin, Hangui, Liu, Yexin +7 · 2 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  5. VidText: Towards Comprehensive Evaluation for Video Text Understanding
    2025/05/28 by Yang, Zhoufaran, Shu, Yan, Wang, Jing +8 · 3 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  6. Gather and Trace: Rethinking Video TextVQA from an Instance-oriented Perspective
    2025/08/06 by Zhang, Yan, Zeng, Gangyan, Wu, Daiqing +5 · 2 citations
    #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences