vix.ing · top · new · best · stats · spec

3MVRD: Multimodal Multi-task Multi-teacher Visually-Rich Form Document Understanding

2024/02/28 by Yihao Ding, Lorenzo Vaiani, Ding, Yihao +10 · 1 citation
Computer Science · #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Hand Gesture Recognition Systems #Handwritten Text Recognition Techniques #Natural Language Processing Techniques

paper · pdf · doi:10.48550/arxiv.2402.17983

openalex publication_date 2024/02/28 · openalex created_date 2024/05/22 · openalex updated_date 2026/07/28

Abstract

This paper presents a groundbreaking multimodal, multi-task, multi-teacher joint-grained knowledge distillation model for visually-rich form document understanding. The model is designed to leverage insights from both fine-grained and coarse-grained levels by facilitating a nuanced correlation between token and entity representations, addressing the complexities inherent in form documents. Additionally, we introduce new inter-grained and cross-grained loss functions to further refine diverse multi-teacher knowledge distillation transfer process, presenting distribution gaps and a harmonised understanding of form documents. Through a comprehensive evaluation across publicly available form document understanding datasets, our proposed model consistently outperforms existing baselines, showcasing its efficacy in handling the intricate structures and content of visually complex form documents.

Cited by

Related