2020/03/04 by Fabian Wolf, Wolf, Fabian, Gernot A. Fink +1
Computer Science · #Annotation #Artificial intelligence #Computer Vision and Pattern Recognition (cs.CV) #Computer science #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Handwritten Text Recognition Techniques #Information retrieval #Keyword spotting #Lexicon #Machine learning #Multimodal Machine Learning Applications #Natural language processing #Scheme (mathematics) #Selection (genetic algorithm) #Spotting #String (physics) #Task (project management) #Word (group theory) #cs.CV
paper · pdf · doi:10.48550/arxiv.2003.01989
Accepted to Workshop on Document Analysis Systems (DAS) 2020
openalex publication_date 2020/03/04 · arxiv created 2020/05/25 · arxiv updated 2020/05/26 · openalex created_date 2022/07/18 · openalex updated_date 2026/08/05
Word spotting is a popular tool for supporting the first exploration of\nhistoric, handwritten document collections. Today, the best performing methods\nrely on machine learning techniques, which require a high amount of annotated\ntraining material. As training data is usually not available in the application\nscenario, annotation-free methods aim at solving the retrieval task without\nrepresentative training samples. In this work, we present an annotation-free\nmethod that still employs machine learning techniques and therefore outperforms\nother learning-free approaches. The weakly supervised training scheme relies on\na lexicon, that does not need to precisely fit the dataset. In combination with\na confidence based selection of pseudo-labeled training samples, we achieve\nstate-of-the-art query-by-example performances. Furthermore, our method allows\nto perform query-by-string, which is usually not the case for other\nannotation-free methods.\n