vix.ing · top · new · best · stats · spec

Geometric Risk Control for Vision-Language Model OCR

2026/03/31 by Weile Gong, Zijian Lu, Mingcai Chen +3
Computer Science · #cs.CV

paper · pdf

9 pages, 5 figures, Code: https://github.com/phare111/GRC

arxiv created 2026/07/30 · arxiv updated 2026/07/31

Abstract

Vision-language models (VLMs) enable flexible generative optical character recognition (OCR), while their open-ended decoders can expose wrong but fluent text with weak visual support. In audit-sensitive records, such an output can be more costly than abstention. Frozen or externally served VLMs therefore require an external decision layer that can determine whether a transcription has sufficient visual evidence for release. We introduce the Geometric Risk Controller (GRC), a model-agnostic controller that treats controlled geometric transformations as repeatable black-box probes, screens structurally implausible continuations, and releases the unique candidate supported by coherent cross-view evidence. The protocol provides empirical selective exposure control with explicit coverage and query cost under a reproducible fixed decision rule. Experiments across frozen VLMs and standard scene-text benchmarks consistently reduce mean, upper-tail, and catastrophic error among released outputs while retaining high coverage.

Citations

Related