2021/12/24 by Yuyu Luo, Jiawei Tang, Luo, Yuyu +3 · 11 citations
Computer Science · Engineering · #Artificial Intelligence (cs.AI) #Artificial intelligence #Benchmark (surveying) #Computer science #Data Visualization and Analytics #Database #Deep learning #Domain (mathematical analysis) #Engineering #FOS: Computer and information sciences #Field (mathematics) #Human-Computer Interaction (cs.HC) #Natural language #Natural language processing #Natural language understanding #Online Learning and Analytics #SQL #Scale (ratio) #Task (project management) #Video Analysis and Summarization #Visualization #cs.AI #cs.HC
paper · pdf · doi:10.48550/arxiv.2112.12926
published in arXiv (Cornell University) (Cornell University)
arxiv created 2021/12/24 · openalex publication_date 2021/12/24 · arxiv updated 2021/12/28 · openalex created_date 2022/05/05 · openalex updated_date 2026/08/05
NL2VIS - which translates natural language (NL) queries to corresponding visualizations (VIS) - has attracted more and more attention both in commercial visualization vendors and academic researchers. In the last few years, the advanced deep learning-based models have achieved human-like abilities in many natural language processing (NLP) tasks, which clearly tells us that the deep learning-based technique is a good choice to push the field of NL2VIS. However, a big balk is the lack of benchmarks with lots of (NL, VIS) pairs. We present nvBench, the first large-scale NL2VIS benchmark, containing 25,750 (NL, VIS) pairs from 750 tables over 105 domains, synthesized from (NL, SQL) benchmarks to support cross-domain NL2VIS task. The quality of nvBench has been extensively validated by 23 experts and 300+ crowd workers. Deep learning-based models training using nvBench demonstrate that nvBench can push the field of NL2VIS.