2025/02/14 by Baichuan Mo, Belinda Mo, Mo, Belinda +18 · 1 voice · 21 citations
Computer Science · #Artificial intelligence #Computer science #Linguistics #Natural Language Processing Techniques #Natural language processing #Philosophy #Plain language #Plain text #Semantic Web and Ontologies #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2502.09956
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2025/02/14 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
Recent interest in building foundation models for KGs has highlighted a fundamental challenge: knowledge-graph data is relatively scarce. The best-known KGs are primarily human-labeled, created by pattern-matching, or extracted using early NLP techniques. While human-generated KGs are in short supply, automatically extracted KGs are of questionable quality. We present a solution to this data scarcity problem in the form of a text-to-KG generator (KGGen), a package that uses language models to create high-quality graphs from plaintext. Unlike other KG extractors, KGGen clusters related entities to reduce sparsity in extracted KGs. KGGen is available as a Python library (pip install kg-gen), making it accessible to everyone. Along with KGGen, we release the first benchmark, Measure of of Information in Nodes and Edges (MINE), that tests an extractor's ability to produce a useful KG from plain text. We benchmark our new tool against existing extractors and demonstrate far superior performance.