vix.ing · top · new · best · stats · spec

Learning Language from a Large (Unannotated) Corpus

2014/01/14 by Linas Vepstas, Vepstas, Linas, Ben Goertzel +1
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #I.2.0 #I.2.6 #I.2.7 #I.5.4 #Machine Learning (cs.LG) #cs.CL #cs.LG

paper · pdf · doi:10.48550/arxiv.1401.3372

29 pages, 5 figures, research proposal

arxiv created 2014/01/14 · arxiv updated 2014/01/16

Abstract

A novel approach to the fully automated, unsupervised extraction of dependency grammars and associated syntax-to-semantic-relationship mappings from large text corpora is described. The suggested approach builds on the authors' prior work with the Link Grammar, RelEx and OpenCog systems, as well as on a number of prior papers and approaches from the statistical language learning literature. If successful, this approach would enable the mining of all the information needed to power a natural language comprehension and generation system, directly from a large, unannotated corpus.

Related