2013/08/22 by Robert A. Bridges, Bridges, Robert A., Corinne L. Jones +7 · 3 citations
Computer Science · Decision Sciences · #Computation and Language (cs.CL) #Data Quality and Management #FOS: Computer and information sciences #Information Retrieval (cs.IR) #Natural Language Processing Techniques #Topic Modeling
paper · pdf · doi:10.48550/arxiv.1308.4941
openalex publication_date 2013/08/22 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Timely analysis of cyber-security information necessitates automated information extraction from unstructured text. While state-of-the-art extraction methods produce extremely accurate results, they require ample training data, which is generally unavailable for specialized applications, such as detecting security related entities; moreover, manual annotation of corpora is very costly and often not a viable solution. In response, we develop a very precise method to automatically label text from several data sources by leveraging related, domain-specific, structured data and provide public access to a corpus annotated with cyber-security entities. Next, we implement a Maximum Entropy Model trained with the average perceptron on a portion of our corpus (∼750,000 words) and achieve near perfect precision, recall, and accuracy, with training times under 17 seconds.