vix.ing · top · new · best · stats · spec

A Support Tool for Tagset Mapping

1995/06/08 by Simone Teufel, Teufel, Simone
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Mathematics, Computing, and Information Processing #Natural Language Processing Techniques #Topic Modeling #cmp-lg #cs.CL

paper · pdf · doi:10.48550/arxiv.cmp-lg/9506005

EACL-Sigdat 95, contains 4 ps figures (minor graphic changes)

openalex publication_date 1995/06/08 · arxiv created 1995/09/07 · arxiv updated 2009/11/30 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Many different tagsets are used in existing corpora; these tagsets vary according to the objectives of specific projects (which may be as far apart as robust parsing vs. spelling correction). In many situations, however, one would like to have uniform access to the linguistic information encoded in corpus annotations without having to know the classification schemes in detail. This paper describes a tool which maps unstructured morphosyntactic tags to a constraint-based, typed, configurable specification language, a ``standard tagset''. The mapping relies on a manually written set of mapping rules, which is automatically checked for consistency. In certain cases, unsharp mappings are unavoidable, and noise, i.e. groups of word forms \sl not conforming to the specification, will appear in the output of the mapping. The system automatically detects such noise and informs the user about it. The tool has been tested with rules for the UPenn tagset \citeup and the SUSANNE tagset \citegarside, in the framework of the EAGLES\footnoteLRE project EAGLES, cf. \citeeagles. validation phase for standardised tagsets for European languages.

Related