2025/06/07 by Buğra Kılıçtaş, Faruk Alpay, Kilictas, Bugra +1 · 1 citation
Computer Science · Social Sciences · #03B70 #18M05 #68T50 #Artificial Intelligence (cs.AI) #F.4.1 #FOS: Computer and information sciences #I.2.7 #Language and cultural evolution #Logic in Computer Science (cs.LO) #Natural Language Processing Techniques #Text Readability and Simplification
paper · pdf · doi:10.48550/arxiv.2506.06870
openalex publication_date 2025/06/07 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
ISO 639:2023 unifies the ISO language-code family and introduces contextual metadata, but it lacks a machine-native mechanism for handling dialectal drift and creole mixtures. We propose a formalisation of recursive semantic anchoring, attaching to every language entity χ a family of fixed-point operators ϕn,m that model bounded semantic drift via the relation ϕn,m(χ) = χ⊕ Δ(χ), where Δ(χ) is a drift vector in a latent semantic manifold. The base anchor ϕ0,0 recovers the canonical ISO 639:2023 identity, whereas ϕ99,9 marks the maximal drift state that triggers a deterministic fallback. Using category theory, we treat the operators ϕn,m as morphisms and drift vectors as arrows in a category DriftLang. A functor Φ: DriftLang → AnchorLang maps every drifted object to its unique anchor and proves convergence. We provide an RDF/Turtle schema (BaseLanguage, DriftedLanguage, ResolvedAnchor) and worked examples -- e.g., ϕ8,4 (Standard Mandarin) versus ϕ8,7 (a colloquial variant), and ϕ1,7 for Nigerian Pidgin anchored to English. Experiments with transformer models show higher accuracy in language identification and translation on noisy or code-switched input when the ϕ-indices are used to guide fallback routing. The framework is compatible with ISO/TC 37 and provides an AI-tractable, drift-aware semantic layer for future standards.