vix.ing · top · new · best · stats · spec

Semantic Alignment of AI Models: Concept Collapse, Checkpoint Dynamics, and Cross-Lingual Transfer

2026/08/03 by Tyler Ashoff, Jordan Rodu
Computer Science · #cs.CL #cs.LG

paper · pdf

Code available at github.com/tylerashoff/persiscope (PyPI: persiscope)

arxiv created 2026/08/03 · arxiv updated 2026/08/04

Abstract

Language model benchmarking is a difficult task. Outcome reasoning alone does not test the model's conceptualization of language and popular open-source benchmarks are quickly saturated or ingested as training data. It is important to test the model's output, but augmenting these tests by characterizing semantic structure gives more insight to how models relate abstract concepts. However, the high dimensional embedding spaces are not easy to interpret. This work demonstrates how topological methods can be used to rigorously compare these spaces to low dimensional and interpretable baselines like ontologies and curated knowledge graphs. These multi-modal alignment tests make it possible to track model adaptations and test phrase understanding across multiple languages.

Citations