2026/01/28 by Murad Farzulla · 1 voice
Economics, Econometrics and Finance · Computer Science · #q-fin.CP #cs.LG
arxiv published 2026/01/28 · arxiv updated 2026/07/11
Do the functional narratives in cryptocurrency whitepapers correspond to how their tokens behave in markets? We develop a content-verified, contamination-aware pipeline for measuring structural correspondence between project narratives and market structure, and report two results. The first is a cautionary one. An apparent entity-level signal in an earlier version of our corpus -- specialised tokens appearing to align more strongly than broad infrastructure tokens -- was entirely an artifact of corpus contamination: roughly a quarter of the documents were failed-download stubs or wrong-document whitepapers (for example, a "Cosmos" entry that was in fact Binance Smart Chain text), and the apparent ordering does not survive content verification: on the clean corpus no token registers as helping alignment. We therefore report it as a contamination diagnosis, not a finding. The second is an honest null. Combining zero-shot NLP classification of 43 content-verified whitepapers across 10 semantic categories with seven cross-sectional market-structure statistics computed from hourly data (17,543 timestamps, 2023-2024), and aligning the two spaces with Procrustes rotation and Tucker's congruence coefficient (φ), we do not detect a significant claims-market alignment in this n = 43 sample (dimension-matched φ= 0.303, zero-padded φ= 0.223; both non-significant). A positive-control and power analysis shows the binding constraint is the low reliability of the text instrument: the minimum detectable effect is φ≈ 0.66, well above the observed ≈ 0.22. This is absence of evidence for alignment, not evidence of its absence -- we can reject strong alignment (φ≥ 0.70) but cannot distinguish weak alignment (φ≈ 0.3) from none.