Imagine you're building a parts catalog for a factory that manufactures products in 104 different countries. The standard approach — BPE and its cousins — is like sorting parts by how often they appear on the assembly line and merging the most common pairs. You get a compact inventory, but the parts for your Nigerian factory might be awkwardly chopped while the German factory gets clean, natural sub-assemblies. Compression happened, but it didn't happen fairly. The Latent Core Tokenizer (LCT) flips the workflow. Instead of letting frequency statistics drive what gets merged, it first discovers the reusable structural units of each language — morphemes, essentially — using Minimum Description Length (MDL), entropy-based boundary signals, and morphotactic constraints. Only after those structural building blocks are identified does it construct a shared 200K-token vocabulary. Structure first, frequency second. This is the core claim: separating structural discovery from vocabulary allocation produces tokens that are more linguistically meaningful, not just more compact. The results are solid if unspectacular in absolute magnitude. Across 104 languages, LCT achieves lower fertility (fewer tokens per word — meaning the tokenizer carves words into more natural pieces) and higher MorphScore (a metric measuring alignment with actual morphological boundaries) than BPE, Unigram, and parity-aware BPE. On four multilingual downstream benchmarks, LCT improves aggregate scores by 1.48, 1.83, and 2.00 points over those three baselines respectively. These are meaningful but not transformative deltas — the kind of improvement that compounds across a pipeline but won't make headlines on its own. The architecture sits squarely in the unsupervised subword tokenization family, but with a two-stage pipeline that is genuinely distinct from the merge-then-prune loop of BPE or the probabilistic pruning of Unigram. MDL provides the information-theoretic backbone for deciding what constitutes a reusable unit; entropy spikes at character boundaries help identify where morphemes begin and end; morphotactic constraints prevent merges that violate known structural patterns. The method leans on linguistic signal rather than raw co-occurrence frequency, which is the key differentiator. Integrity is reasonable but has gaps. The baselines are current and correctly named — BPE, Unigram, and the more recent parity-aware BPE (XLM-R-style) — which is the right comparison set. Four downstream benchmarks across multilingual tasks provide a real signal, not just intrinsic tokenization metrics. However, the paper is under review, no code release is mentioned, and independent replication is absent. The cross-lingual disparity metric is acknowledged honestly: LCT maintains 'comparable' disparity, meaning it doesn't solve the equity problem, it just doesn't make it worse while improving quality. The bigger implication is methodological. The NLP field has treated tokenization as a solved preprocessing step for years — BPE works, ship it. This paper argues that the way you discover structure before allocating vocabulary capacity matters for representation quality downstream. If this result holds under replication, it suggests that the entire tokenizer-as-compression-tool framing leaves performance on the table, especially for morphologically rich and low-resource languages. The next test is whether LCT's structural discovery phase scales to the vocabulary sizes and training regimes of frontier LLMs, where the tokenizer's interaction with the model's embedding layer is the real bottleneck.