Recent advancements in byte-level large language models (LLMs) have made tokenizer-free modeling increasingly competitive by grouping bytes into dynamically sized patches. However, existing byte-patch architectures traditionally apply uniform dense feed-forward computation to every patch, failing to adapt model capacity to variations in patch semantics and granularity.
The Shift to EntropyMoE
Researchers have introduced EntropyMoE, a Mixture-of-Experts (MoE) architecture designed to address these limitations. By replacing dense feed-forward modules in the global patch Transformer with Top-K expert layers, the system allows for more specialized and efficient computation tailored to dynamic byte patches.
Routing via Patch Entropy
The core innovation of EntropyMoE lies in its routing mechanism. Instead of traditional token-based signals, the router selects experts directly from patch entropy. This leverages the same granularity signal that underlies the dynamic patch construction process. Key features of this architecture include:
- Dynamic patches serve as the fundamental unit for expert routing.
- The byte coverage of each patch determines its contribution to workload accounting.
- Patch entropy and length jointly define the feature space for regulating expert specialization.
Experiments demonstrate that EntropyMoE achieves the lowest held-out bits-per-byte compared to matched dense and sparse baselines. Crucially, the model maintains comparable downstream accuracy, establishing patch entropy as an effective routing coordinate for sparse conditional computation in models beyond traditional tokenizer-based representations.