Byte Language Models: Scaling, Emergent Abstractions, and Information Allocation
This paper challenges the conventional wisdom that language models require subword tokenizers by demonstrating that standard Transformers can effectively process raw byte sequences and outperform trad...