A new arXiv paper introduces Toten, a knowledge-based ontological tokenization approach for physical quantities and technical notation in Brazilian Portuguese. The work focuses on a narrow but important NLP problem: preserving the meaning of specialized notation during language processing.

Technical text often breaks standard tokenization assumptions, especially when units, quantities, symbols, and multilingual context appear together. Better tokenization can improve downstream extraction, search, and domain-specific language understanding.

The paper is another reminder that practical AI systems often depend on small preprocessing decisions that determine whether specialized knowledge survives the pipeline.