google/sentencepiece ? reverse-engineered prompt

Reverse engineered prompt

Build me a fast text tokenizer and detokenizer that can learn directly from raw text, then turn sentences into subword pieces or token IDs and back again.

I want it to work well for language model style text, so it should handle whitespace in a reversible way, and support both a unigram based approach and BPE. It should be able to train a model from a plain text file, save it to a single model file, then load that same model later for encoding and decoding in a consistent way.

Please include a simple Python interface too, since that’s the easiest way for me to try it out, and make sure the core is efficient enough for large text files. If there are existing docs or current best practices online that help, check those while you build it.

Are you gonna build this?

make sure you review the code using coderabbit

Try freeSponsored — opens CodeRabbit in a new tab