CS336 TikTokenizer

See exactly how your custom-trained BPE Tokenizer breaks down text into subwords. Built from scratch for Stanford CS336, trained on the TinyStories dataset with a vocabulary size of 10,000.

Input Text

Tokens

0 tokens
Tokens will appear here in real-time...