CS336 TikTokenizer
See exactly how your custom-trained BPE Tokenizer breaks down text into subwords. Built from scratch for Stanford CS336, trained on the TinyStories dataset with a vocabulary size of 10,000.
Input Text
Tokens
0 tokensTokens will appear here in real-time...