── CharTokenizer ───────────────────────────────────────────────────
One token per Unicode codepoint. Vocab is the set of distinct codepoints encountered during training (or supplied at construction time).
── CharTokenizer ───────────────────────────────────────────────────
One token per Unicode codepoint. Vocab is the set of distinct codepoints encountered during training (or supplied at construction time).