Constructors
1Methods
9── Tokenizer overrides ────────────────────────────────────────
Apply the merge sequence to one pre-tokenized chunk. Operates on a working buffer of (id, next-pair-rank) pairs and repeatedly picks the lowest-rank pair to merge. O(N · log K) per chunk where N = chunk length and K = #active merges (typically << M because most merges' pairs never appear in the chunk).
Rebuild id_to_token_ + pair_to_merge_ from the current vocab_ + merges_str_. Called after construction and after train.