ttok switches default tokenizer from GPT-4 to GPT-5/GPT-6

On 9 October 2026 the author of the ttok tool described a small slip in a post on their blog. They had released ttok 0.4, run uv tool upgrade ttok, and piped a file into the new version. Only then did they realise it was defaulting to the GPT-4 tokenizer when it should, in their words, "very clearly" default to GPT-5/GPT-6 instead.

The author figured that switching the default was a reasonable excuse to finally ship a 1.0, so the change is tied to the move to version 1.0.

There is a catch. OpenAI has not confirmed that GPT-6 uses the same tokenizer as the GPT-5 family, and the author notes there is an angry issue about it. To fill the gap, the author points to a commit by William Liu, which reports on an experiment he ran and which indicates the tokenizers are likely the same.

The commit text, quoted in the post, says all seven GPT models tested (5.5, 5.6 Sol/Terra/Luna, and 6 Astra/Sol/Luna) report 44,794 tokens and match each other on every one of the 31 fixtures. It adds that GPT-6 introduces no input-count change on this corpus.

So the default now rests on an inference: identical counts across seven models on one test corpus, not a statement from OpenAI.

Key facts

  • ttok 0.4 defaulted to the GPT-4 tokenizer, which the author noticed right after upgrading with uv tool upgrade ttok.
  • The author used the default switch to GPT-5/GPT-6 as a reason to ship ttok 1.0.
  • A commit by William Liu reports that seven GPT models (5.5, 5.6 Sol/Terra/Luna, 6 Astra/Sol/Luna) all report 44,794 tokens and match on every one of 31 fixtures.
  • OpenAI has not confirmed that GPT-6 uses the same tokenizer as the GPT-5 family; the experiment only suggests the tokenizers are likely the same.

Why it matters

A tool that reads piped text and works with a tokenizer is only as useful as its default. Here the default had slipped back to the GPT-4 tokenizer, which the author considers clearly wrong for current models. The experiment also points to something broader: seven models across the GPT-5.5, 5.6 and 6 variants gave identical counts of 44,794 tokens on the test corpus, which suggests a single tokenizer may serve all of them.

Who it affects

Anyone using ttok, especially people on version 0.4 who rely on the default tokenizer, since the author says that default was GPT-4 rather than GPT-5/GPT-6. It is also relevant to anyone checking whether GPT-6 counts tokens the same way as the GPT-5 family.

How to use it

The author's own workflow was to upgrade with uv tool upgrade ttok and pipe a file into the tool, which is how the wrong default surfaced. The post does not give further usage instructions. The practical takeaway is that the GPT-5/GPT-6 tokenizer is meant to be the default from the 1.0 line the author plans to ship.

How solid is it

The account is a first-person blog post by the tool's author, so the bug and the decision to switch the default are credible as the author's own report. The claim about GPT-6 is weaker: it rests on William Liu's commit, a third-party experiment with seven models and 31 fixtures. The author calls the tokenizers "likely" the same, and OpenAI has not confirmed it.

Risks and caveats

The result is scoped to one test corpus, and the commit itself says GPT-6 introduces no input-count change "on this corpus". If OpenAI's GPT-6 tokenizer later turns out to differ, counts from the new default could drift from the real ones. The post also mentions an angry issue about the lack of confirmation, but does not identify it.

“GPT-6 introduces no input-count change on this corpus.”

— William Liu's commit, as quoted in the post