Pimp My IDE / Garage dispatch
Back to garage
September 26, 2026 | typography / tokenizers / interface truth

Make tokens visible. Keep the meter honest.

A font can turn token boundaries into physical width. That makes model text feel expensive. It does not make the browser your billing system.

The take. Token-space type is a useful perception instrument. Pin the tokenizer, test the exact reading surface, and keep authoritative counts on a separate instrument.

A font can make token pressure visible.

The Token-space Font Compiler combines a base font with a tokenizer. It writes token logic into font shaping rules, then gives each emitted token a width of 3 em inside a supported shaping run.[1] Ordinary text becomes a row of equal-pitch units without a tokenizer script running beside it.

That changes how a reader feels a response. A short word can occupy one full unit. A visually compact string can consume several units. The usual line length stops pretending that characters and model work are the same thing.

Use the font to feel token pressure. Use the tokenizer to count tokens.

The artifact is real, large, and inspectable.

The garage downloaded the compiler's current default TrueType font, CSS, and build report. The font identified as valid TrueType data. The report names the family as DeepSeek V4.1 Flash x Inter and records 35,249 glyphs, 82,734 applied merges, and 1,691 passes. The CSS enables the required ligature feature and disables normal kerning.

This smoke check proves that the published files existed and could be inspected here. It does not prove every browser shapes every message into the same units. The project's own report and discrepancy log are much more careful than that.[1]

Text shaping and tokenization meet at a fragile seam.

HarfBuzz describes shaping as the conversion of Unicode code points into positioned glyphs. The result depends on the input string, font, script, and language.[2] A browser can also split text into separate shaping runs when formatting, links, scripts, direction, or fallback fonts change.

The compiler warns that hard breaks, whitespace handling, normalization, missing glyphs, complex scripts, and formatting boundaries can change the visual result. Its Claude-compatible options are unofficial approximations. Some experimental presets have search or byte limits. The page says plainly that the output is not an exact whole-message token counter.[1]

The model still owns the real boundary.

OpenAI's tiktoken documentation tells callers to select an encoding directly or resolve the encoding for a specific model. Its BPE explanation also notes that common subwords can become tokens.[3] Hugging Face Tokenizers treats normalization, pre-tokenization, truncation, padding, and special tokens as parts of a tokenizer pipeline.[4]

A font can approximate some of that pipeline for a reading surface. It cannot know every chat template, hidden message frame, multimodal payload, provider revision, or server-side billing rule. If money or a context limit depends on the number, call the tokenizer that matches the model and save the result with the request.

Put both instruments in the cockpit.

Use token-space type for review, teaching, prompt editing, or a deliberately strange agent transcript. Keep it scoped to the text that benefits from the effect. Then keep an exact counter beside it for limits, tests, and cost records.

The split is practical. The font answers, "Where does this text feel dense?" The tokenizer answers, "What sequence will this model receive?" The provider receipt answers, "What did this request actually cost?"

Interactive makeover / token display handoff

Token pitch gauge

Traditional purpose replaced: show one token number and hide how it was produced. Better version: stress the reading surface, close each evidence circuit, and print the gaps that still need real data.

Stress the shaping run

The cells below are a teaching proxy. They demonstrate equal pitch and run breaks. They do not tokenize your text.

Specimen condition
Evidence circuits
Teaching proxy / plain shaping run0 of 4 circuits closed
TokenizerOpen
SurfaceOpen
CasesOpen
ReceiptOpen
Font witnessEqual pitch is visible in one illustrative run.
Exact count[RUN MATCHING TOKENIZER]

The display is not the meter.

No evidence circuit is selected. Use the cells to inspect the idea, then attach real tokenizer and surface evidence.

Print the token display handoff

Replace every required field. Four selected checks mean the template structure is complete, not that a tokenizer ran.

Sources read, not vibes

Open the source log
  1. Token-space Font Compiler, read September 26, 2026. Compiler behavior, downloadable artifacts, supported presets, 3 em token width, browser-local build path, audit results, and documented shaping, normalization, coverage, framing, and counting limits.
  2. HarfBuzz manual, "What is HarfBuzz?", read September 26, 2026. Definition of text shaping and the input, font, script, and language factors that affect glyph selection and placement.
  3. OpenAI tiktoken repository and README, read September 26, 2026. Model-specific encoding selection, reversible BPE behavior, and common-subword explanation.
  4. Hugging Face Tokenizers documentation, read September 26, 2026. Tokenizer pipeline scope, including normalization, alignment tracking, truncation, padding, and special tokens.
  5. Hacker News discussion 49851883, read September 26, 2026. Discovery route and early reader reports about browser performance, Safari rendering, empathy, and mixed-script curiosity. These comments do not verify compiler correctness.

Source boundary. The compiler's page and build report support claims about its implementation and stated audits. The garage downloaded and inspected the current default TTF, CSS, and JSON report. HarfBuzz, tiktoken, and Hugging Face document neighboring mechanisms. The interactive gauge is a teaching model. It does not execute a tokenizer or inspect a real font.