Skip to main content
Sebastian Spicker
Articles
Topics
Search
About
Dark
Light
Topic
Tokenisation
3 articles
RSS feed
all topics
2026
1 article
4 Mar
2026
The Model Has No Seahorse: Vocabulary Gaps and What They Reveal About LLMs
There is no seahorse emoji in Unicode. A language model may nevertheless claim there is one or emit a plausible substitute. That failure is real, but the output alone does not tell us whether tokenisation, training data, prompting, decoding, or verification caused it.
ai
language-models
tokenisation
2025
1 article
22 Mar
2025
The Papertrail: AI PDF Renaming and the Tokens That Make It Interesting
Everyone has a Downloads folder full of “scan0023.pdf” and “document(3)-final-FINAL.pdf”. Renaming them by content sounds trivial — read the file, understand what it is, give it a name. The implementation reveals something useful about how LLMs actually handle text: what a token is, why context windows matter in practice, why you want structured output instead of prose, and why heuristics should go first. The project, now called Folionym, is at github.com/sebastianspicker/folionym.
llm
pdf
automation
2024
1 article
7 Oct
2024
Three Rs in Strawberry: What the Viral Counting Test Actually Reveals
The viral “three Rs in strawberry” test exposes a real mismatch between token-level input and character-level operations. It does not, by itself, identify one tokenizer boundary or explain any particular model response.
llm
tokenisation
language-models