- Briefing
- Stage Test · 3/3
Test
Interactive labs
Knife, +75, MoE, 2ⁿ, video toll, mixer, thermostat and the whole loop.
- You will understand
- Predict, move, observe.
- One lab at a time.
- The tables remain if the graphic does not load.
Lab: the whole loop
Six stations. Each is a click. At the end you have three seals: constant +75, same cashier, bill = active params.
Text in
You type “Hello world”. Still letters. The model does not see them.
Lab: cut the text
This is not the real GLM vocab. It is a simulator: two knives, two counts. In the real test, Ox and GLM-5.3 match; Kimi/MiMo/Qwen do not.
Why
Compact style (CJK/GLM family, illustrative) · 11 pieces
Western, more split (illustrative) · 14 pieces
If Ox Alpha uses the GLM knife, its count is 86 = 11 pieces + 75 letterhead. The Western knife gives 14.
Rebuilt table · +75 test
Published pattern (dax / r/singularity / MaxForAI / Davis): every text is exactly 75 tokens more on Ox than on GLM-5.3.
| Prompt | GLM | Ox | Δ | MiMo | Kimi |
|---|---|---|---|---|---|
| EN Hello, how are you today? | 8 | 83 | 75 | 10 | 9 |
| DE Guten Morgen, wie geht es di | 11 | 86 | 75 | 13 | 14 |
| CN 你好,今天过得怎么样? | 9 | 84 | 75 | 11 | 8 |
| code def fib(n): return n if n<2 | 24 | 99 | 75 | 26 | 28 |
| emoji 🚀🔥💡✨🎯 | 10 | 85 | 75 | 14 | 15 |
| EN The quick brown fox jumps ov | 12 | 87 | 75 | 14 | 13 |
Illustrative numbers of the public pattern, not our own run. The constant is the finding.
Lab: subtract the letterhead
Type the token counts two APIs returned for the SAME text. If Δ = 75 every time, the knife is the same.
Why
Δ = tok_ox − tok_glm Δ = 87 − 12 = 75 H0 same tokenizer + wrapper → Δ = 75 for every P H1 different tokenizer → Var(Δ) ≫ 0
Fits the hypothesis: 75 letterhead tokens. If this repeats in English, Chinese, code and emoji, the dictionary is GLM.
Lab: wake the specialists
The building always has 744 “musicians”. You only pay the ones who play. A dense 2T model bills everyone, all the time.
Why
Only 5.4% of the building works on each word. Real GLM-5.3 uses ~40B of 744B.
Lab: 2ⁿ is not marketing
GPUs love powers of two. 1,048,576 is not a round “million”. It is exactly 2²⁰.
Why
220 = 1,048,576
≈ 10.5 books of 100k tokens. This is Ox Alpha’s desk, and GLM-5.2/5.3’s.
- 2^10 · 1,024 · a long SMS
- 2^14 · 16,384 · an essay
- 2^17 · 131,072 · Ox max output
- 2^20 · 1,048,576 · Ox / GLM-5.2 context
Video toll booth
Four published clips (Davis / Railway / BohuTANG). Ox and GLM-5V-Turbo charge the same. MiMo and Qwen do not. 5 fps = 30 fps is the trick: the booth bills time, not frames.
Why
Stamp: SAME CASHIER. 296 = 296 tokens.
Published numbers, not our own run. GLM-4.6V bills differently: “GLM-ish” is not enough.
Lab: the thermostat
Temperature 0 = always the most likely token (disciplined parrot). 1 = jazz. Clone fingerprinting uses 0. Pedagogical simulator, not the real API.
Why
Chat/code zone. Useful for writing, useless as a passport.
Confidence mixer
A rubric, not Bayes. Drag how much you trust each test. The ceiling is locked: no live run and no lab statement means no passport. Manifold is the market, not this bar.
Why
Z.ai / GLM-5.x
82%
Your rubric · 90% ceiling
Xiaomi MiMo-V3
24%
Playbook helps; the knife does not
Fine-tune / Composer
22%
Tokenizer yes, z.ai kitchen less
Inkling · Mira
6%
Audio ≠ video. Ruled out.