Skip to content
ChristopherOx Alpha
Index
  1. Briefing
  2. Stage Test · 3/3

Test

Interactive labs

Knife, +75, MoE, 2ⁿ, video toll, mixer, thermostat and the whole loop.

  1. You will understand
  2. Predict, move, observe.
  3. One lab at a time.
  4. The tables remain if the graphic does not load.

Lab: the whole loop

Six stations. Each is a click. At the end you have three seals: constant +75, same cashier, bill = active params.

Text in

You type “Hello world”. Still letters. The model does not see them.

Lab: cut the text

This is not the real GLM vocab. It is a simulator: two knives, two counts. In the real test, Ox and GLM-5.3 match; Kimi/MiMo/Qwen do not.

Before you touch

If the knife were the same and there was a letterhead, what does Δ do?

Pick one. Then move the control and compare.

Why

Fixed. A wrapper adds a constant extra. If Δ jitters, the knife is not the same.
Two-knife simulator on the same text. The real test is the +75 constant, not this vocabulary.. Two-knife simulator on the same text. The real test is the +75 constant, not this vocabulary.

Compact style (CJK/GLM family, illustrative) · 11 pieces

Hola,¿cómoestáshoy?

Western, more split (illustrative) · 14 pieces

Hola,¿cómoestáshoy?

If Ox Alpha uses the GLM knife, its count is 86 = 11 pieces + 75 letterhead. The Western knife gives 14.

Rebuilt table · +75 test

Published pattern (dax / r/singularity / MaxForAI / Davis): every text is exactly 75 tokens more on Ox than on GLM-5.3.

PromptGLMOxΔMiMoKimi
EN
Hello, how are you today?
88375109
DE
Guten Morgen, wie geht es di
1186751314
CN
你好,今天过得怎么样?
98475118
code
def fib(n): return n if n<2
2499752628
emoji
🚀🔥💡✨🎯
1085751415
EN
The quick brown fox jumps ov
1287751413

Illustrative numbers of the public pattern, not our own run. The constant is the finding.

Lab: subtract the letterhead

Type the token counts two APIs returned for the SAME text. If Δ = 75 every time, the knife is the same.

Before you touch

If GLM returns 12 and Ox returns 87, Δ is...

Pick one. Then move the control and compare.

Why

75. 87 − 12 = 75. That is the letterhead.
Δ = tok_ox − tok_glm
Δ = 87 − 12 = 75

H0  same tokenizer + wrapper   →  Δ = 75  for every P
H1  different tokenizer        →  Var(Δ) ≫ 0

Fits the hypothesis: 75 letterhead tokens. If this repeats in English, Chinese, code and emoji, the dictionary is GLM.

Lab: wake the specialists

The building always has 744 “musicians”. You only pay the ones who play. A dense 2T model bills everyone, all the time.

Before you touch

If you wake fewer experts, the bill...

Pick one. Then move the control and compare.

Why

Falls. cost ≈ k · P_active · tokens. The whole building is not billed on every word.
744 musicians in the building. You only pay the ones who play. Real GLM-5.3 uses ~40B of 744B.. 744 musicians in the building. You only pay the ones who play. Real GLM-5.3 uses ~40B of 744B.
Building (total)
744B
Awake (the bill)
40B
Cheaper than dense 2T
×50.0

Only 5.4% of the building works on each word. Real GLM-5.3 uses ~40B of 744B.

Lab: 2ⁿ is not marketing

GPUs love powers of two. 1,048,576 is not a round “million”. It is exactly 2²⁰.

Before you touch

If n goes up by 1, the cube...

Pick one. Then move the control and compare.

Why

It doubles. That is why 2¹⁹ to 2²⁰ looks like a desk, not a marketing round number.
Each step multiplies by two. 2²⁰ = 1,048,576, Ox Alpha’s desk and GLM-5.2/5.3’s.. Each step multiplies by two. 2²⁰ = 1,048,576, Ox Alpha’s desk and GLM-5.2/5.3’s.

220 = 1,048,576

≈ 10.5 books of 100k tokens. This is Ox Alpha’s desk, and GLM-5.2/5.3’s.

  • 2^10 · 1,024 · a long SMS
  • 2^14 · 16,384 · an essay
  • 2^17 · 131,072 · Ox max output
  • 2^20 · 1,048,576 · Ox / GLM-5.2 context

Video toll booth

Four published clips (Davis / Railway / BohuTANG). Ox and GLM-5V-Turbo charge the same. MiMo and Qwen do not. 5 fps = 30 fps is the trick: the booth bills time, not frames.

Before you touch

If 5 fps bills the same as 30 fps, the cashier is counting...

Pick one. Then move the control and compare.

Why

Time. The identical Ox = 5V-Turbo toll is the fingerprint; fps invariance is common.
Published numbers, not our own run. Ox and GLM-5V-Turbo charge the same on four clips.. Published numbers, not our own run. Ox and GLM-5V-Turbo charge the same on four clips.
Ox Alpha
296
GLM-5V-Turbo
296
MiMo v2.5
910
Qwen
408

Stamp: SAME CASHIER. 296 = 296 tokens.

Published numbers, not our own run. GLM-4.6V bills differently: “GLM-ish” is not enough.

Lab: the thermostat

Temperature 0 = always the most likely token (disciplined parrot). 1 = jazz. Clone fingerprinting uses 0. Pedagogical simulator, not the real API.

Before you touch

To compare clones character by character, which T do you use?

Pick one. Then move the control and compare.

Why

Near 0. Close to argmax. At high T, pizza competes and the passport dissolves.

Chat/code zone. Useful for writing, useless as a passport.

Confidence mixer

A rubric, not Bayes. Drag how much you trust each test. The ceiling is locked: no live run and no lab statement means no passport. Manifold is the market, not this bar.

Before you touch

If you raise only the stealth playbook, who should rise more?

Pick one. Then move the control and compare.

Why

Xiaomi rises more on this rubric. The playbook pushes; the tokenizer does not. Lower the playbook and GLM stays near the ceiling.
Confidence mixer. Four rubric bars. The 90% ceiling is locked: no live run and no lab statement.

Z.ai / GLM-5.x

82%

Your rubric · 90% ceiling

Xiaomi MiMo-V3

24%

Playbook helps; the knife does not

Fine-tune / Composer

22%

Tokenizer yes, z.ai kitchen less

Inkling · Mira

6%

Audio ≠ video. Ruled out.