Skip to content
ChristopherOx Alpha
Index
  1. Briefing
  2. Stage Case · 1/3

Case

Every suspect, one by one

GLM, Xiaomi, ByteDance, Cursor, Mira, the West. This file’s rubric, not a market.

  1. You will understand
  2. One suspect at a time, for and against.
  3. GLM leads as a hypothesis, not a verdict.
  4. A “kill” cell is an instrument.

How to read the score. 85 is not “it is GLM”. It is “if Z.ai claims it tomorrow, nobody is shocked; if Murati claims it, we rewrite the file”.

85% · leads · Z.ai / Zhipu AI (China)

GLM-5.x multimodal

Ox Alpha is, with high probability, an unpublished variant of the GLM-5 family (same tokenizer, same API errors, same video encoder, same style). Public GLM-5.3 is text-only; Ox Alpha adds image and video.

For

  • Tokenizer: constant +75 tokens vs GLM-5.3 on every text (EN, DE, CN, code, emoji).
  • Error codes 1210 / 1214 identical to chat.z.ai.
  • Video encoder with the same token budgets as GLM-5V-Turbo.
  • Prefix cache: one pi.dev tester saw ~97%; OpenRouter catalog lists ~87% Ox vs ~92% GLM-5.3. Useful, not a passport.
  • Jailbreaks that leak “I'm GLM, a large language model developed by Z.ai”.
  • Prior playbook: Pony Alpha (Feb 2026) → revealed as GLM-5 (744B MoE).
  • Z.ai just stood up a 1 GW data center and is stress-testing Huawei chips with huge limits.
  • Kingbench 87.5% (2nd) vs GLM-5.3 91.25% (1st): siblings, not twins.
  • Manifold: ~85% Z.ai as of 21 Aug 2026. Pliny: “Ox-alpha is from Zai, GLM-5.X family”.
  • 21 Aug: @JoshRadDev got Chinese error logs + ModelScope GLM strings on invalid params.
  • Tony Zhou: tokenizer vector [76, 384, 114, 64, 49, 313] L1=0 vs GLM-5.3; MiMo 176; MiniMax 172.
  • Independent +75 replications: aitrackerbot N=25, RocketmanSh 88/106/97/108 vs 13/31/22/33, r/singularity N=6, BohuTANG N=4, KatilaMarriete N=7, Sevi N=6. unclecode 4/4, Sushaanth 11/11, Pritish 17.
  • Video 4 clips: Ox = GLM-5V-Turbo token-for-token (296/296/884/1064). MiMo 910/910/1364/1134. Qwen 408/408/1288/1156.

Against

  • Public GLM-5.3 is text-only. If it is GLM, it has to be an unreleased V / Flash / Air.
  • Sharing a tokenizer does not prove identical weights: someone could fine-tune GLM.
  • Dylan Fetch: on OpenTTD-Bench 5.2 and 5.3 look so alike he doubts a “GLM-5.4” and points at MiMo-V3.
  • 17% timeouts in one tester; slower than GLM-5.3 on coding (18.3s vs 6.7s).
  • Nobody at Z.ai has confirmed it. Flash/~300B is a rumor, not a datasheet.
  • The SKU (5.3V vs Flash vs Air) is still open. Manifold “who” ~80% Z.ai; “name” GLM-5.3V ~46%.
Analogy, example, trap

Analogy. A fingerprint, the same shoe, the same voice and the same zip code. Not a passport photo — but the detective already has a clear suspect.

8% · weak · Xiaomi (China)

MiMo-V3

Xiaomi already used the stealth trick: Hunter Alpha and Healer Alpha unmasked as MiMo-V2-Pro and MiMo-V2-Omni. Ox Alpha matches their sheet: 1M context, multimodal, agents and coding.

For

  • Identical playbook: free stealth → real traffic → reveal. Nostalgia, not instruments.
  • The 100T tokens/day promise sounds like a prior MiMo promo.
  • Some “dirty token” glitches look more MiMo than GLM, per one Reddit thread.
  • Dylan Fetch sees no 5.2→5.3 jump and suspects a lab he has not tested: Xiaomi.

Against

  • Tokenizer does not match MiMo on 6–25 prompt tests (it diverges; GLM does not).
  • API errors 1210/1214 belong to z.ai, not Xiaomi.
  • Pritish Yuvraj (ex-Llama/Meta): 17 multilingual prompts tokenize like GLM, not MiMo.
  • Manifold: ~6%.
  • Tokenizer L1=176 vs GLM=0 (Zhou). unclecode: mimo-v2.5 2/4, v2.5-pro 0/4; emoji 90 vs 57.
  • Video 2s/30fps/360p: Ox 296 vs MiMo v2.5 910. Audio: MiMo accepts; Ox rejects like GLM-5V.
  • Pritish 17 prompts: “MiMo didn’t match”. Hunter-Alpha-was-MiMo is an anti-prior: they already burned that animal.
Analogy, example, trap

Analogy. The modus operandi fits. The knife, the toll and the audio do not.

5% · weak · Tencent (China)

Hy4

Showed up in a Reddit thread as a candidate, with no independent technical tests.

For

  • Tencent has compute and coding/agent models.
  • A dedicated r/opencode thread named it.
  • Hy4 gray-test in Yuanbao the same day (20 Aug): calendar coincidence, not instruments.

Against

  • Zero public fingerprints tying it.
  • No recent comparable stealth history on OpenRouter.
  • Hy3 context 262,144 vs Ox 1,048,576.
  • Hy4 gray-test in Yuanbao the SAME day (20 Aug) is a calendar coincidence: Tencent launched Hy4 through another door, not as a stealth animal.
  • Tokenizer is not Hunyuan. L1 vs GLM=0 kills Hy4-as-weights.
Analogy, example, trap

Analogy. A name someone shouted in a crowded room. Nobody else recognized it.

4% · weak · MiniMax (China)

MiniMax (unannounced)

Timing of an unannounced model could fit. Circumstantial.

For

  • Chinese lab with frontier and multimodal models.
  • The stealth series has been 4/4 Chinese.

Against

  • MiniMax tokenizer diverges in public tests.
  • No matching API errors or encoder.
  • Zhou L1=172 (random knife). aitrackerbot: MiniMax M3 did not match. unclecode MiniMax-M3 tokenizer 0/4.
  • M3 already ships under its own name on OpenRouter. An anonymous twin with a GLM vocab is not their playbook.
Analogy, example, trap

Analogy. They were in the city on the day of the crime. That is not evidence.

12% · plausible · Third party (Cursor, Anomaly, etc.)

GLM fine-tune

Someone could take GLM weights, post-train them and serve them under another brand. Cursor has done something like this. An Anomaly employee used an internal 1M-context model.

For

  • Explains GLM tokenizer + behavior not identical to 5.3.
  • Kingbench: 87.5% vs GLM-5.3’s 91.25% — similar, not a clone.
  • Slower than GLM-5.3 at coding: possible different stack.

Against

  • Serving 100T tokens/day free needs a hyperscaler, not a small fine-tuner.
  • 1210 errors and 97% cache point at z.ai’s stack, not a foreign wrapper.
  • ZCode is not in OpenRouter’s public top (Hermes, Claude Code, pi). The kitchen tell is 1210 + encoder, not the app ranking.
  • ByteDance-as-CDN or Cursor-wrapping-GLM still resolve to Z.ai on Manifold’s “who is behind” unless the wrapper puts its logo on the reveal.
Analogy, example, trap

Analogy. You can paint a Ferrari another color. The engine, chassis and serial number are still Ferrari — and here the workshop also looks original.

9% · weak · Cursor / Anysphere

Composer 3 (Cursor)

On 21 Aug @TimMacc argued the only story that checks every box is Composer 3: GLM-tuned, huge compute, right timing. Cursor announced Composer 3 in June 2026 as a from-scratch 1.5T-class model. The from-scratch thesis dies on the GLM vocab; only a GLM wrapper survives — and that still collides with 1210 and the 5V encoder.

For

  • Cursor already ships Composer, historically GLM/Kimi-adjacent fine-tunes.
  • Timing: a new Composer drop plus a free OpenRouter week.
  • Would explain “GLM underneath” plus extra coding RL.

Against

  • Composer 3 was sold as from-scratch 1.5T on Colossus. A from-scratch would not share a GLM vocab, a 5V-Turbo encoder or error 1210.
  • Chinese error logs + ModelScope GLM strings (Josh rad) do not look like Cursor’s US stack.
  • 1210/1214 and 97% cache are z.ai kitchen, not Cursor’s.
  • No Cursor employee has claimed Ox Alpha.
  • Grok 4.6 already launched 12 Aug (500K ctx, 0 emoji/1k). Two anonymous frontier drops from the same cluster in 8 days, one free on a rival IDE, is a stretch.
  • Product: SpaceX bought Anysphere. Gifting the flagship on OpenCode (a competitor) with ZDR theater is odd; the stealth channel is a Chinese ritual.
Analogy, example, trap

Analogy. Someone says the tarped car is a Tesla with a Toyota engine. Maybe. The toll, dashboard error and radio are still Toyota.

6% · weak · Unknown

Multi-model router

A system that swaps models by task would explain inconsistency (sometimes brilliant, sometimes timeout).

For

  • Testers report irregular behavior across task types.
  • 17% timeouts in one run.

Against

  • A router does not produce a constant +75 tokenizer versus ONE concrete model.
  • A fixed video encoder does not fit a cocktail of backends.
Analogy, example, trap

Analogy. If you mix three cooks, the menu changes. Here the token recipe is always the same.

2% · ruled out · Mira Murati

Inkling (Thinking Machines)

Inkling shipped 15–16 Jul 2026 as Thinking Machines Lab’s open-weights multimodal. It is not stealth, already has a name, paper and weights. It is not Ox Alpha.

For

  • Mira leads a frontier lab. Inkling is strong at agents and 1M context.

Against

  • Inkling is PUBLIC (open-weights). Ox Alpha is anonymous on purpose.
  • Inkling: text+image+audio. Ox Alpha: text+image+video, no audio (rejects audio like GLM).
  • Different architecture: Inkling ~975B/41B active vs GLM-5 ~744B/40B.
  • Tokenizer is not GLM. Nobody serious on X ties it to Murati.
  • Dates: Inkling July; Ox Alpha 20 August, OpenRouter stealth channel used by Chinese labs.
  • On 21 Aug Thinking Machines tweeted Inkling FREE on OpenRouter WITH ITS NAME. Co-listed. Audio-native vs video-native. Opposite data policies (Inkling trains on traces; Ox’s sheet says no).
Analogy, example, trap

Analogy. Confusing Inkling with Ox Alpha is seeing a red Ferrari in the showroom and thinking the tarped car in the lot is the same. One already has plates.

2% · ruled out · Western labs

OpenAI / Anthropic / Grok / Gemini

Some said it as a joke (“it’s Gemini 3.5”). Western tokenizers do not match.

For

  • 100T/day capacity sounds like a hyperscaler (Google, Microsoft, xAI).

Against

  • GPT / Claude / Grok / Gemini tokenizers diverge hard.
  • Backend messages in Chinese.
  • 4 of 4 prior stealths on this channel were Chinese labs.
  • OpenRouter lists xAI separately (Grok 4.6). No reason to hide it as Ox.
  • Grok 4.6 = 500K ctx vs Ox 1,048,576. Grok style ~0 emoji/1k; Ox ~1.3 (GLM/Qwen house).
  • PRC canary: some testers saw Tiananmen denial / Tibet “inalienable part”; that is a PRC host filter, not a Grok signature.
Analogy, example, trap

Analogy. The accent, the dictionary and the office do not match. Case closed for the West.

6% · weak · ByteDance (China)

ByteDance Seed / Doubao

Dan (@DanDr1s) argued capacity: Ox served ~3.85T in a day (OR+OpenCode). Doubao claims 120T+/day and ~46% of Chinese cloud AI traffic. Video+coding is a Seed product. A US host would dodge the boycott. Escape: ByteDance serving an improved GLM.

For

  • Best ECONOMIC argument of any rival: 100T/day is rounding error for Volcano, a moonshot for a mid-size lab.
  • Seed-2.0 multimodal + Seedance video + Seed-2.0-Code fit Ox I/O (native video, agents).
  • “Not who you think”: GLM is the favorite; ByteDance was barely in the market.

Against

  • Tokenizer L1=0 vs GLM-5.3, encoder = GLM-5V-Turbo, error 1210 is z.ai code — not Volcano.
  • Seed-2.0-Code already has a public bytedance-seed/* SKU on OpenRouter (Inkling problem: co-listed).
  • If ByteDance only rents GPUs, Manifold “who is behind it” is still Z.ai. If they fine-tuned GLM and serve it on the z.ai stack, that is a conspiracy without evidence.
  • Style + DeepSWE steps (~117 vs GLM 124) track GLM, not consumer Doubao.
Analogy, example, trap

Analogy. The best invoice argument. The worst instruments argument. The GPU landlord does not sign the tenant’s ID.

4% · weak · Anomaly / OpenCode

Anomaly (OpenCode) in-house

OpenCode announced Ox in the first person (“We have capacity for 100T/day”). An employee used an internal 1M model. Theory: a GLM fine-tune on harvested free-agent traces.

For

  • They are the channel. Zen + OpenCode Go. Real traffic (Ox #4, ~3.6–3.9T, ~79k users in a day).
  • Fine-tune-on-harvested-traces is a coherent small-lab story after four free Chinese drops.

Against

  • They do not own 100T/day. OpenCode is a client. Zen is a router. 40B-active MoE at 49–78 tok/s is hyperscaler class. r/opencode: Z.ai stood up the 1 GW center last month.
  • Error 1210 / invalid zstd / chat.z.ai 1210/1214 are z.ai codes, not Anomaly’s. A self-host does not leak those.
  • Tokenizer + encoder = a GLM stack serving, not an Anomaly checkpoint with a new scaffold.
  • OpenCode’s ZDR clashes with OpenRouter stealth terms (the provider retains prompts). A reseller mismatch, not a first-party one.
Analogy, example, trap

Analogy. Anomaly is the loudspeaker. Z.ai is the singer. The “we” of 100T is wholesaler language, same as the four prior stealths.

Kill matrix

Instruments, not vibes. A “kill” cell is not hate: it is a test that suspect cannot explain.

TestGLMMiMoMiniMaxHy4ComposerInklingWestSeed
Tokenizer +75 / L1
L1=0 vs GLM-5.3. MiMo 176, MiniMax 172, others 148–302. From-scratch does not inherit the knife.
296 toll
Ox = GLM-5V-Turbo token-for-token. MiMo 910, Qwen 408. fps-invariance is common; the budget is not.
Audio
Ox rejects. MiMo v2.5 accepts. Inkling is audio-native. Softer than 296 (few raw dumps).
1210 + Java PaaS
Bilingual 1210 screenshot on Ox. com.wd.paas is Zhipu. 1214-on-Ox is a claim (Pritish), not a photo.
Stealth playbook
4/4 2026 animals = Chinese labs. Xiaomi already burned Hunter. Cursor from-scratch does not use this channel for a flagship.
100T factory
Capacity, not a meter. ByteDance already claims 120T/day — best waiter. Does not sign the chef’s passport.

Pick a cell to read the instrument.