Mastodon Skip to content
live markets
S&P 5007,674.37▲ 2.20%NASDAQ26,180.46▲ 1.33%DOW53,277.00▲ 2.02%GOLD4,680.60▲ 14.97%WTI87.06▲ 2.53%BRENT94.39▲ 3.71%EUR/USD1.1678▲ 2.28%USD/JPY158.94▼ 2.18%DXY98.84▼ 2.31%BTC$77,096▼ 0.60%ETH$2,417▲ 0.80%SOL$93.59▲ 1.30%TOTAL CRYPTO$2.62T▼ 2.23%
pulseofnations.
UTC --:--NYC --:--LON --:--WAW --:-- bluesky ↗ Join the wire

Mystery AI Model ‘Ox Alpha’ Tops GPT-5.6 in Coding Tests

An anonymous model on OpenRouter beats GPT-5.6 and Claude in coding benchmarks, with fingerprinting pointing 90% to Zhipu AI’s GLM-5.3.

Partner Surfshark VPN

An anonymous AI model designated “stealth/ox-alpha” appeared on OpenRouter on August 20, offered free for a limited preview window, and immediately outperformed GPT-5.6 and Claude on coding benchmarks. Independent researchers have since run technical fingerprinting tests that point with roughly 90% confidence to Zhipu AI’s GLM-5.3 as the model behind the mask, though no party has officially confirmed the link.

The model offers a 1-million-token context window, accepts text, image, and video inputs, and charges nothing during its preview period. Within days, OpenRouter’s dashboard showed it processing billions of tokens from coding tools including Claude Code and Hermes Agent, making it one of the most heavily used free models on the platform.

Fingerprinting the anonymous model

Researcher @aitrackerbot on X ran 25 diverse prompts through Ox Alpha and compared exact token counts against known tokenizers for candidate models. The result: Ox Alpha’s native token counts matched GLM-5.3 exactly, offset by a constant 75-token system prompt wrapper that appeared on every request.

The same researcher then tested four controlled video inputs, and Ox Alpha’s video token spending matched Zhipu’s GLM-5V-Turbo model token-for-token across three independent design parameters: frame sampling, duration scaling at roughly 147 tokens per second, and per-frame resolution scaling. Three separate engineering choices matching exactly is difficult to attribute to coincidence.

The team also tested MiMo v2.5, Qwen 3.8 Max, and GLM-4.6V against the same suite. All three produced clearly different signatures. MiMo was additionally ruled out by a binary behavioral difference: Ox Alpha rejects audio input entirely, while MiMo v2.5 accepts and tokenizes it.

Why go anonymous?

The strategy follows a pattern established earlier in 2026. Two previous OpenRouter stealth models, Hunter Alpha and Healer Alpha, were both eventually revealed as Xiaomi MiMo releases. Anonymous preview listings let labs collect massive-scale usage data, benchmark results, and failure mode information without the reputational pressure of an official launch.

Zhipu AI has stayed silent. OpenRouter lists the model under a generic “Stealth” provider with no attribution. OpenCode’s co-founder posted a cryptic message on X suggesting familiarity with the model’s identity but offered no official confirmation.

The timing is notable. Zhipu committed to releasing GLM-5.3 open weights roughly two weeks after its August 14 launch, pointing at a reveal around the week of August 24. If the stealth listing is indeed a GLM-5.3 checkpoint, the free preview window may be designed to generate momentum ahead of that formal release.

The broader trend matters more than any single model’s identity. Anonymously released frontier-class models achieving production adoption within hours signal a shift in how AI labs test and distribute capabilities. The era of quiet model drops beating established players on benchmarks suggests the competitive gap between named labs and unknown entrants is narrower than publicly available leaderboards indicate.

Sources: explainx.ai; OpenRouter model catalog; @aitrackerbot fingerprinting analysis on X; benchable.ai benchmark data; capitalandcompute.net model release tracker.

React to this dispatch
Share this dispatch X WhatsApp Report an error

discussion

Join the discussion

Your email address will not be published. Required fields are marked *

Next dispatch France Picks Mistral for AI Security Tests, Bans OpenAI Read →