Reflection AI, a Brooklyn-based startup founded by former Google DeepMind researchers, launched Beam on October 5, its first frontier open-weight AI model, built to challenge Chinese open models like DeepSeek, Qwen and GLM on their own terms: performance at a fraction of the compute cost.
Beam is a text-only mixture-of-experts model with 501 billion total parameters and 23 billion active per token, giving it the serving cost of a much smaller model while keeping frontier-scale capacity. Reflection pretrained it on 23.8 trillion tokens in under four weeks on 6,144 Nvidia GB300 GPUs, reporting 92.3 percent goodput, a measure of how much GPU time went to useful computation rather than failures and restarts. A second run used another 10,500 GB300 chips for reinforcement learning, generating more than 100 million rollouts across roughly a million synthetic coding, agentic and STEM environments. The context window reaches one million tokens.
The company claims Beam matches Z.ai’s GLM-5.2 on advanced reasoning benchmarks and outperforms today’s leading Western open models while using three to four times less inference compute. On coding benchmarks it reports 80.9 on SWE-Bench Verified and 80.1 on Terminal Bench v2.1. Weights and full technical details follow later this month under the Apache 2.0 license, distributed through hyperscalers and neoclouds, with early access open through a waitlist.
Read the fine print on the benchmarks
Reflection’s own comparison table shows the honest picture: DeepSeek V4.1 Flash, Moonshot’s Kimi K3 and GLM-5.3 beat Beam on most rows of raw capability, including Terminal Bench v2.1, where DeepSeek’s model scored 90.6 against Beam’s 80.1. The company’s pitch is not that Beam is the smartest model available. It is that Beam delivers comparable intelligence per token, with active-parameter counts far below rivals, at a price point enterprises can run at scale. A sparse model with 23 billion active parameters costs roughly as much to serve as a 23B dense model, which matters for anyone self-hosting.
All figures are self-reported and await independent verification. GLM-5.2 itself carries roughly 753 billion total parameters with 40 billion active, so Beam’s active-parameter advantage is real but narrower than the total-parameter gap suggests. Reflection’s efficiency figures also exclude prompt prefill and serving overhead, according to analysts who reviewed the announcement, so real-world savings depend on the workload. Kimi K3, for its part, runs a 2.8-trillion-parameter architecture with 104 billion active parameters, which puts the scale difference between the two camps in perspective.
A Western answer to DeepSeek
The gap Reflection is attacking is real. Chinese labs, led by DeepSeek and Alibaba’s Qwen, have dominated open-weight leaderboards since early 2025, leaving Western developers who want open models with few competitive options. The nearest American entries before Beam, Nvidia’s Nemotron 3 Ultra and Thinking Machines’ Inkling, trailed the Chinese frontier on most agentic benchmarks. Reflection, founded by Misha Laskin, who led reward modeling on DeepMind’s Gemini project, and Ioannis Antonoglou, who co-created AlphaGo, is backed by Nvidia and Sequoia and valued at $25 billion. Laskin argued on CNBC that enterprises want models they can inspect, fine-tune and deploy on their own infrastructure rather than rent through an API, and that open weights are a national-capability question as much as a business one.
Reflection’s timing also rides a policy tailwind. The US government’s push to keep frontier AI development domestic has translated into procurement interest for open-weight American models, and companies like Reflection have positioned themselves explicitly as the domestic alternative to Chinese open weights for government and regulated-industry workloads. Whether Beam wins that business will depend less on benchmark deltas of a few points and more on security reviews, deployment support and the credibility of its efficiency claims under independent measurement.
Lands amid an efficiency race
The launch arrives as AI startups shift toward cheaper models under price pressure, and days after Meta published a study showing two smaller coding agents can catch more bugs than one larger one. The trend across the industry points the same way: capability per dollar is becoming the metric that matters, not capability alone.
For developers deciding what to build on, the practical questions are cost, latency and control, and Beam’s numbers are aimed at all three. The Apache 2.0 license means commercial use is unrestricted from day one, the full fine-tuning stack ships with the weights, and integration with open-source libraries is planned at launch. Enterprises that hesitated on Chinese-licensed models over compliance concerns, or on Western closed APIs over cost, now have a third option to benchmark. The catch is that benchmarks age fast in this market; a model that leads on price-performance this month can be displaced by the next DeepSeek or Qwen release within a quarter, so the durable advantage, if there is one, will be in distribution and enterprise trust rather than in any single score.
There is also the hardware angle. Beam was trained end to end on Nvidia’s GB300 platform, and its launch doubles as a demonstration of what that hardware can do in a single month of pretraining. Nvidia’s backing of Reflection has always been about more than equity: a competitive Western open model keeps the training and serving demand on Nvidia silicon rather than migrating to domestic Chinese accelerators, which is exactly what happened with several recent Chinese open releases.
Beam’s full release later this month will show whether the benchmarks hold up outside Reflection’s own testing. If they do, Western open-weight developers finally have a domestic option near the frontier, and Chinese open models face their first serious competition in over a year. If they do not, Beam becomes another entry in a leaderboard it did not win. Either way, the efficiency argument is now the center of the open-model race.
