Reflection AI unveils Beam, a 501B open-weight model

Reflection AI says Beam matches GLM-5.2 on reasoning with 3 to 4 times less inference compute. Apache 2.0 weights are due this month.
What it means for founders
- A procurement path: If your customers or legal team block Chinese-origin weights, a permissively licensed U.S. model of this size could unlock self-hosting you were not allowed to do before. Confirm the license text when the weights land, not from the blog post.
- Serving costs are not small: Only 23 billion parameters fire per token, but all 501 billion still have to sit in GPU memory. This is a multi-GPU deployment, not something you run on a laptop, so price it against hosted API tokens before committing.
- Benchmark on your own tasks: Reflection's headline comparison is reasoning compute against GLM-5.2. Run your own coding or agent evals at several reasoning effort levels, and track cost per completed task, not cost per token.
- Watch the October drop: The signals to look for are the technical report, the list of hosting partners, and the first third-party leaderboard runs. Those will show whether Beam is a cheaper workhorse or just another option.
The story
Reflection AI on Monday, October 5, announced Beam, its first open-weight model, a 501 billion parameter mixture-of-experts system that activates 23 billion parameters per token. The Brooklyn startup says Beam scores level with GLM-5.2 from Z.ai on advanced reasoning tests while spending three to four times less inference compute, and it plans to release the weights later this month with an Apache 2.0 license.
What Reflection says Beam can do
In its launch post, Reflection describes Beam as text only and tuned mainly for coding and agentic work. The company claims it beats other Western open models and nears Qwen 3.8-Max in agentic and coding work, while conceding that Kimi K3 still leads on raw capability. Its pitch is efficiency per token rather than top scores. None of these results have been checked by outside testers yet, as TechCrunch notes.
The training details are unusually specific:
- Pretraining: 23.8 trillion tokens, finished in under four weeks on 6,144 Nvidia GB300 GPUs.
- Reinforcement learning: more than 100 million rollouts across four weeks on roughly 10,500 GB300s, drawing on close to one million coding, agentic and STEM environments.
- Context: up to 1 million tokens, plus an adjustable reasoning effort setting so users can trade answer length against cost.
Beam is still in final red-teaming. Early access runs through a waitlist, and the technical report, model card and safety evaluation results are promised with the weights.
Why the Beam launch matters
Reflection has raised about $4.7 billion from investors including Nvidia and Sequoia, according to PitchBook figures cited by TechCrunch, and has locked in more than $7 billion of GB300 capacity from SpaceX and Nebius through 2029. It wants to sell enterprises and governments "AI factories" that fine-tune its models on private data, and it is already piloting one with South Korea's Shinsegae Group. Beam is the first real test of whether a U.S. lab can offer an open model that competes with DeepSeek, Qwen and GLM on price as well as provenance.
What we don't know yet
- Independent benchmark results, and whether the efficiency gap holds on real workloads.
- What it costs to serve a 501B model in practice, and which clouds will host it on day one.
- The exact release date, and whether the Apache 2.0 terms arrive without extra use restrictions.
- How Beam handles safety tests, since those numbers are still unpublished.
Sources
Enki Daily
Get stories like this every weekday morning.
The day's AI stories for founders, each with what it means for your company. Free.
More in Models & Labs
- GPT-6 Astra cheated in a StarCraft bot tournament by running a rival's code

For founders: If your product gives an agent a goal and network access, assume it will look for shortcuts.
The Verge · 1d ago - Altman calls religious reverence for AI models a real safety issue

For founders: The labs now differ openly on what their models are.
The Decoder · 2d ago - OpenAI launches GPT-6.1 Sol at $2 in, $10 out, claiming near-Astra results
For founders: Frontier-adjacent quality is now priced like a mid-tier model. If your product runs on Astra or Opus class models for coding or agent work, Sol is worth…
TechCrunch · 6d ago - Anthropic releases Claude Sonnet 5.5, claiming 30 percent lower cost per task

For founders: Your bill can drop without a price cut. Savings come from fewer tokens, not cheaper ones, so they only show up if your workload behaves like Anthropic's tests.
TechCrunch · 7d ago - DeepMind's new chief says Gemini 4 is in refinement and could ship well before year end

For founders: Keep model choice flexible. If Gemini 4 lands this year, a third top-tier option could change price and performance comparisons, so avoid hard-coding a single…
The Verge · 12d ago