AMD announced the Threadripper Halo Station at IFA 2026 in Berlin, a liquid-cooled workstation designed to run AI models with more than one trillion parameters locally, without a cloud connection.
The system pairs a 96-core Threadripper PRO 9995WX processor codenamed “Shimada Peak” with up to four AMD Instinct MI350P accelerators. AMD’s SVP Jack Huynh described it as a new class of workstation that brings supercomputer-class compute to individual users and developers.
AMD called it “the most powerful workstation in the world.” That claim comes with caveats – the system shown at IFA was a prototype, and pricing and availability have not been announced – but the specs are real and the target market is specific: AI developers who want to train, fine-tune, and run large models without paying for cloud GPU time.
What is inside
The base configuration shown at IFA included two MI350P PCIe cards for a total of 288 GB of HBM3E memory. Users can upgrade to four cards and 576 GB of HBM3E. AMD says that amount of accelerator memory is enough to hold a trillion-parameter model entirely on-device.
Each MI350P card features 144 GB of HBM3E running at 4 TB/s of memory bandwidth, which AMD says is 14 times higher than any LPDDR5X variant. The cards have a TBP of up to 600 W each and are liquid-cooled independently, as is the CPU.
The Threadripper PRO 9995WX is a Zen 5 chip with 96 cores and 192 threads, boosting to 5.4 GHz. It supports up to 2 TB of DDR5 system memory across eight channels at up to 6400 MT/s. Combined with the accelerator memory, total system memory reaches up to 2.6 TB, with aggregate GPU memory bandwidth of 16 TB/s and RDIMM bandwidth of 410 GB/s.
The system runs on standard desktop power infrastructure but will require dedicated liquid cooling loops for both the CPU and each accelerator card. That means the Halo Station is not a plug-and-play upgrade for existing workstations – it requires specific cooling infrastructure and physical space.
How it compares to Nvidia
AMD published a detailed comparison against Nvidia’s DGX Station, which uses a single GB300 Grace Blackwell Ultra Superchip. The Nvidia system offers 252 GB of HBM3e GPU memory at 7.1 TB/s peak bandwidth and 496 GB of LPDDR5X CPU memory at 396 GB/s.
The AMD configuration with four MI350P cards provides 576 GB of HBM3E at 16 TB/s aggregate bandwidth – roughly double the memory capacity and more than double the bandwidth of the Nvidia setup. The tradeoff is power and complexity: four liquid-cooled accelerator cards plus a liquid-cooled CPU versus a single integrated superchip.
Neither system has announced final pricing. Nvidia’s DGX Station has historically been priced in the six-figure range, and the AMD Halo Station will likely land in similar territory given the component costs.
The local AI pitch
AMD’s pitch is straightforward: cloud GPU access is expensive, shared, and sometimes unavailable. Running models locally eliminates latency, removes per-token costs, and gives developers full control over their training environment.
The practical use cases AMD highlighted include training and fine-tuning large language models, running agentic workflows that need continuous compute, and processing sensitive data that cannot leave a physical location. The company positioned the system at the intersection of traditional workstation workloads – 3D rendering, simulation, video editing – and AI development.
The challenge is that most AI development still happens in the cloud, where companies like AWS, Google Cloud, and Azure offer elastic scaling that a single workstation cannot match. AMD is betting that a segment of developers – particularly those working with sensitive data, on-premises requirements, or who simply want to avoid cloud bills – will pay for local compute.
Whether that bet pays off depends on pricing, availability, and whether the AI developer community sees enough value in local inference to justify the cost. For now, the Halo Station is a statement of intent from AMD: it wants a piece of the AI infrastructure market, and it is willing to build hardware that pushes the boundaries of what a single desk can hold.
The Halo Station also signals AMD’s broader strategy of competing with Nvidia not just on price but on openness. The system uses standard PCIe interfaces and AMD’s ROCm software stack, which is open source. That contrasts with Nvidia’s CUDA ecosystem, which dominates AI development but locks developers into proprietary tooling. AMD has been pushing ROCm compatibility for years, and the Halo Station gives it a flagship hardware platform to demonstrate that the software stack can handle serious workloads.
The timing matters too. AI development is shifting from training large models in the cloud to running inference and fine-tuning on local hardware. Companies like Apple, Google, and Microsoft are all investing in on-device AI capabilities, and AMD wants to capture the developer workstation market that feeds into those ecosystems. The Halo Station is not a consumer product – it is aimed at researchers, AI engineers, and creative professionals who need more power than a laptop but less than a data center.
Pricing will determine whether the Halo Station becomes a niche product or a genuine alternative to cloud GPU clusters. If AMD can undercut Nvidia’s DGX Station significantly while offering comparable or better specs, there is a real market for developers who want to own their compute rather than rent it. The liquid cooling requirement adds complexity, but many AI labs already have cooling infrastructure in place.

discussion