ASUS ExpertCenter Pro ET900N G3: local models for $100 thousand (Category AI4SDLC)
Looked at the published 30 July 2026 The Alex Ziskind Test of the YearThis was a data center a year ago… Now it's on my desk? It seems that their local models have never been so close. The little thing left: find $100 Thousands on your own desktop supercomputer ASUS ExpertCenter Pro ET900N G3.
And that's hardly an exaggeration. One of the American sellers indicated a price for the system $99,999.99. For this money we get a tower on the architecture of NVIDIA DGX Station: 72-nuclear Grace CPU, Blackwell Ultra GPU, before 20 PFLOPS in FP4 according to NVIDIA specifications 748 GB of coherent memory. A year ago, this amount of AI computing was more associated with a rack, and now you can put a box next to a desk. The maximum energy consumption of the system 1,6 KW is more of an office server than a new home PC.)
The most interesting thing here is memory. In the advertising line 748 GB looks like one huge common pool, but physically it consists of two very different parts: - 252 GB HBM3e throughput 7,1 TB/s; - 496 GB LPDDR5 X throughput 396 GB/s; Between Grace and Blackwell is a coherent NVLink-C2C, so the GPU can access both areas in a shared address space.
This removes the hard line “the model didn’t fit into VRAM – launch is impossible,” but it doesn’t negate physics. Quantized to NVFP4 GLM-5.2 about 465 GB and entirely in HBM does not fit. It runs locally, but parts of the scales have to be read from slower memory, so the speed is noticeably inferior to models that remain completely in HBM. Coherent memory is not magical. 748 GB is equally fast VRAM.
The second important part of the video is not the speed of one chat, but the competitive mode of the models. On NVIDIA Nemotron 3 Super 120B Alex launches up to 128 parallel agents. Depending on the model and the number of requests, the total performance reaches several thousand tokens per second. This works due to continuous batching: the GPU collects many requests and processes them together more efficiently. In this case, an individual agent waits longer - the total throughput does not grow for free.
And here the system becomes interesting not to the rich amateur, but to the team. A single local node can handle coding agents, research tasks, and internal AI services without sending code and data to an external API. It is possible to hold dozens of long-lived agents, experiment with large open models and get predictable infrastructure under high constant load.
But I would argue with the phrase “zero price token”. The price simply moves from the API account to capital expenditure, electricity, cooling, storage, upgrades and operation. And the local model must pass the same quality gate as the cloud. 128 Agents are useless if each of them quickly produces mediocre results.
For me, the main conclusion is that local AI has really changed class. Now "locally" can mean not a small model on a laptop, but hundreds of billions of parameters and a whole team of agents on one machine. Iron will not design agent architecture for us, nor will it solve the quality problem, but it raises the ceiling of what can be assembled in its own circuit.
Their local models have never been so close. Especially if you have a lot of money.100 thousands, a separate power line and a very compelling business case:)
P.S. And my toad only forked for the computer.Beelink GTR9 Pro AMD Ryzen™ AI Max+ 395" (here heybut already in Russia)Which is also useful for experiments:)
#AI #AI4SDLC #Agents #Engineering #Architecture #Hardware