🇬🇧 English · Home · Deutsche Version
Fleet Navigator runs on consumer hardware – no data centre required
Market overview: hardware for local LLM inference
As of 26 August 2026 · Price sources: Geizhals.de, Apple Store Germany, manufacturers‘ direct sales
| Device | Memory | Bandwidth | Street price (Germany) | Status / source |
|---|---|---|---|---|
| NVIDIA DGX Spark FE (GB10) | 128 GB LPDDR5X unified | 273 GB/s | from €5,785 | Geizhals, price as of 26 Aug 2026 |
| ASUS Ascent GX10 | 128 GB unified | 273 GB/s | from €4,292 | Geizhals, cheapest GB10 variant |
| Lenovo ThinkStation PGX | 128 GB unified | 273 GB/s | from €4,729 | Geizhals |
| Gigabyte AI TOP ATOM | 128 GB unified | 273 GB/s | from €5,030 | Geizhals |
| Mac Studio M5 Ultra 256 GB | 256 GB unified | ~1,200 GB/s | €10,999 | Apple list price; 512 GB from the end of October |
| Mac Studio M5 Ultra (base) | 96 GB unified, 1 TB SSD | ~1,200 GB/s | €6,599 | 30-core CPU / 64-core GPU; 256 GB = +€4,400 |
| Mac Studio M5 Max (base) | 36 GB unified, 512 GB SSD | lower | €2,999 | Entry-level model |
| RTX PRO 6000 Blackwell WS | 96 GB GDDR7 ECC | 1,792 GB/s | from €12,999 | Geizhals, card only (600 W) |
| AMD Ryzen AI Halo (reference) | 128 GB LPDDR5X-8000 + 2 TB | 256 GB/s | ~€3,700 (US) | Micro Center exclusive, since 6 Jul 2026; no German listing |
| GMKtec EVO-X2 | 128 GB + 2 TB SSD | 256 GB/s | from €3,599 | Geizhals, price as of 17 Aug 2026 |
| Framework Desktop 128 GB | 128 GB LPDDR5X | 256 GB/s | €3,449 | Framework direct sales |
| Beelink GTR9 Pro | 128 GB + 2 TB, 2x 10GbE | 256 GB/s | ~$4,349 | currently no German listing on Geizhals |
| Framework Desktop 192 GB (Max+ 495) | 192 GB LPDDR5X-8533 | 273 GB/s | expected €4,000–5,000+ | Q3 2026, price and date not yet confirmed |
| Xiaomi AI Cube | up to 160 GB shared | O100: 1.22 TB/s | no price | Prototype; O100/D100 not before 2027 |
All prices are snapshots. The memory market is currently moving week by week – always check again before you order.
What the numbers really mean
Bandwidth determines tokens per second, not TOPS
During token generation, the entire model has to be read from memory once for every single token. That puts the theoretical upper limit at roughly: tokens/s ≈ memory bandwidth / model size in bytes
A 70B model in Q4 quantisation takes up around 40 GB. On Strix Halo with 256 GB/s, that works out at about 6 tokens per second in theory. In practice, reviewers measure around 10 to 14 tokens per second on 70B models. The discrepancy arises because mixture-of-experts models do not touch all of their weights for every token. On the Mac Studio M5 Ultra with around 1.2 TB/s, you get four to five times that. This is the real difference between the three price classes, not the advertised NPU performance.
Prefill is the second, often underestimated factor
When a long context is being read in, what counts is compute, not bandwidth. This is exactly where the DGX Spark justifies its premium, and the RTX PRO 6000 with 24,064 CUDA cores and 1,792 GB/s is in a league of its own. For RAG scenarios with large document contexts, prefill performance matters more than raw generation speed. If you load many chunks into the context, you will notice the difference in time to first token far more than in output speed.
96 GB usable, not 128 GB
On all Strix Halo boxes, about 96 of the 128 GB can be addressed as GPU memory under Linux. The rest is reserved for the system. Apple’s split is more generous. If you plan with the full 128 GB as your usable model budget, your sums will not add up.
The AMD Halo family
AMD’s own reference box is called the Ryzen AI Halo Developer Platform: Ryzen AI Max+ 395 (16 Zen 5 cores, 32 threads), Radeon 8060S iGPU with 40 compute units (RDNA 3.5), 128 GB LPDDR5X-8000 unified, 2 TB SSD, 10GbE, Wi-Fi 7, 120 W system power. Officially $3,999.99 – so far a Micro Center exclusive in the USA, with no German offers.
What you can actually buy are the OEM variants with identical silicon: GMKtec EVO-X2, Framework Desktop, Beelink GTR9 Pro, HP Z2 Mini G1a. Since they all use the same chip, they differ only in cooling, networking, repairability and support.
| Criterion | Recommendation |
|---|---|
| Cheapest way into 128 GB | GMKtec EVO-X2 |
| Homelab, quiet under sustained load, 2x 10GbE | Beelink GTR9 Pro |
| Repairability and upgradeability | Framework Desktop |
| Business IT with manufacturer warranty | HP Z2 Mini G1a |
The more powerful configuration with Ryzen AI Max+ 495, 192 GB RAM and 160 GB VRAM arrives in the third quarter of 2026. If you need 192 GB, it is worth the wait. If 128 GB is enough, there is no reason to wait, because the Max+ 395 stays in the range unchanged.
Xiaomi: two stories that keep getting mixed up
Star Core Super AI Computer: Listed on Xiaomi’s crowdfunding platform Youpin, but made not by Xiaomi but by the Chinese manufacturer Linglong. Also a Ryzen AI Max+ 395 with 128 GB, from CNY 13,999 – so no chip of its own, just another Strix Halo enclosure. An international launch is not planned.
Xiaomi AI Cube: Unveiled on 24 August 2026 at Xiaomi’s technology conference for Xring chips – a prototype with three in-house chips and a total power draw of 150 W:
| Chip | Role | Key specs |
|---|---|---|
| Xring O3 | Main processor | 10-core CPU, G2 Ultra NX GPU with 16 cores, NPU with 200 TOPS, 3 nm |
| Xring O100 | Bandwidth-bound workloads | 6 nm, 3D wafer-level stacking, up to 1.22 TB/s near-memory bandwidth |
| Xring D100 | Compute-intensive AI | 3 nm, 20-core CPU, 16-core NPU, up to 160 GB shared memory, models up to 200B |
Our assessment: The O100 and D100 will not reach the market until 2027, and there is no price or date for the AI Cube. So it is not relevant for a purchasing decision this year. For watching the market, it is: if China builds its own memory architecture with 1.22 TB/s, NVIDIA’s pricing structure will come under pressure in the medium term.
Why launch prices and today’s prices are so far apart
The cause is the DRAM and NAND price crisis. The driver is the DRAM contract market, not manufacturer margins. For SSDs, price levels in August 2026 are on average 126 per cent above where they stood in mid-September 2025.
| Product | Launch price | Today |
|---|---|---|
| Framework Desktop 128 GB | $1,999 | $3,449 / €3,449 |
| GMKtec EVO-X2 128 GB | ~$2,000 | $3,649.99 |
| Beelink GTR9 Pro 128 GB | ~$1,999 | $4,349 |
| RTX PRO 6000 Blackwell (list) | $13,250 | $16,000 (+20.8%) |
What this means for purchasing: If you need 128 GB of unified memory and can afford the price, you are better off buying now than hoping for prices to fall. Forecasts for the rest of 2026 point upwards.
Decision guide
| Use case | Recommendation | Why |
|---|---|---|
| Development, fine-tuning experiments | ASUS Ascent GX10 (€4,292) | CUDA compatibility, which neither Strix Halo nor Apple offers |
| Inference demos at the customer’s site | Framework Desktop / GMKtec EVO-X2 (from €3,449) | An ordinary x86 machine with Windows or Linux, no special distribution |
| High generation speed (>15 tok/s at 70B) | Mac Studio M5 Ultra (from €6,599) | The only option with ~1.2 TB/s, but costs two to three times as much |
| Maximum prefill performance, short contexts | RTX PRO 6000 (from €12,999) | 1,792 GB/s, but only 96 GB and 600 W |
| 192 GB memory requirement | Wait for the Max+ 495 (Q3 2026) | Expected price €4,000–5,000+ |
Sources (prices as of 26 Aug 2026)
- Geizhals.de: NVIDIA DGX Spark Founders Edition, ASUS Ascent, Lenovo ThinkStation PGX, Gigabyte AI TOP ATOM, RTX PRO 6000 Blackwell, GMKtec EVO-X2, AMD Ryzen AI Halo
- ComputerBase: Mac Studio with M5 Ultra · Macwelt: Mac Studio now with M5 Max and M5 Ultra
- PCGamesHardware: Ryzen AI Halo: AMD’s mini PC at a maxi price · Memory and storage prices at record levels
- Notebookcheck: Xiaomi mini PC with 3 processors · Framework Desktop with 192 GB
- ComputingForGeeks: Ryzen AI Max+ 395 Mini PCs Compared · MiniXPC: Xiaomi crowdfunding Star Core
Our recommendation: Mac Mini M4 or M5 Pro
For most practices and law firms, the Mac Mini Pro with 64 GB of unified memory is the ideal solution:
- Price, Mac Mini M4 Pro: under €2,800 as a one-off investment – because of the 2026 memory price rally, check the current price on the day before you order
- Price, Mac Mini M5 Pro: Apple’s current model, prices on request
- Performance: Runs 30B models smoothly at about 30 tokens per second
- Unified memory: GPU and CPU share the same memory, ideal for large language models
- Power consumption: 60 to 80 watts, far less than a GPU server
- Noise: Whisper-quiet, fits on any desk
You keep working at your usual workstation
You do not have to work on the Mac Mini itself. Fleet Navigator routes inference across your local network to other computers. You use your Windows PC, your laptop or your office Mac as usual, while the AI computes in the background on the Mac Mini.
For this you need a Fleet Navigator Enterprise licence on the Mac Mini and a Fleet Navigator Maat licence on each additional workstation. That way several workstations share the same local AI instance, and the data never leaves your office.
It works just the same with Linux machines and multi-GPU workstations as the inference server. Here too, the workstations need a Fleet Navigator Maat licence and can use the routed inference on several other computers. This also applies to Linux machines with multiple GPUs.
What do 7B, 12B, 30B and 70B mean?
| Model size | Good for | Hardware needed |
|---|---|---|
| 3B–7B | Short summaries, keywords, simple dictation | Any modern PC or laptop (16 GB RAM) |
| 12B | Contract drafts, evaluating patient histories, patient letters | Mac Mini M4 standard model, mid-range GPU with 12 GB VRAM |
| 30B | Demanding legal submissions, medical differential diagnosis, complex tax cases | Mac Mini M4 Pro 64 GB, workstation with 24 GB+ VRAM |
| 70B | Very long contracts, comparative research | Two GPUs with 24 GB each, or a Mac Studio |
The number stands for the parameters of the AI model – more parameters = more knowledge, but more hardware required
Configurations for every requirement – our recommendations for your hardware
From entry level to a professional multi-user solution
Entry level
from ~€1,500
Standard PC or laptop
- RAM: 32 GB
- GPU: at least 6 GB VRAM
- SSD storage: 50 GB free
- OS: Windows 11, Linux or macOS
- Suitable for 7B and 12B models
- Ideal for: Light Edition
⭐ Recommended
Recommended
from ~€2,800
Mac Mini M4/M5 Pro
- RAM: 32–64 GB
- GPU: NVIDIA RTX 3060+ (min. 12 GB VRAM)
- Alternative: Apple Mac M4+
- Storage: 20 GB free
- Suitable for 12B–30B models
- Ideal for: Enterprise Edition
Professional
from €3,449
Strix Halo mini PC (Ryzen AI Max+ 395)
- RAM: 128 GB unified memory (~96 GB usable for models)
- GPU: Radeon 8060S iGPU (40 compute units, RDNA 3.5)
- Recommendation: Framework Desktop €3,449 or GMKtec EVO-X2 from €3,599
- Storage: 50 GB+ free
- Suitable for 30B–70B models (70B: 10–14 tok/s)
- Ideal for: Navigator Enterprise
High end: maximum performance for 70B models
For very large models, several concurrent users and maximum speed
🍎 Apple Silicon
Mac Studio
from €6,599
Apple Mac Studio M5 Ultra
- RAM: 96–256 GB unified memory
- Chip: Apple M5 Ultra, ~1.2 TB/s memory bandwidth
- Storage: from 1 TB SSD
- Suitable for 70B models (>15 tok/s)
- Power consumption: low, whisper-quiet
- Ideal for: Enterprise Edition
⚡ Maximum performance
NVIDIA DGX Spark
from €4,292
GB10 family: ASUS Ascent GX10 from €4,292, NVIDIA FE from €5,785
- Memory: 128 GB unified memory
- Chip: NVIDIA GB10 Grace Blackwell
- AI performance: up to 1 petaFLOP (FP4)
- Suitable for models up to 70B+
- Ecosystem: full CUDA / NVIDIA AI
- Ideal for: Enterprise & multi-user
Not sure which hardware?
We advise you free of charge – tailored to your edition and your requirements