Beliebte Suchanfragen:

🇬🇧 English · Home · Deutsche Version

Fleet Navigator runs on consumer hardware – no data centre required


Market overview: hardware for local LLM inference

As of 26 August 2026 · Price sources: Geizhals.de, Apple Store Germany, manufacturers‘ direct sales

DeviceMemoryBandwidthStreet price (Germany)Status / source
NVIDIA DGX Spark FE (GB10)128 GB LPDDR5X unified273 GB/sfrom €5,785Geizhals, price as of 26 Aug 2026
ASUS Ascent GX10128 GB unified273 GB/sfrom €4,292Geizhals, cheapest GB10 variant
Lenovo ThinkStation PGX128 GB unified273 GB/sfrom €4,729Geizhals
Gigabyte AI TOP ATOM128 GB unified273 GB/sfrom €5,030Geizhals
Mac Studio M5 Ultra 256 GB256 GB unified~1,200 GB/s€10,999Apple list price; 512 GB from the end of October
Mac Studio M5 Ultra (base)96 GB unified, 1 TB SSD~1,200 GB/s€6,59930-core CPU / 64-core GPU; 256 GB = +€4,400
Mac Studio M5 Max (base)36 GB unified, 512 GB SSDlower€2,999Entry-level model
RTX PRO 6000 Blackwell WS96 GB GDDR7 ECC1,792 GB/sfrom €12,999Geizhals, card only (600 W)
AMD Ryzen AI Halo (reference)128 GB LPDDR5X-8000 + 2 TB256 GB/s~€3,700 (US)Micro Center exclusive, since 6 Jul 2026; no German listing
GMKtec EVO-X2128 GB + 2 TB SSD256 GB/sfrom €3,599Geizhals, price as of 17 Aug 2026
Framework Desktop 128 GB128 GB LPDDR5X256 GB/s€3,449Framework direct sales
Beelink GTR9 Pro128 GB + 2 TB, 2x 10GbE256 GB/s~$4,349currently no German listing on Geizhals
Framework Desktop 192 GB (Max+ 495)192 GB LPDDR5X-8533273 GB/sexpected €4,000–5,000+Q3 2026, price and date not yet confirmed
Xiaomi AI Cubeup to 160 GB sharedO100: 1.22 TB/sno pricePrototype; O100/D100 not before 2027

All prices are snapshots. The memory market is currently moving week by week – always check again before you order.

What the numbers really mean

Bandwidth determines tokens per second, not TOPS

During token generation, the entire model has to be read from memory once for every single token. That puts the theoretical upper limit at roughly: tokens/s ≈ memory bandwidth / model size in bytes

A 70B model in Q4 quantisation takes up around 40 GB. On Strix Halo with 256 GB/s, that works out at about 6 tokens per second in theory. In practice, reviewers measure around 10 to 14 tokens per second on 70B models. The discrepancy arises because mixture-of-experts models do not touch all of their weights for every token. On the Mac Studio M5 Ultra with around 1.2 TB/s, you get four to five times that. This is the real difference between the three price classes, not the advertised NPU performance.

Prefill is the second, often underestimated factor

When a long context is being read in, what counts is compute, not bandwidth. This is exactly where the DGX Spark justifies its premium, and the RTX PRO 6000 with 24,064 CUDA cores and 1,792 GB/s is in a league of its own. For RAG scenarios with large document contexts, prefill performance matters more than raw generation speed. If you load many chunks into the context, you will notice the difference in time to first token far more than in output speed.

96 GB usable, not 128 GB

On all Strix Halo boxes, about 96 of the 128 GB can be addressed as GPU memory under Linux. The rest is reserved for the system. Apple’s split is more generous. If you plan with the full 128 GB as your usable model budget, your sums will not add up.

The AMD Halo family

AMD’s own reference box is called the Ryzen AI Halo Developer Platform: Ryzen AI Max+ 395 (16 Zen 5 cores, 32 threads), Radeon 8060S iGPU with 40 compute units (RDNA 3.5), 128 GB LPDDR5X-8000 unified, 2 TB SSD, 10GbE, Wi-Fi 7, 120 W system power. Officially $3,999.99 – so far a Micro Center exclusive in the USA, with no German offers.

What you can actually buy are the OEM variants with identical silicon: GMKtec EVO-X2, Framework Desktop, Beelink GTR9 Pro, HP Z2 Mini G1a. Since they all use the same chip, they differ only in cooling, networking, repairability and support.

CriterionRecommendation
Cheapest way into 128 GBGMKtec EVO-X2
Homelab, quiet under sustained load, 2x 10GbEBeelink GTR9 Pro
Repairability and upgradeabilityFramework Desktop
Business IT with manufacturer warrantyHP Z2 Mini G1a

The more powerful configuration with Ryzen AI Max+ 495, 192 GB RAM and 160 GB VRAM arrives in the third quarter of 2026. If you need 192 GB, it is worth the wait. If 128 GB is enough, there is no reason to wait, because the Max+ 395 stays in the range unchanged.

Xiaomi: two stories that keep getting mixed up

Star Core Super AI Computer: Listed on Xiaomi’s crowdfunding platform Youpin, but made not by Xiaomi but by the Chinese manufacturer Linglong. Also a Ryzen AI Max+ 395 with 128 GB, from CNY 13,999 – so no chip of its own, just another Strix Halo enclosure. An international launch is not planned.

Xiaomi AI Cube: Unveiled on 24 August 2026 at Xiaomi’s technology conference for Xring chips – a prototype with three in-house chips and a total power draw of 150 W:

ChipRoleKey specs
Xring O3Main processor10-core CPU, G2 Ultra NX GPU with 16 cores, NPU with 200 TOPS, 3 nm
Xring O100Bandwidth-bound workloads6 nm, 3D wafer-level stacking, up to 1.22 TB/s near-memory bandwidth
Xring D100Compute-intensive AI3 nm, 20-core CPU, 16-core NPU, up to 160 GB shared memory, models up to 200B

Our assessment: The O100 and D100 will not reach the market until 2027, and there is no price or date for the AI Cube. So it is not relevant for a purchasing decision this year. For watching the market, it is: if China builds its own memory architecture with 1.22 TB/s, NVIDIA’s pricing structure will come under pressure in the medium term.

Why launch prices and today’s prices are so far apart

The cause is the DRAM and NAND price crisis. The driver is the DRAM contract market, not manufacturer margins. For SSDs, price levels in August 2026 are on average 126 per cent above where they stood in mid-September 2025.

ProductLaunch priceToday
Framework Desktop 128 GB$1,999$3,449 / €3,449
GMKtec EVO-X2 128 GB~$2,000$3,649.99
Beelink GTR9 Pro 128 GB~$1,999$4,349
RTX PRO 6000 Blackwell (list)$13,250$16,000 (+20.8%)

What this means for purchasing: If you need 128 GB of unified memory and can afford the price, you are better off buying now than hoping for prices to fall. Forecasts for the rest of 2026 point upwards.

Decision guide

Use caseRecommendationWhy
Development, fine-tuning experimentsASUS Ascent GX10 (€4,292)CUDA compatibility, which neither Strix Halo nor Apple offers
Inference demos at the customer’s siteFramework Desktop / GMKtec EVO-X2 (from €3,449)An ordinary x86 machine with Windows or Linux, no special distribution
High generation speed (>15 tok/s at 70B)Mac Studio M5 Ultra (from €6,599)The only option with ~1.2 TB/s, but costs two to three times as much
Maximum prefill performance, short contextsRTX PRO 6000 (from €12,999)1,792 GB/s, but only 96 GB and 600 W
192 GB memory requirementWait for the Max+ 495 (Q3 2026)Expected price €4,000–5,000+
Sources (prices as of 26 Aug 2026)

Our recommendation: Mac Mini M4 or M5 Pro

For most practices and law firms, the Mac Mini Pro with 64 GB of unified memory is the ideal solution:

You keep working at your usual workstation

You do not have to work on the Mac Mini itself. Fleet Navigator routes inference across your local network to other computers. You use your Windows PC, your laptop or your office Mac as usual, while the AI computes in the background on the Mac Mini.

For this you need a Fleet Navigator Enterprise licence on the Mac Mini and a Fleet Navigator Maat licence on each additional workstation. That way several workstations share the same local AI instance, and the data never leaves your office.

It works just the same with Linux machines and multi-GPU workstations as the inference server. Here too, the workstations need a Fleet Navigator Maat licence and can use the routed inference on several other computers. This also applies to Linux machines with multiple GPUs.


What do 7B, 12B, 30B and 70B mean?

Model sizeGood forHardware needed
3B–7BShort summaries, keywords, simple dictationAny modern PC or laptop (16 GB RAM)
12BContract drafts, evaluating patient histories, patient lettersMac Mini M4 standard model, mid-range GPU with 12 GB VRAM
30BDemanding legal submissions, medical differential diagnosis, complex tax casesMac Mini M4 Pro 64 GB, workstation with 24 GB+ VRAM
70BVery long contracts, comparative researchTwo GPUs with 24 GB each, or a Mac Studio

The number stands for the parameters of the AI model – more parameters = more knowledge, but more hardware required


Configurations for every requirement – our recommendations for your hardware

From entry level to a professional multi-user solution

Entry level

from ~€1,500

Standard PC or laptop

  • RAM: 32 GB
  • GPU: at least 6 GB VRAM
  • SSD storage: 50 GB free
  • OS: Windows 11, Linux or macOS
  • Suitable for 7B and 12B models
  • Ideal for: Light Edition

Professional

from €3,449

Strix Halo mini PC (Ryzen AI Max+ 395)

  • RAM: 128 GB unified memory (~96 GB usable for models)
  • GPU: Radeon 8060S iGPU (40 compute units, RDNA 3.5)
  • Recommendation: Framework Desktop €3,449 or GMKtec EVO-X2 from €3,599
  • Storage: 50 GB+ free
  • Suitable for 30B–70B models (70B: 10–14 tok/s)
  • Ideal for: Navigator Enterprise

High end: maximum performance for 70B models

For very large models, several concurrent users and maximum speed

🍎 Apple Silicon

Mac Studio

from €6,599

Apple Mac Studio M5 Ultra

  • RAM: 96–256 GB unified memory
  • Chip: Apple M5 Ultra, ~1.2 TB/s memory bandwidth
  • Storage: from 1 TB SSD
  • Suitable for 70B models (>15 tok/s)
  • Power consumption: low, whisper-quiet
  • Ideal for: Enterprise Edition

Not sure which hardware?

We advise you free of charge – tailored to your edition and your requirements

Book your appointment