192GB is becoming a product category rather than a one-off experiment

IFA 2026 brought several compact systems built around AMD's Ryzen AI Max+ PRO 495.

GMKtec has the EVO-X5 Pro.

ACEMAGIC is offering an F9A PRO 495, while MINISFORUM is using the same processor across new AI-focused systems.

Lenovo also uses the chip in its ThinkCentre X Ultra, although its announced configurations currently stop at 128GB.

The common theme is no longer simply a faster mini PC. It is unusually large unified-memory pools aimed directly at local AI.

Up to 160GB can be assigned to Radeon 8065S graphics

The important architectural feature is shared memory.

On 192GB configurations from GMKtec and ACEMAGIC, as much as 160GB can be allocated to the integrated Radeon 8065S.

That is not 160GB of separate physical VRAM sitting beside 192GB of system memory.

CPU and GPU share the same LPDDR5X pool, and software can reserve a very large portion of it for graphics and AI workloads.

A flagship discrete GPU is faster, but it does not have 160GB of VRAM

This is where conventional performance comparisons become misleading.

A high-end discrete graphics card has much more raw compute and substantially greater memory bandwidth.

When a model fits comfortably in its VRAM, the discrete GPU will often be the faster option.

A model cannot take advantage of that compute if its weights do not fit in memory in the first place.

That is the problem these new systems are designed to avoid.

Large-model inference often hits capacity before compute

Language models contain billions of parameters that need to reside somewhere during inference.

Quantization reduces their storage requirement by representing those weights at lower precision.

That can make a 70-billion-parameter model relatively practical on local hardware.

At 200 or 300 billion parameters, memory requirements quickly become much larger even under aggressive quantization.

A graphics card with 24GB or 32GB of VRAM then needs to offload data to system memory or split the workload across several GPUs.

A 160GB GPU-accessible pool changes that constraint dramatically.

GMKtec claims support for local 300-billion-parameter models

GMKtec markets the EVO-X5 Pro as capable of running LLMs with as many as 300 billion parameters entirely offline.

That claim needs context.

Making a model fit into memory and making that model run quickly are two very different achievements.

Actual usability depends on quantization, backend software, context length and how much memory the operating system and other applications need.

A dense 300B model should not be expected to generate tokens at the same rate as a much smaller 30B model just because both technically load.

273GB/s is huge for integrated graphics and modest for a flagship GPU

GMKtec quotes approximately 273GB/s of memory bandwidth from its LPDDR5X-8533 implementation.

That is substantial for a compact unified-memory PC.

It remains far below the bandwidth available to some premium discrete graphics cards, which can exceed one terabyte per second.

Large language models can become strongly bandwidth-bound because generating each token requires repeatedly reading model weights.

The new architecture therefore answers “does the model fit?” more convincingly than it answers “is this the fastest way to run it?”

Ryzen AI Max+ PRO 495 is still an evolution of the existing Halo architecture

AMD has not replaced every core architecture.

Ryzen AI Max+ PRO 495 retains 16 Zen 5 CPU cores and 32 threads.

Radeon 8065S continues to use RDNA 3.5 with 40 compute units and reaches up to 3GHz.

The XDNA 2 NPU is rated at up to 55 TOPS.

The defining upgrade over Ryzen AI Max+ 395 is therefore memory: up to 192GB of LPDDR5X-8533 instead of the earlier platform's 128GB LPDDR5X-8000 ceiling.

131 TOPS is not the useful number for a giant LLM

Some system vendors advertise as much as 131 TOPS across the complete platform.

That figure combines several compute engines under specific conditions.

It is not the token-generation speed of a language model.

The dedicated XDNA 2 NPU itself is rated at 55 TOPS.

For large models running primarily on Radeon 8065S, memory bandwidth, precision and software efficiency can matter much more than an aggregate theoretical TOPS figure.

EVO-X5 Pro looks more like a small workstation than a traditional NUC

GMKtec pairs its large memory configuration with a vapor chamber and three active cooling fans.

It offers several performance modes and says the system has been validated for continuous 24/7 operation.

Storage can reach 24TB.

Connectivity includes dual front USB4 ports and dual rear USB4 v2 ports, with external GPU support as well.

AMD DASH remote management and a dedicated TPM also make the machine look increasingly like compact professional infrastructure rather than a living-room PC.

Multiple boxes can become a small distributed system

Vendors are also preparing to scale beyond a single chassis.

GMKtec supports USB4 v2 multi-system connectivity for distributed workloads.

Lenovo takes a similar approach with ThinkCentre X Ultra, allowing as many as four machines to be linked for larger models, longer contexts and multi-agent workloads.

That does not automatically combine all memory into one giant virtual GPU.

Software still has to partition a model or distribute independent inference requests across the individual nodes.

ACEMAGIC fits the same memory concept into roughly two liters

ACEMAGIC's F9A PRO 495 confirms that the 192GB configuration is not exclusive to one vendor.

It also supports LPDDR5X-8533 and up to 160GB of GPU allocation.

The enclosure is roughly two liters.

Dual PCIe 4.0 NVMe slots, USB4 and OCuLink provide storage and external expansion.

The physical dimensions still say mini PC.

The memory specification no longer does.

Soldered 192GB memory has a major upgradeability tradeoff

High-speed LPDDR5X is integrated into the platform rather than installed as ordinary replaceable desktop DIMMs.

That helps bandwidth and efficiency.

It also means buyers cannot assume a cheaper 64GB model can be upgraded to 192GB later by replacing memory modules.

Memory capacity becomes a configuration decision that may last for the useful lifetime of the machine.

Huge memory does not solve cooling

Running a large model continuously is not a bursty office workload.

CPU and GPU components can remain heavily utilized for long periods.

Inside a compact enclosure, sustained thermals eventually determine real clock speeds.

That is why a mini PC like EVO-X5 Pro ends up with a vapor chamber and three fans.

Inference throughput after half an hour will matter far more than a maximum boost clock printed on a specifications page.

Price may be the largest practical obstacle

Final prices for all of the most interesting 192GB configurations have not yet been published.

That amount of high-speed LPDDR5X is expensive, and the rest of these systems is hardly entry-level hardware.

Lenovo's ThinkCentre X Ultra, which currently tops out at 128GB in announced configurations, is already expected to start at roughly $3,700 in November.

It would therefore be unrealistic to expect the first 192GB mini PCs to compete with $1,000 mainstream desktops.

They are attacking a different cost equation

A conventional AI workstation may need several discrete GPUs simply to obtain enough accelerator memory for a very large model.

That quickly increases hardware cost, power consumption, cooling and chassis size.

A few-liter machine with 160GB accessible to the GPU may not win on raw throughput.

It may run a workload that would otherwise require a far more expensive multi-GPU system.

For developers, independent researchers and organizations working with sensitive local data, that is a fundamentally different proposition.

The mini PC is no longer trying only to replace a tower

Traditional mini PCs were designed to provide enough office, media or light workstation performance in the smallest possible space.

These Ryzen AI Max+ PRO 495 systems are starting to look like personal inference servers.

They can stay powered on, keep models local and process documents without necessarily sending them to a cloud provider.

The specification enabling that shift is not 55 TOPS or a 5.2GHz boost clock.

It is 192GB.

For once, the most consequential upgrade in a new PC category looks less like a processor race and more like a victory for memory capacity.