Thanks to its future driver, RADEON Software 25.8.1, AMD will enhance the artificial intelligence processing capabilities of its Strix Halo processors. We’re talking about LLM (Large Language Model) management of 128B parameters, all locally. Similarly, the Reds improve the possible context size: up to 25,6000 tokens!
AMD makes it possible to run LLMs with 128B parameters locally!
With this new graphics driver, via the Variable Graphics Memory function, we can increase the amount of RAM available to the iGPU tenfold. In this context, up to 96GB of memory can be allocated to the iGPU to boost AI performance. It will therefore be possible to run models with 128 billion parameters locally. For example, Mistral Large 123B will be able to run directly from your PC, as will Llama 4 Scout 109B (17B active).
Another improvement to take into account is the support for wider contexts. Much wider, in fact, since 25,6000 tokens will be supported, as opposed to the usual 4,096. This will enable much longer conversations to be managed without “forgetting”. Likewise, it also allows the processing of much larger documents.
Of course, the underlying condition for taking advantage of all this is to be equipped with Strix Halo hardware. These are not yet widely available, and the first machines are very expensive. We’re thinking in particular of the Corsair AI Workstation 300 PCs, with list prices starting from $1,599.99.
If NVIDIA has clearly taken the lead in the AI sector thanks to its graphics cards, AMD could well win a battle in the CPU sector. From here, we await Intel’s response. For the moment, AMD seems to be hitting hard in this fast-growing sector.

