blog

Part 1 of 5: Kernel

For running LLM inference on Intel ARC GPUs with up-to-date engines like vLLM, you pretty much have to use the latest kernel. Intel XE kernel drivers and userspace libraries have been evolving quickly over the past year.

Only within the past few weeks have there been enough fixes for me to finally run multi-GPU inference on vLLM.

Update your kernel to at least 7.1.x

For Debian this means upgrading to forky (testing) as of this writing.

← all posts