uPCIe#
xNVMe provides kernel-bypassing backends implemented using uPCIe, a minimal header-only user space NVMe driver. Unlike SPDK, uPCIe has no reactor, threading model, or application framework — just direct PCIe BAR access and DMA.
Two backend configs are available, differing in where I/O buffers are allocated:
uPCIe (host memory) — DMA buffers in host memory, backed by hugepages.
uPCIe CUDA — DMA buffers in GPU device memory, enabling PCIe peer-to-peer (P2P) transfers directly between the NVMe device and the GPU.
Additionally, the host-memory backend supports a uPCIe (multi-process) mode in which several processes share a single NVMe controller via a designated primary process.
Device Identifiers#
When using user space NVMe drivers, the operating-system kernel NVMe
driver is detached and the device is bound either to uio_pci_generic for the
legacy no-IOMMU path or to vfio-pci for the IOMMU-backed path. Thus, the
device files in /dev/, such as /dev/nvme0n1, are not available. Devices
are instead identified by their PCI id (0000:03:00.0), and namespace
identifier via the --nsid option.
System Configuration#
Driver Attachment#
Use the xnvme-driver script to unbind the kernel NVMe driver and bind
devices to uio_pci_generic or vfio-pci:
xnvme-driver
upcie, upcie-cuda and upcie-hip all support uio_pci_generic
(no-IOMMU path) and vfio-pci (IOMMU-backed path). For the GPU backends see
Enforcing IOMMU.
Attachment Modes#
The attachment mode follows the driver bound to the device, and decides how DMA memory is wired up:
uio_pci_genericUIO_LUT. BAR mapping plus a hugepage-backed lookup-table translator. The only mode that supports uPCIe (multi-process).
vfio-pciwith/dev/iommuVFIO_CDEV. One iommufd IOAS shared across controllers, memory imported from a hugepage-backed memfd.
vfio-pciwithout/dev/iommuVFIO_TYPE1. The legacy vfio container, one per runtime, with the hugepage mapped into it.
Between the two vfio-pci modes, the backend prefers VFIO_CDEV when
/dev/iommu is available and the device has a vfio-dev entry, falling
back to VFIO_TYPE1. Setting XNVME_UPCIE_VFIO_MODE to iommufd or
type1 forces the choice.
Privileges#
uPCIe requires root for two reasons:
Reading
/proc/self/pagemapto translate virtual to physical addresses for DMA.Accessing VFIO device nodes and writing PCI config space via sysfs to enable Bus Master when needed.
Dual-Backend Operation#
The same PCIe device can be opened simultaneously with both upcie and
upcie-cuda. This lets one handle send I/O through host-memory buffers while
another sends I/O through GPU device-memory buffers, both going to the same NVMe
controller. Opening order does not matter; the controller is torn down only when
the last handle across both backends is closed. Attempting to open the same URI
with an unrelated backend while either uPCIe backend holds it returns -EBUSY.
Before opening both handles, complete the system configuration for each backend: hugepages for uPCIe (host memory) and the additional CUDA and kernel requirements for uPCIe CUDA.
Example#
struct xnvme_opts host_opts = xnvme_opts_default();
struct xnvme_opts cuda_opts = xnvme_opts_default();
host_opts.be = "upcie";
cuda_opts.be = "upcie-cuda";
struct xnvme_dev *host_dev = xnvme_dev_open("0000:03:00.0", &host_opts);
struct xnvme_dev *cuda_dev = xnvme_dev_open("0000:03:00.0", &cuda_opts);