Device-initiated I/O: I/O Size Scaling#
In device-initiated I/O, GPU threads are the resource that drives commands, and
those threads compete with compute workloads for GPU resources. The PCIe
bandwidth saturation experiment characterized the transition from IOPS-bound to
bandwidth-bound regimes and the protocol overhead ratio for CPU-initiated P2P.
This experiment examines the same questions under device-initiated I/O by sweeping
I/O size across a wide range, from 512 bytes up to 64 KiB, with queue depth
as the secondary variable. The number of queues, nqueues, is fixed at 1. The
aim is to identify the minimum thread count required to saturate the PCIe link
at each I/O size.
xnvmeperf drives device-initiated I/O via the cuda-run subcommand with
the upcie-cuda backend. With nqueues=1, each CUDA thread handles one
in-flight command, so the total thread count equals queue depth × number of
devices. As I/O size grows, each command transfers more payload bytes, so fewer
in-flight commands are needed to fill the PCIe link. The minimum saturating
queue depth is therefore expected to decrease as I/O size increases, revealing
the thread count required at each I/O size.
Independent Variables#
Variable |
Parameter Set |
|---|---|
I/O size |
{ 512, 1024, 2048, 4096, 8192, 16384, 32768, 65536 } |
Queue depth |
{ 1, 2, 4, 8, 16, 32, 64, 128 } |
Number of queues |
1 |
Number of devices |
16 |
Total CUDA threads |
queue depth × 16 = { 16, 32, 64, 128, 256, 512, 1024, 2048 } |
Tool and backend |
xnvmeperf (cuda-run) + upcie-cuda |
Metrics Collected#
Metric |
Reported by |
|---|---|
Payload bandwidth (GB/s) |
xnvmeperf |
SM activity (fraction of SMs occupied) |
DCGM field 1002 (SM_ACTIVE) |
Warp slot occupancy |
DCGM field 1003 (SM_OCCUPANCY) |
GPU memory bandwidth utilization |
DCGM field 1005 (DRAM_ACTIVE) |
SM clock |
DCGM field 100 (SM_CLOCK) |
Run validity guards |
DCGM fields 101/112/202/237/238 |
Host-to-device PCIe bandwidth (GB/s) |
|
The DCGM fields are collected as described in GPU Telemetry.
Environment#
The benchmarks were run on the High-Performance Compute Server. NVMe devices are bound
to user space drivers. The CPU governor is set to performance with turbo
boost and SMT enabled. Each configuration is run five times and results are
reported as arithmetic means.
The namespaces are formatted before the run, as described in Device Fill State.
Execution of the Experiment#
Instructions for running bench_cuda_iosize.yaml are provided in
Experimental Framework.
Results#
Results are presented as payload bandwidth vs. I/O size, with one line per
queue depth (qdepth ∈ { 1, 2, 4, 8, 16, 32, 64, 128 }). All configurations
use xnvmeperf with the cuda-run subcommand and the upcie-cuda
backend, nqueues=1, and 16 NVMe devices. Total CUDA thread count equals
queue depth × 16. The dashed reference line marks the host-to-device PCIe
bandwidth from nvbandwidth.
Payload bandwidth vs. I/O size for xnvmeperf (cuda-run), 16 NVMe devices, nqueues=1. The minimum queue depth to saturate the ~45 GB/s practical ceiling drops from >128 at 512 B to 2 at 64 KiB.#
The queue depth required to saturate the PCIe link decreases monotonically as
I/O size grows. At 512-byte I/O, even the maximum configuration of 2048 CUDA
threads (qdepth=128, 16 devices) delivers only 21.6 GB/s. The workload is
constrained by per-device IOPS limits, not by available PCIe bandwidth. As I/O
size increases, each command carries more payload, and the saturation threshold
drops accordingly.
All lines converge at a practical ceiling of approximately 44–45 GB/s, which
falls roughly 84% of the way to the 53.7 GB/s nvbandwidth reference. This
gap is consistent with the overhead of NVMe command processing and PCIe
protocol framing on top of raw DMA throughput, as characterized in
Results.
The saturation queue depths are:
I/O size |
Min. |
Total CUDA threads |
|---|---|---|
512 B |
> 128 (not reached) |
> 2048 |
1024 B |
> 128 (not reached) |
> 2048 |
2048 B |
64 |
1024 |
4096 B |
32 |
512 |
8192 B |
16 |
256 |
16384 B |
8 |
128 |
32768 B |
4 |
64 |
65536 B |
2 |
32 |
At 64 KiB, a single queue depth of 2, 32 CUDA threads across 16 devices, is
sufficient to sustain ~45 GB/s of storage bandwidth with no CPU involvement
in the command path. qdepth=1 still achieves 41.9 GB/s at this I/O size,
demonstrating that device-initiated I/O can approach the practical link ceiling
with minimal thread-count overhead.
GPU Compute Cost of the Polling Kernel#
GPU engine activity vs. I/O size for xnvmeperf (cuda-run), 16 NVMe devices, nqueues=1, queue depth fixed at 128 (qdepth 128 × 16 devices = 2048 CUDA threads in total). SM active and SM occupancy are charted with each point carrying its reading; the SM clock (field 100) is on the right axis.#
The activity plot holds the queue depth at the largest configuration of the sweep, so what it charts across I/O sizes is the GPU-side cost of driving the I/O. Unlike the CPU-initiated P2P path, the device-initiated path keeps a persistent kernel resident.
SM activity is flat at approximately 13.3% from 512 bytes to 64 KiB, and warp slot occupancy holds near 0.87%. Neither follows the work being done, since achieved bandwidth doubles across the sweep while the compute footprint moves by about two tenths of a percentage point. The cost follows the launch geometry instead. One resident block per queue per device is 16 blocks for the single queue per device this sweep uses, a footprint that stays flat however large the I/O. The queue-depth sweep confirms the reading by varying the queue count, which scales the footprint in proportion, as reported in Results.
DRAM activity stays at 0% up to 8 KiB and reaches only approximately 2.1% at the largest I/O sizes, so the data the SSDs read into GPU memory consumes a negligible share of HBM bandwidth and the constraint stays on the PCIe link rather than on GPU memory, as the CPU-initiated measurements in Results also found.
Saturating the link through device-initiated I/O therefore costs roughly an eighth of the GPU’s streaming multiprocessors and almost none of its memory bandwidth, leaving the remainder for the compute the data was fetched for. That eighth is the price of the single queue per device used throughout this sweep, and it is the queue count that moves it. The price of each further queue is reported in Results.
Run Validity#
Between 73% and 77% of each monitoring window qualified as transferring.
The guard fields agree that the runs are comparable. The SM and memory clocks hold at 1755 MHz and 1593 MHz with the throttle reason bits clear, the replay counter stays at zero, and the link fields report Gen5 x16 throughout.
Summary#
The minimum thread count required to saturate the PCIe link decreases monotonically with I/O size. At small I/O sizes the constraint is per-device IOPS rather than link capacity, and even 2048 CUDA threads across 16 devices cannot saturate the link at 512 or 1024 bytes. At 64 KiB, just 32 threads suffice. The practical bandwidth ceiling for device-initiated I/O is approximately 44–45 GB/s, consistent with the ~28% protocol overhead characterized in the preceding bandwidth experiment.