Unless noted otherwise, every result above uses unit stride (incx = incy = 1) — the normal case, and the coalesced, best-case GPU access pattern. Real usage sometimes passes a non-unit stride (e.g. operating on a row or column of a larger matrix, where incx = lda), which breaks memory coalescing and costs measurably more. This section sweeps a few representative strides to characterize that cost separately, collapsed below by default — expand a stride to see its table and chart.
Benchmark results for isamax on Intel R Iris R Xe Graphics Tgl Gt2.
Intel R Iris R Xe Graphics Tgl Gt2
See also
Stride sweep
Unless noted otherwise, every result above uses unit stride (
incx = incy = 1) — the normal case, and the coalesced, best-case GPU access pattern. Real usage sometimes passes a non-unit stride (e.g. operating on a row or column of a larger matrix, whereincx = lda), which breaks memory coalescing and costs measurably more. This section sweeps a few representative strides to characterize that cost separately, collapsed below by default — expand a stride to see its table and chart.Intel R Iris R Xe Graphics Tgl Gt2 — stride = 4
Intel R Iris R Xe Graphics Tgl Gt2 — stride = 5
Intel R Iris R Xe Graphics Tgl Gt2 — stride = 32
Intel R Iris R Xe Graphics Tgl Gt2 — stride = 33
Intel R Iris R Xe Graphics Tgl Gt2 — stride = 255
Intel R Iris R Xe Graphics Tgl Gt2 — stride = 256
See also: