Whether every in-bounds element of a matrix operand is reachable through
the array<vec4<f32>> view that vec4ViewBinding produces for it.
The view truncates the binding down to a multiple of 16 bytes, so a
tightly-uploaded array can hold valid matrix elements past the view's end
even while its stride is 4-aligned (e.g. a column-major m×1 operand with
padded lda uploaded without padding cells). Call this with the kernel-side
dimensions and only take a shader's vectorized path when it returns true;
the scalar fallback reads the full storage and is always correct.
Whether every in-bounds element of a matrix operand is reachable through the
array<vec4<f32>>view that vec4ViewBinding produces for it. The view truncates the binding down to a multiple of 16 bytes, so a tightly-uploaded array can hold valid matrix elements past the view's end even while its stride is 4-aligned (e.g. a column-major m×1 operand with padded lda uploaded without padding cells). Call this with the kernel-side dimensions and only take a shader's vectorized path when it returns true; the scalar fallback reads the full storage and is always correct.