GPUDevice from init()
'no-transpose' for A, 'transpose' for A^T
number of rows in A
number of columns in A
scalar multiplier for op(A)*x
Float64Array, row-major or column-major (see layout), at least
(m-1)*lda+n elements for row-major or (n-1)*lda+m elements for column-major
leading dimension of A (>= n for row-major, >= m for column-major)
Float64Array input vector
stride for x (must be a positive integer)
scalar multiplier for y
Float64Array input/output vector
stride for y (must be a positive integer)
Optionallayout: "column-major" | "row-major"
storage layout of A (default: 'row-major'); column-major
swaps the effective m/n and flips trans internally (op(A) stays
what you asked for either way — x/y keep their original lengths)
Performs the matrix-vector operation $$y \leftarrow \alpha \mathrm{op}(A) x + \beta y$$ in double precision (double-double emulation).
x and y are kept resident on the GPU. A must be a GpuMatrix (Float64Array-
backed); its own layout (set at GpuMatrix.from time) determines the
operation — there is no separate layout argument here.
import { init, cleanup } from "wgblas";
import { dgemv } from "wgblas/dgemv";
import { GpuVector } from "wgblas/classes/GpuVector";
import { GpuMatrix } from "wgblas/classes/GpuMatrix";
const device = await init();
// A diagonal A makes the chained result easy to check by eye.
const n = 3;
const A = new Float64Array([1, 0, 0, 0, 2, 0, 0, 0, 3]);
const x = new Float64Array([1, 1, 1]);
const AGpu = GpuMatrix.from(A, n, n, n, "row-major");
const xGpu = GpuVector.from(x);
const yGpu = GpuVector.from(new Float64Array(n));
console.log("A = diag(1, 2, 3) =");
console.table([A.slice(0, 3), A.slice(3, 6)]);
console.log("x =", x);
// Results stay on the GPU between the two calls — no readback in between.
await dgemv(
device,
"no-transpose",
n,
n,
1,
AGpu,
AGpu.lda,
xGpu,
1,
0,
yGpu,
1,
); // y = A*x
await dgemv(
device,
"no-transpose",
n,
n,
1,
AGpu,
AGpu.lda,
yGpu,
1,
0,
xGpu,
1,
); // x = A*y
// Single readback at the end.
console.log("A*A*x =", await xGpu.read()); // [1*1, 2*2, 3*3] = [1, 4, 9]
AGpu.destroy();
xGpu.destroy();
yGpu.destroy();
if (typeof process !== "undefined") cleanup();
GPUDevice from init()
'no-transpose' for A, 'transpose' for A^T
number of rows in A
number of columns in A
scalar multiplier for op(A)*x
GpuMatrix (Float64Array-backed)
leading dimension of A (must equal A.lda)
GpuVector input vector (Float64Array-backed, not mutated)
stride for x (must be a positive integer)
scalar multiplier for y
GpuVector input/output vector (Float64Array-backed, mutated in place)
stride for y (must be a positive integer)
Performs the matrix-vector operation $$y \leftarrow \alpha \mathrm{op}(A) x + \beta y$$ in double precision (double-double emulation — WGSL has no native f64 type).
trans='no-transpose': op(A) = A, x is length n, y is length mtrans='transpose': op(A) = A^T, x is length m, y is length nA is an m×n matrix stored in row-major order.
ldais the leading dimension (number of doubles between the start of consecutive rows — must be >= n).Browser (standalone HTML):