GPUDevice from init()
'no-transpose' for A, 'transpose' for A^T
'no-transpose' for B, 'transpose' for B^T
rows of op(A) and C
columns of op(B) and C
columns of op(A), rows of op(B)
scalar multiplier for op(A)*op(B)
Float32Array, row-major or column-major (see layout)
leading dimension of A as stored
Float32Array, row-major or column-major (see layout)
leading dimension of B as stored
scalar multiplier for C
Float32Array input/output matrix, row-major or column-major
leading dimension of C as stored
Optionallayout: "column-major" | "row-major"
storage layout shared by A/B/C when they're Float32Array
(default: 'row-major'); column-major A/B flips the respective trans
flag internally, column-major C computes C^T = op(B)^T*op(A)^T instead
(same underlying bytes) — op(A)*op(B) stays what you asked for either way
updated C as a Float32Array
Performs the matrix-matrix operation $$C \leftarrow \alpha \mathrm{op}(A) \mathrm{op}(B) + \beta C$$
A, B, and C are all kept GPU-resident. Each matrix's own layout (set at
GpuMatrix.from time) determines the operation — there is no separate
layout argument here. A and B must be GpuMatrix whenever C is, and vice
versa — mixing a GpuMatrix with a plain Float32Array is not supported.
import { init, cleanup } from "wgblas";
import { sgemm } from "wgblas/sgemm";
import { GpuMatrix } from "wgblas/classes/GpuMatrix";
const device = await init();
// C = A*B with A 2x3 and B 3x2, so C is 2x2.
const m = 2,
n = 2,
k = 3;
const A = new Float32Array([1, 2, 3, 4, 5, 6]);
const B = new Float32Array([1, 0, 0, 1, 1, 1]);
const AGpu = GpuMatrix.from(A, m, k, k, "row-major");
const BGpu = GpuMatrix.from(B, k, n, n, "row-major");
const CGpu = GpuMatrix.from(new Float32Array(m * n), m, n, n, "row-major");
console.log("A =");
console.table([A.slice(0, 3), A.slice(3, 6)]);
await sgemm(
device,
"no-transpose",
"no-transpose",
m,
n,
k,
1,
AGpu,
AGpu.lda,
BGpu,
BGpu.lda,
0,
CGpu,
CGpu.lda,
);
const result = await CGpu.read();
console.log("C = A*B =");
console.table([result.slice(0, 2), result.slice(2, 4)]); // [[4,5],[10,11]]
AGpu.destroy();
BGpu.destroy();
CGpu.destroy();
if (typeof process !== "undefined") cleanup();
GPUDevice from init()
'no-transpose' for A, 'transpose' for A^T
'no-transpose' for B, 'transpose' for B^T
rows of op(A) and C
columns of op(B) and C
columns of op(A), rows of op(B)
scalar multiplier for op(A)*op(B)
GpuMatrix
leading dimension of A (must equal A.lda)
GpuMatrix
leading dimension of B (must equal B.lda)
scalar multiplier for C
GpuMatrix (mutated in place)
leading dimension of C (must equal C.lda)
no C — it stays GPU-resident; call C.read() yourself for a CPU readback (see the example)
Performs the matrix-matrix operation $$C \leftarrow \alpha \mathrm{op}(A) \mathrm{op}(B) + \beta C$$
transA/transB='no-transpose': op(A) = A (m×k), op(B) = B (k×n)transA/transB='transpose': op(A) = A^T, op(B) = B^TA, B, C are row-major or column-major (see
layout) — backed by one of two shared-memory-tiled, register-blocked kernels chosen by shape:sgemm_small.wgsl(BM=BN=32) below a 6x6 workgroup grid,sgemm_large.wgsl(BM=BN=64) above it.Browser (standalone HTML):