GPUDevice from init()
'lower' to update only col <= row, 'upper' for col >= row
'no-transpose' for A, 'transpose' for A^T
'no-transpose' for B, 'transpose' for B^T
rows of op(A) and C
columns of op(B) and C
columns of op(A), rows of op(B)
scalar multiplier for op(A)*op(B)
Float32Array, row-major or column-major (see layout)
leading dimension of A as stored
Float32Array, row-major or column-major (see layout)
leading dimension of B as stored
scalar multiplier for C
Float32Array input/output matrix, row-major or column-major
leading dimension of C as stored
Optionallayout: "column-major" | "row-major"
storage layout shared by A/B/C when they're Float32Array
(default: 'row-major') — same handling as sgemm, plus uplo is
flipped internally for column-major C so it still names the triangle
you asked for
updated C as a Float32Array (only the requested triangle changed)
Performs the matrix-matrix operation $$C \leftarrow \mathrm{uplo}(\alpha \mathrm{op}(A) \mathrm{op}(B) + \beta C)$$
A, B, and C are all kept GPU-resident. Each matrix's own layout (set at
GpuMatrix.from time) determines the operation — there is no separate
layout argument here. A and B must be GpuMatrix whenever C is, and vice
versa — mixing a GpuMatrix with a plain Float32Array is not supported.
import { init, cleanup } from "wgblas";
import { sgemmtr } from "wgblas/sgemmtr";
import { GpuMatrix } from "wgblas/classes/GpuMatrix";
const device = await init();
// A all ones times the identity is all ones, so the masking is obvious:
// only the upper triangle of C is written.
const n = 3;
const A = new Float32Array([1, 1, 1, 1, 1, 1, 1, 1, 1]);
const B = new Float32Array([1, 0, 0, 0, 1, 0, 0, 0, 1]);
const AGpu = GpuMatrix.from(A, n, n, n, "row-major");
const BGpu = GpuMatrix.from(B, n, n, n, "row-major");
const CGpu = GpuMatrix.from(new Float32Array(n * n), n, n, n, "row-major");
console.log("A = all ones, B = identity, so A*B is all ones");
await sgemmtr(
device,
"upper",
"no-transpose",
"no-transpose",
n,
n,
n,
1,
AGpu,
AGpu.lda,
BGpu,
BGpu.lda,
0,
CGpu,
CGpu.lda,
);
const result = await CGpu.read();
console.log("C = upper(A*B) =");
console.table([result.slice(0, 3), result.slice(3, 6), result.slice(6, 9)]); // [[1,1,1],[0,1,1],[0,0,1]]
AGpu.destroy();
BGpu.destroy();
CGpu.destroy();
if (typeof process !== "undefined") cleanup();
GPUDevice from init()
'lower' to update only col <= row, 'upper' for col >= row
'no-transpose' for A, 'transpose' for A^T
'no-transpose' for B, 'transpose' for B^T
rows of op(A) and C
columns of op(B) and C
columns of op(A), rows of op(B)
scalar multiplier for op(A)*op(B)
GpuMatrix
leading dimension of A (must equal A.lda)
GpuMatrix
leading dimension of B (must equal B.lda)
scalar multiplier for C
GpuMatrix (mutated in place; only the requested triangle changes)
leading dimension of C (must equal C.lda)
no C — it stays GPU-resident; call C.read() yourself for a CPU readback (see the example)
Performs the matrix-matrix operation $$C \leftarrow \mathrm{uplo}(\alpha \mathrm{op}(A) \mathrm{op}(B) + \beta C)$$
sgemm's operation, but only the triangle of C named byuplois read or written ('lower':col <= row,'upper':col >= row; C need not be square — the test applies over the full m×n grid).Same kernels as
sgemm(sgemmtr_small.wgsl/sgemmtr_large.wgsl, identical tiling), with the final output write masked to one triangle.Browser (standalone HTML):