GPUDevice from init()
'left' for op(A)B, 'right' for Bop(A)
'lower' if only A's lower triangle is stored, 'upper' for upper
'no-transpose' for A, 'transpose' for A^T
'unit' to treat A's diagonal as all-ones (A's diagonal is not read), 'non-unit' to read it
rows of B
columns of B
scalar multiplier for the matrix product
Float32Array, triangular, row-major or column-major (see layout)
leading dimension of A as stored
Float32Array input/output matrix, overwritten with the result, row-major or column-major
leading dimension of B as stored
Optionallayout: "column-major" | "row-major"
storage layout shared by A/B when they're Float32Array
(default: 'row-major') — column-major A is a genuine transpose (A isn't
symmetric like ssymm's), so both transA and uplo are adjusted
internally to compensate
updated B as a Float32Array
Performs the triangular matrix-matrix operation $$B \leftarrow \alpha \mathrm{op}(A) B \quad (\texttt{side='left'})$$ $$B \leftarrow \alpha B \mathrm{op}(A) \quad (\texttt{side='right'})$$
A and B are both kept GPU-resident; B is mutated in place. Each matrix's
own layout (set at GpuMatrix.from time) determines the operation —
there is no separate layout argument here. A and B must both be
GpuMatrix or both be Float32Array — mixing is not supported.
import { init, cleanup } from "wgblas";
import { strmm } from "wgblas/strmm";
import { GpuMatrix } from "wgblas/classes/GpuMatrix";
const device = await init();
// B := A*B in place, A lower triangular. Starting from the identity, B ends
// up holding the dense triangle of A.
const n = 3;
const A = new Float32Array([2, 0, 0, 3, 4, 0, 5, 6, 8]);
const B = new Float32Array([1, 0, 0, 0, 1, 0, 0, 0, 1]);
const AGpu = GpuMatrix.from(A, n, n, n, "row-major");
const BGpu = GpuMatrix.from(B, n, n, n, "row-major");
console.log("A (lower triangular) =");
console.table([A.slice(0, 3), A.slice(3, 6), A.slice(6, 9)]);
await strmm(
device,
"left",
"lower",
"no-transpose",
"non-unit",
n,
n,
1,
AGpu,
AGpu.lda,
BGpu,
BGpu.lda,
);
const result = await BGpu.read();
console.log("B = A*I = A =");
console.table([result.slice(0, 3), result.slice(3, 6), result.slice(6, 9)]);
AGpu.destroy();
BGpu.destroy();
if (typeof process !== "undefined") cleanup();
GPUDevice from init()
'left' for op(A)B, 'right' for Bop(A)
'lower' if only A's lower triangle is stored, 'upper' for upper
'no-transpose' for A, 'transpose' for A^T
'unit' to treat A's diagonal as all-ones (A's diagonal is not read), 'non-unit' to read it
rows of B
columns of B
scalar multiplier for the matrix product
GpuMatrix, triangular
leading dimension of A (must equal A.lda)
GpuMatrix (mutated in place)
leading dimension of B (must equal B.lda)
no B — it stays GPU-resident; call B.read() yourself for a CPU readback (see the example)
Performs the triangular matrix-matrix operation $$B \leftarrow \alpha \mathrm{op}(A) B \quad (\texttt{side='left'})$$ $$B \leftarrow \alpha B \mathrm{op}(A) \quad (\texttt{side='right'})$$
Ais triangular, only itsuplotriangle stored;Bis a general m×n matrix, overwritten in place.side='left':Ais m×m —ApremultipliesBside='right':Ais n×n —ApostmultipliesBNo dedicated fused kernel — a
triangularizepass materializes a dense copy ofop(A)(zero-filling the unstored triangle, substituting the implicit diagonal whendiag='unit'), then a plainsgemmpass (sgemm_small.wgsl/sgemm_large.wgsl, unmodified) does the actual multiply, both on one command encoder.Browser (standalone HTML):