GPUDevice from init()
'left' for AB, 'right' for BA
'lower' if only A's lower triangle is stored, 'upper' for upper
rows of B and C
columns of B and C
scalar multiplier for the matrix product
Float32Array, symmetric, row-major or column-major (see layout)
leading dimension of A as stored
Float32Array, row-major or column-major (see layout)
leading dimension of B as stored
scalar multiplier for C
Float32Array input/output matrix, row-major or column-major
leading dimension of C as stored
Optionallayout: "column-major" | "row-major"
storage layout shared by A/B/C when they're Float32Array
(default: 'row-major') — column-major A keeps representing the same
symmetric matrix but flips which physical triangle looks stored, so
uplo is adjusted internally to compensate
updated C as a Float32Array
Performs the symmetric matrix-matrix operation $$C \leftarrow \alpha A B + \beta C \quad (\texttt{side='left'})$$ $$C \leftarrow \alpha B A + \beta C \quad (\texttt{side='right'})$$
A, B, and C are all kept GPU-resident. Each matrix's own layout (set at
GpuMatrix.from time) determines the operation — there is no separate
layout argument here. A and B must be GpuMatrix whenever C is, and vice
versa — mixing a GpuMatrix with a plain Float32Array is not supported.
import { init, cleanup } from "wgblas";
import { ssymm } from "wgblas/ssymm";
import { GpuMatrix } from "wgblas/classes/GpuMatrix";
const device = await init();
// A is symmetric with only its upper triangle stored. Taking B = identity
// makes C the full symmetric matrix A stands for.
const n = 3;
const A = new Float32Array([2, 1, 0, 0, 2, 1, 0, 0, 2]);
const B = new Float32Array([1, 0, 0, 0, 1, 0, 0, 0, 1]);
const AGpu = GpuMatrix.from(A, n, n, n, "row-major");
const BGpu = GpuMatrix.from(B, n, n, n, "row-major");
const CGpu = GpuMatrix.from(new Float32Array(n * n), n, n, n, "row-major");
console.log("A (upper triangle stored) =");
console.table([A.slice(0, 3), A.slice(3, 6), A.slice(6, 9)]);
await ssymm(
device,
"left",
"upper",
n,
n,
1,
AGpu,
AGpu.lda,
BGpu,
BGpu.lda,
0,
CGpu,
CGpu.lda,
);
const result = await CGpu.read();
console.log("C = A*I = the full symmetric A =");
console.table([result.slice(0, 3), result.slice(3, 6), result.slice(6, 9)]); // [[2,1,0],[1,2,1],[0,1,2]]
AGpu.destroy();
BGpu.destroy();
CGpu.destroy();
if (typeof process !== "undefined") cleanup();
GPUDevice from init()
'left' for AB, 'right' for BA
'lower' if only A's lower triangle is stored, 'upper' for upper
rows of B and C
columns of B and C
scalar multiplier for the matrix product
GpuMatrix, symmetric
leading dimension of A (must equal A.lda)
GpuMatrix
leading dimension of B (must equal B.lda)
scalar multiplier for C
GpuMatrix (mutated in place)
leading dimension of C (must equal C.lda)
no C — it stays GPU-resident; call C.read() yourself for a CPU readback (see the example)
Performs the symmetric matrix-matrix operation $$C \leftarrow \alpha A B + \beta C \quad (\texttt{side='left'})$$ $$C \leftarrow \alpha B A + \beta C \quad (\texttt{side='right'})$$
Ais symmetric, only itsuplotriangle stored;BandCare general m×n matrices.side='left':Ais m×m —ApremultipliesBside='right':Ais n×n —ApostmultipliesBNo dedicated fused kernel — a
symmetrizepass materializes a dense copy ofA(mirroring the unstored triangle), then a plainsgemmpass (sgemm_small.wgsl/sgemm_large.wgsl, unmodified) does the actual multiply, both on one command encoder.Browser (standalone HTML):