GPUDevice from init()
'lower' to update only col <= row, 'upper' for col >= row
'no-transpose' for AB^T + BA^T, 'transpose' for A^TB + B^TA
order of C (C is n×n); rows of op(A) and op(B)
columns of op(A) and op(B) — the shared/contracted dimension
scalar multiplier for both rank-k terms
Float32Array, row-major or column-major (see layout)
leading dimension of A as stored
Float32Array, row-major or column-major (see layout)
leading dimension of B as stored
scalar multiplier for C
Float32Array input/output matrix, row-major or column-major
leading dimension of C as stored
Optionallayout: "column-major" | "row-major"
storage layout shared by A/B/C when they're Float32Array
(default: 'row-major') — same handling as sgemmtr; column-major C
flips uplo internally so it still names the triangle you asked for
updated C as a Float32Array (only the requested triangle changed)
Performs the symmetric rank-2k update $$C \leftarrow \mathrm{uplo}(\alpha \mathrm{op}(A) \mathrm{op}(B)^{T} + \alpha \mathrm{op}(B) \mathrm{op}(A)^{T} + \beta C)$$
A, B, and C are all kept GPU-resident. Each matrix's own layout (set at
GpuMatrix.from time) determines the operation — there is no separate
layout argument here. A and B must be GpuMatrix whenever C is, and vice
versa — mixing a GpuMatrix with a plain Float32Array is not supported.
import { init, cleanup } from "wgblas";
import { ssyr2k } from "wgblas/ssyr2k";
import { GpuMatrix } from "wgblas/classes/GpuMatrix";
const device = await init();
// With B = A the two rank-k terms are equal, so the result is exactly twice
// ssyrk's on the same A.
const n = 3,
k = 2;
const A = new Float32Array([1, 0, 0, 1, 1, 1]);
const AGpu = GpuMatrix.from(A, n, k, k, "row-major");
const BGpu = GpuMatrix.from(A.slice(), n, k, k, "row-major");
const CGpu = GpuMatrix.from(new Float32Array(n * n), n, n, n, "row-major");
console.log("A = B =");
console.table([A.slice(0, 2), A.slice(2, 4), A.slice(4, 6)]);
await ssyr2k(
device,
"upper",
"no-transpose",
n,
k,
1,
AGpu,
AGpu.lda,
BGpu,
BGpu.lda,
0,
CGpu,
CGpu.lda,
);
const result = await CGpu.read();
console.log("C = upper(A*B^T + B*A^T) =");
console.table([result.slice(0, 3), result.slice(3, 6), result.slice(6, 9)]); // 2x ssyrk: [[2,0,2],[0,2,2],[0,0,4]]
AGpu.destroy();
BGpu.destroy();
CGpu.destroy();
if (typeof process !== "undefined") cleanup();
GPUDevice from init()
'lower' to update only col <= row, 'upper' for col >= row
'no-transpose' for AB^T + BA^T, 'transpose' for A^TB + B^TA
order of C (C is n×n); rows of op(A) and op(B)
columns of op(A) and op(B) — the shared/contracted dimension
scalar multiplier for both rank-k terms
GpuMatrix
leading dimension of A (must equal A.lda)
GpuMatrix
leading dimension of B (must equal B.lda)
scalar multiplier for C
GpuMatrix (mutated in place; only the requested triangle changes)
leading dimension of C (must equal C.lda)
no C — it stays GPU-resident; call C.read() yourself for a CPU readback (see the example)
Performs the symmetric rank-2k update $$C \leftarrow \mathrm{uplo}(\alpha \mathrm{op}(A) \mathrm{op}(B)^{T} + \alpha \mathrm{op}(B) \mathrm{op}(A)^{T} + \beta C)$$
only the triangle of C named by
uplois read or written ('lower':col <= row,'upper':col >= row). C is always n×n.trans='no-transpose': op(A) = A, op(B) = B (both n×k)trans='transpose': op(A) = A^T, op(B) = B^T (A, B stored k×n)No dedicated kernel — two passes of
sgemmtr's kernel (sgemmtr_small.wgsl/sgemmtr_large.wgsl) on one command encoder, the second accumulating withbeta=1so C never leaves the GPU between them.Browser (standalone HTML):