GPUDevice from init()
'left' to solve op(A)X=alphaB, 'right' to solve Xop(A)=alphaB
'lower' if only A's lower triangle is stored, 'upper' for upper
'no-transpose' for A, 'transpose' for A^T
'unit' to treat A's diagonal as all-ones (A's diagonal is not read), 'non-unit' to read it
rows of B
columns of B
scalar multiplier for B
Float32Array, triangular, row-major or column-major (see layout)
leading dimension of A as stored
Float32Array input/output matrix, overwritten with the solution, row-major or column-major
leading dimension of B as stored
Optionallayout: "column-major" | "row-major"
storage layout shared by A/B when they're Float32Array
(default: 'row-major') — column-major A is a genuine transpose (A isn't
symmetric like ssymm's), so both transA and uplo are adjusted
internally to compensate
the solution X, written into B, as a Float32Array
Solves the triangular matrix equation
$$\mathrm{op}(A) X = \alpha B \quad (\texttt{side='left'})$$
$$X \mathrm{op}(A) = \alpha B \quad (\texttt{side='right'})$$
overwriting B in place with X.
A and B are both kept GPU-resident. Each matrix's own layout (set at
GpuMatrix.from time) determines the operation — there is no separate
layout argument here. A and B must both be GpuMatrix or both be
Float32Array — mixing is not supported.
import { init, cleanup } from "wgblas";
import { strsm } from "wgblas/strsm";
import { GpuMatrix } from "wgblas/classes/GpuMatrix";
const device = await init();
// Solves A*X = B in place on the GPU. Passing B = A recovers the identity,
// undoing what the strmm example does.
const n = 3;
const A = new Float32Array([2, 0, 0, 3, 4, 0, 5, 6, 8]);
const AGpu = GpuMatrix.from(A, n, n, n, "row-major");
const BGpu = GpuMatrix.from(A.slice(), n, n, n, "row-major");
console.log("A (lower triangular) = B =");
console.table([A.slice(0, 3), A.slice(3, 6), A.slice(6, 9)]);
await strsm(
device,
"left",
"lower",
"no-transpose",
"non-unit",
n,
n,
1,
AGpu,
AGpu.lda,
BGpu,
BGpu.lda,
);
const result = await BGpu.read();
console.log("X solving A*X = B, i.e. the identity =");
console.table([result.slice(0, 3), result.slice(3, 6), result.slice(6, 9)]);
AGpu.destroy();
BGpu.destroy();
if (typeof process !== "undefined") cleanup();
GPUDevice from init()
'left' to solve op(A)X=alphaB, 'right' to solve Xop(A)=alphaB
'lower' if only A's lower triangle is stored, 'upper' for upper
'no-transpose' for A, 'transpose' for A^T
'unit' to treat A's diagonal as all-ones (A's diagonal is not read), 'non-unit' to read it
rows of B
columns of B
scalar multiplier for B
GpuMatrix, triangular
leading dimension of A (must equal A.lda)
GpuMatrix (overwritten in place with the solution)
leading dimension of B (must equal B.lda)
no B — it stays GPU-resident; call B.read() yourself for a CPU readback (see the example)
Solves the triangular matrix equation $$\mathrm{op}(A) X = \alpha B \quad (\texttt{side='left'})$$ $$X \mathrm{op}(A) = \alpha B \quad (\texttt{side='right'})$$ overwriting
Bwith the solutionX—Ais triangular, only itsuplotriangle stored;Bis a general m×n matrix.side='left':Ais m×m — solvesop(A)*X = alpha*Bside='right':Ais n×n — solvesX*op(A) = alpha*BBlocked substitution (strsv's own technique, generalized to a matrix RHS): strsv_invert_block + sgemm, both unmodified. A near-zero diagonal entry amplifies error, same as any triangular solve.
Browser (standalone HTML):