wgblas
    Preparing search index...

    Function strmm

    • Performs the triangular matrix-matrix operation $$B \leftarrow \alpha \mathrm{op}(A) B \quad (\texttt{side='left'})$$ $$B \leftarrow \alpha B \mathrm{op}(A) \quad (\texttt{side='right'})$$ A is triangular, only its uplo triangle stored; B is a general m×n matrix, overwritten in place.

      • side='left': A is m×m — A premultiplies B
      • side='right': A is n×n — A postmultiplies B

      No dedicated fused kernel — a triangularize pass materializes a dense copy of op(A) (zero-filling the unstored triangle, substituting the implicit diagonal when diag='unit'), then a plain sgemm pass (sgemm_small.wgsl/sgemm_large.wgsl, unmodified) does the actual multiply, both on one command encoder.

      import { init, cleanup } from "wgblas";
      import { strmm } from "wgblas/strmm";

      const device = await init();

      // B := alpha*op(A)*B with A lower triangular (entries above the diagonal are
      // ignored). Taking B = identity makes the result the dense triangle of A.
      const n = 3,
      lda = n;
      const A = new Float32Array([2, 0, 0, 3, 4, 0, 5, 6, 8]);
      const B = new Float32Array([1, 0, 0, 0, 1, 0, 0, 0, 1]);

      console.log("A (lower triangular) =");
      console.table([A.slice(0, 3), A.slice(3, 6), A.slice(6, 9)]);

      const { B: result } = await strmm(
      device,
      "left",
      "lower",
      "no-transpose",
      "non-unit",
      n,
      n,
      1,
      A,
      lda,
      B,
      lda,
      );
      console.log("B = A*I = A =");
      console.table([result.slice(0, 3), result.slice(3, 6), result.slice(6, 9)]);

      if (typeof process !== "undefined") cleanup();

      Browser (standalone HTML):

      <!doctype html>
      <html lang="en">
      <head>
      <meta charset="UTF-8" />
      <title>strmm — wgblas browser example</title>
      <script src="https://unpkg.com/wgblas/dist/wgblas.browser.js"></script>
      </head>
      <body>
      <pre id="out">Running…</pre>
      <script>
      const { init, strmm, cleanup } = window.wgblas;

      (async () => {
      const device = await init();

      // B := A*B in place, A lower triangular. Starting from the identity,
      // B ends up holding the dense triangle of A.
      const n = 3, lda = n;
      const A = new Float32Array([2, 0, 0,
      3, 4, 0,
      5, 6, 8]);
      const B = new Float32Array([1, 0, 0,
      0, 1, 0,
      0, 0, 1]);

      const { B: result } = await strmm(device, "left", "lower", "no-transpose", "non-unit", n, n, 1, A, lda, B, lda);

      document.getElementById("out").textContent = [
      "A (lower triangular) =",
      " [" + A.slice(0, 3).join(", ") + "]",
      " [" + A.slice(3, 6).join(", ") + "]",
      " [" + A.slice(6, 9).join(", ") + "]",
      "B = A*I = A =",
      " [" + result.slice(0, 3).join(", ") + "]",
      " [" + result.slice(3, 6).join(", ") + "]",
      " [" + result.slice(6, 9).join(", ") + "]",
      ].join("\n");

      cleanup();
      })();
      </script>
      </body>
      </html>

      Parameters

      • device: GPUDevice

        GPUDevice from init()

      • side: "right" | "left"

        'left' for op(A)B, 'right' for Bop(A)

      • uplo: "lower" | "upper"

        'lower' if only A's lower triangle is stored, 'upper' for upper

      • transA: "no-transpose" | "transpose"

        'no-transpose' for A, 'transpose' for A^T

      • diag: "unit" | "non-unit"

        'unit' to treat A's diagonal as all-ones (A's diagonal is not read), 'non-unit' to read it

      • m: number

        rows of B

      • n: number

        columns of B

      • alpha: number

        scalar multiplier for the matrix product

      • A: Float32Array

        Float32Array, triangular, row-major or column-major (see layout)

      • lda: number

        leading dimension of A as stored

      • B: Float32Array

        Float32Array input/output matrix, overwritten with the result, row-major or column-major

      • ldb: number

        leading dimension of B as stored

      • Optionallayout: "column-major" | "row-major"

        storage layout shared by A/B when they're Float32Array (default: 'row-major') — column-major A is a genuine transpose (A isn't symmetric like ssymm's), so both transA and uplo are adjusted internally to compensate

      Returns Promise<{ B: Float32Array; gpuTimeMs?: number }>

      updated B as a Float32Array

    • Performs the triangular matrix-matrix operation $$B \leftarrow \alpha \mathrm{op}(A) B \quad (\texttt{side='left'})$$ $$B \leftarrow \alpha B \mathrm{op}(A) \quad (\texttt{side='right'})$$

      A and B are both kept GPU-resident; B is mutated in place. Each matrix's own layout (set at GpuMatrix.from time) determines the operation — there is no separate layout argument here. A and B must both be GpuMatrix or both be Float32Array — mixing is not supported.

      import { init, cleanup } from "wgblas";
      import { strmm } from "wgblas/strmm";
      import { GpuMatrix } from "wgblas/classes/GpuMatrix";

      const device = await init();

      // B := A*B in place, A lower triangular. Starting from the identity, B ends
      // up holding the dense triangle of A.
      const n = 3;
      const A = new Float32Array([2, 0, 0, 3, 4, 0, 5, 6, 8]);
      const B = new Float32Array([1, 0, 0, 0, 1, 0, 0, 0, 1]);

      const AGpu = GpuMatrix.from(A, n, n, n, "row-major");
      const BGpu = GpuMatrix.from(B, n, n, n, "row-major");

      console.log("A (lower triangular) =");
      console.table([A.slice(0, 3), A.slice(3, 6), A.slice(6, 9)]);

      await strmm(
      device,
      "left",
      "lower",
      "no-transpose",
      "non-unit",
      n,
      n,
      1,
      AGpu,
      AGpu.lda,
      BGpu,
      BGpu.lda,
      );

      const result = await BGpu.read();
      console.log("B = A*I = A =");
      console.table([result.slice(0, 3), result.slice(3, 6), result.slice(6, 9)]);

      AGpu.destroy();
      BGpu.destroy();
      if (typeof process !== "undefined") cleanup();

      Parameters

      • device: GPUDevice

        GPUDevice from init()

      • side: "right" | "left"

        'left' for op(A)B, 'right' for Bop(A)

      • uplo: "lower" | "upper"

        'lower' if only A's lower triangle is stored, 'upper' for upper

      • transA: "no-transpose" | "transpose"

        'no-transpose' for A, 'transpose' for A^T

      • diag: "unit" | "non-unit"

        'unit' to treat A's diagonal as all-ones (A's diagonal is not read), 'non-unit' to read it

      • m: number

        rows of B

      • n: number

        columns of B

      • alpha: number

        scalar multiplier for the matrix product

      • A: GpuMatrix

        GpuMatrix, triangular

      • lda: number

        leading dimension of A (must equal A.lda)

      • B: GpuMatrix

        GpuMatrix (mutated in place)

      • ldb: number

        leading dimension of B (must equal B.lda)

      Returns Promise<{ gpuTimeMs?: number }>

      no B — it stays GPU-resident; call B.read() yourself for a CPU readback (see the example)