wgblas
    Preparing search index...

    Function ssymm

    • Performs the symmetric matrix-matrix operation $$C \leftarrow \alpha A B + \beta C \quad (\texttt{side='left'})$$ $$C \leftarrow \alpha B A + \beta C \quad (\texttt{side='right'})$$ A is symmetric, only its uplo triangle stored; B and C are general m×n matrices.

      • side='left': A is m×m — A premultiplies B
      • side='right': A is n×n — A postmultiplies B

      No dedicated fused kernel — a symmetrize pass materializes a dense copy of A (mirroring the unstored triangle), then a plain sgemm pass (sgemm_small.wgsl/sgemm_large.wgsl, unmodified) does the actual multiply, both on one command encoder.

      import { init, cleanup } from "wgblas";
      import { ssymm } from "wgblas/ssymm";

      const device = await init();

      // C = alpha*A*B + beta*C with A symmetric (only its upper triangle stored).
      // Taking B = identity makes C the full symmetric matrix A stands for, so the
      // mirrored entries below the diagonal become visible.
      const n = 3,
      lda = n;
      const A = new Float32Array([2, 1, 0, 0, 2, 1, 0, 0, 2]);
      const B = new Float32Array([1, 0, 0, 0, 1, 0, 0, 0, 1]);
      const C = new Float32Array(n * n);

      console.log("A (upper triangle stored) =");
      console.table([A.slice(0, 3), A.slice(3, 6), A.slice(6, 9)]);

      const { C: result } = await ssymm(
      device,
      "left",
      "upper",
      n,
      n,
      1,
      A,
      lda,
      B,
      lda,
      0,
      C,
      lda,
      );
      console.log("C = A*I = the full symmetric A =");
      console.table([result.slice(0, 3), result.slice(3, 6), result.slice(6, 9)]); // [[2,1,0],[1,2,1],[0,1,2]]

      if (typeof process !== "undefined") cleanup();

      Browser (standalone HTML):

      <!doctype html>
      <html lang="en">
      <head>
      <meta charset="UTF-8" />
      <title>ssymm — wgblas browser example</title>
      <script src="https://unpkg.com/wgblas/dist/wgblas.browser.js"></script>
      </head>
      <body>
      <pre id="out">Running…</pre>
      <script>
      const { init, ssymm, cleanup } = window.wgblas;

      (async () => {
      const device = await init();

      // A is symmetric with only its upper triangle stored. Taking B = identity
      // makes C the full symmetric matrix A stands for.
      const n = 3, lda = n;
      const A = new Float32Array([2, 1, 0,
      0, 2, 1,
      0, 0, 2]);
      const B = new Float32Array([1, 0, 0,
      0, 1, 0,
      0, 0, 1]);
      const C = new Float32Array(n * n);

      const { C: result } = await ssymm(device, "left", "upper", n, n, 1, A, lda, B, lda, 0, C, lda);

      document.getElementById("out").textContent = [
      "A (upper triangle stored) =",
      " [" + A.slice(0, 3).join(", ") + "]",
      " [" + A.slice(3, 6).join(", ") + "]",
      " [" + A.slice(6, 9).join(", ") + "]",
      "C = A*I = the full symmetric A =",
      " [" + result.slice(0, 3).join(", ") + "]",
      " [" + result.slice(3, 6).join(", ") + "]",
      " [" + result.slice(6, 9).join(", ") + "]",
      ].join("\n");

      cleanup();
      })();
      </script>
      </body>
      </html>

      Parameters

      • device: GPUDevice

        GPUDevice from init()

      • side: "right" | "left"

        'left' for AB, 'right' for BA

      • uplo: "lower" | "upper"

        'lower' if only A's lower triangle is stored, 'upper' for upper

      • m: number

        rows of B and C

      • n: number

        columns of B and C

      • alpha: number

        scalar multiplier for the matrix product

      • A: Float32Array

        Float32Array, symmetric, row-major or column-major (see layout)

      • lda: number

        leading dimension of A as stored

      • B: Float32Array

        Float32Array, row-major or column-major (see layout)

      • ldb: number

        leading dimension of B as stored

      • beta: number

        scalar multiplier for C

      • C: Float32Array

        Float32Array input/output matrix, row-major or column-major

      • ldc: number

        leading dimension of C as stored

      • Optionallayout: "column-major" | "row-major"

        storage layout shared by A/B/C when they're Float32Array (default: 'row-major') — column-major A keeps representing the same symmetric matrix but flips which physical triangle looks stored, so uplo is adjusted internally to compensate

      Returns Promise<{ C: Float32Array; gpuTimeMs?: number }>

      updated C as a Float32Array

    • Performs the symmetric matrix-matrix operation $$C \leftarrow \alpha A B + \beta C \quad (\texttt{side='left'})$$ $$C \leftarrow \alpha B A + \beta C \quad (\texttt{side='right'})$$

      A, B, and C are all kept GPU-resident. Each matrix's own layout (set at GpuMatrix.from time) determines the operation — there is no separate layout argument here. A and B must be GpuMatrix whenever C is, and vice versa — mixing a GpuMatrix with a plain Float32Array is not supported.

      import { init, cleanup } from "wgblas";
      import { ssymm } from "wgblas/ssymm";
      import { GpuMatrix } from "wgblas/classes/GpuMatrix";

      const device = await init();

      // A is symmetric with only its upper triangle stored. Taking B = identity
      // makes C the full symmetric matrix A stands for.
      const n = 3;
      const A = new Float32Array([2, 1, 0, 0, 2, 1, 0, 0, 2]);
      const B = new Float32Array([1, 0, 0, 0, 1, 0, 0, 0, 1]);

      const AGpu = GpuMatrix.from(A, n, n, n, "row-major");
      const BGpu = GpuMatrix.from(B, n, n, n, "row-major");
      const CGpu = GpuMatrix.from(new Float32Array(n * n), n, n, n, "row-major");

      console.log("A (upper triangle stored) =");
      console.table([A.slice(0, 3), A.slice(3, 6), A.slice(6, 9)]);

      await ssymm(
      device,
      "left",
      "upper",
      n,
      n,
      1,
      AGpu,
      AGpu.lda,
      BGpu,
      BGpu.lda,
      0,
      CGpu,
      CGpu.lda,
      );

      const result = await CGpu.read();
      console.log("C = A*I = the full symmetric A =");
      console.table([result.slice(0, 3), result.slice(3, 6), result.slice(6, 9)]); // [[2,1,0],[1,2,1],[0,1,2]]

      AGpu.destroy();
      BGpu.destroy();
      CGpu.destroy();
      if (typeof process !== "undefined") cleanup();

      Parameters

      • device: GPUDevice

        GPUDevice from init()

      • side: "right" | "left"

        'left' for AB, 'right' for BA

      • uplo: "lower" | "upper"

        'lower' if only A's lower triangle is stored, 'upper' for upper

      • m: number

        rows of B and C

      • n: number

        columns of B and C

      • alpha: number

        scalar multiplier for the matrix product

      • A: GpuMatrix

        GpuMatrix, symmetric

      • lda: number

        leading dimension of A (must equal A.lda)

      • B: GpuMatrix

        GpuMatrix

      • ldb: number

        leading dimension of B (must equal B.lda)

      • beta: number

        scalar multiplier for C

      • C: GpuMatrix

        GpuMatrix (mutated in place)

      • ldc: number

        leading dimension of C (must equal C.lda)

      Returns Promise<{ gpuTimeMs?: number }>

      no C — it stays GPU-resident; call C.read() yourself for a CPU readback (see the example)