wgblas
    Preparing search index...

    Function ssyrk

    • Performs the symmetric rank-k update $$C \leftarrow \mathrm{uplo}(\alpha \mathrm{op}(A) \mathrm{op}(A)^{T} + \beta C)$$

      only the triangle of C named by uplo is read or written ('lower': col <= row, 'upper': col >= row). C is always n×n.

      • trans='no-transpose': op(A) = A (n×k), computes alphaAA^T + beta*C
      • trans='transpose': op(A) = A^T (A stored k×n), computes alphaA^TA + beta*C

      No dedicated kernel — this is sgemmtr's exact kernel (sgemmtr_small.wgsl/sgemmtr_large.wgsl) with A aliased into both operand slots (B := A).

      import { init, cleanup } from "wgblas";
      import { ssyrk } from "wgblas/ssyrk";

      const device = await init();

      // C = uplo(alpha*A*A^T + beta*C). Entry (i,j) is the dot product of rows i
      // and j of A, and only the upper triangle is written.
      const n = 3,
      k = 2;
      const A = new Float32Array([1, 0, 0, 1, 1, 1]);
      const C = new Float32Array(n * n);

      console.log("A =");
      console.table([A.slice(0, 2), A.slice(2, 4), A.slice(4, 6)]);

      const { C: result } = await ssyrk(
      device,
      "upper",
      "no-transpose",
      n,
      k,
      1,
      A,
      k,
      0,
      C,
      n,
      );
      console.log("C = upper(A*A^T) =");
      console.table([result.slice(0, 3), result.slice(3, 6), result.slice(6, 9)]);
      // row dot products: [[1·1, 1·0, 1·1], [-, 1·1, 1·1], [-, -, 2]] = [[1,0,1],[0,1,1],[0,0,2]]

      if (typeof process !== "undefined") cleanup();

      Browser (standalone HTML):

      <!doctype html>
      <html lang="en">
      <head>
      <meta charset="UTF-8" />
      <title>ssyrk — wgblas browser example</title>
      <script src="https://unpkg.com/wgblas/dist/wgblas.browser.js"></script>
      </head>
      <body>
      <pre id="out">Running…</pre>
      <script>
      const { init, ssyrk, cleanup } = window.wgblas;

      (async () => {
      const device = await init();

      // Entry (i,j) of C is the dot product of rows i and j of A; only the
      // upper triangle is written.
      const n = 3, k = 2;
      const A = new Float32Array([1, 0,
      0, 1,
      1, 1]);
      const C = new Float32Array(n * n);

      const { C: result } = await ssyrk(device, "upper", "no-transpose", n, k, 1, A, k, 0, C, n);

      document.getElementById("out").textContent = [
      "A =",
      " [" + A.slice(0, 2).join(", ") + "]",
      " [" + A.slice(2, 4).join(", ") + "]",
      " [" + A.slice(4, 6).join(", ") + "]",
      "C = upper(A*A^T) =",
      " [" + result.slice(0, 3).join(", ") + "]",
      " [" + result.slice(3, 6).join(", ") + "]",
      " [" + result.slice(6, 9).join(", ") + "]",
      ].join("\n");

      cleanup();
      })();
      </script>
      </body>
      </html>

      Parameters

      • device: GPUDevice

        GPUDevice from init()

      • uplo: "lower" | "upper"

        'lower' to update only col <= row, 'upper' for col >= row

      • trans: "no-transpose" | "transpose"

        'no-transpose' for AA^T, 'transpose' for A^TA

      • n: number

        order of C (C is n×n); rows of op(A)

      • k: number

        columns of op(A) — the shared/contracted dimension

      • alpha: number

        scalar multiplier for op(A)*op(A)^T

      • A: Float32Array

        Float32Array, row-major or column-major (see layout)

      • lda: number

        leading dimension of A as stored

      • beta: number

        scalar multiplier for C

      • C: Float32Array

        Float32Array input/output matrix, row-major or column-major

      • ldc: number

        leading dimension of C as stored

      • Optionallayout: "column-major" | "row-major"

        storage layout shared by A/C when they're Float32Array (default: 'row-major') — same handling as sgemmtr; column-major C flips uplo internally so it still names the triangle you asked for

      Returns Promise<{ C: Float32Array; gpuTimeMs?: number }>

      updated C as a Float32Array (only the requested triangle changed)

    • Performs the symmetric rank-k update $$C \leftarrow \mathrm{uplo}(\alpha \mathrm{op}(A) \mathrm{op}(A)^{T} + \beta C)$$

      A and C are both kept GPU-resident. Each matrix's own layout (set at GpuMatrix.from time) determines the operation — there is no separate layout argument here. A must be a GpuMatrix whenever C is, and vice versa — mixing a GpuMatrix with a plain Float32Array is not supported.

      import { init, cleanup } from "wgblas";
      import { ssyrk } from "wgblas/ssyrk";
      import { GpuMatrix } from "wgblas/classes/GpuMatrix";

      const device = await init();

      // Entry (i,j) of C is the dot product of rows i and j of A; only the upper
      // triangle is written.
      const n = 3,
      k = 2;
      const A = new Float32Array([1, 0, 0, 1, 1, 1]);

      const AGpu = GpuMatrix.from(A, n, k, k, "row-major");
      const CGpu = GpuMatrix.from(new Float32Array(n * n), n, n, n, "row-major");

      console.log("A =");
      console.table([A.slice(0, 2), A.slice(2, 4), A.slice(4, 6)]);

      await ssyrk(
      device,
      "upper",
      "no-transpose",
      n,
      k,
      1,
      AGpu,
      AGpu.lda,
      0,
      CGpu,
      CGpu.lda,
      );

      const result = await CGpu.read();
      console.log("C = upper(A*A^T) =");
      console.table([result.slice(0, 3), result.slice(3, 6), result.slice(6, 9)]); // [[1,0,1],[0,1,1],[0,0,2]]

      AGpu.destroy();
      CGpu.destroy();
      if (typeof process !== "undefined") cleanup();

      Parameters

      • device: GPUDevice

        GPUDevice from init()

      • uplo: "lower" | "upper"

        'lower' to update only col <= row, 'upper' for col >= row

      • trans: "no-transpose" | "transpose"

        'no-transpose' for AA^T, 'transpose' for A^TA

      • n: number

        order of C (C is n×n); rows of op(A)

      • k: number

        columns of op(A) — the shared/contracted dimension

      • alpha: number

        scalar multiplier for op(A)*op(A)^T

      • A: GpuMatrix

        GpuMatrix

      • lda: number

        leading dimension of A (must equal A.lda)

      • beta: number

        scalar multiplier for C

      • C: GpuMatrix

        GpuMatrix (mutated in place; only the requested triangle changes)

      • ldc: number

        leading dimension of C (must equal C.lda)

      Returns Promise<{ gpuTimeMs?: number }>

      no C — it stays GPU-resident; call C.read() yourself for a CPU readback (see the example)