IMPORTANT: To view this page as Markdown, append `.md` to the URL (e.g. /get-started.md). For the complete documentation index, see llms.txt.
Skip to main content
For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).

Mojo module

row_scales

Per-row output scales for the SM100 grouped matmul epilogues.

A grouped matmul scales each output row by its expert's scale. A RowScales carrier adds a second, per-row factor: output row m is multiplied by expert_scale * row_scales[m]. Activations quantized with a per-token tensor scale need it to undo that scale in the epilogue.

Implementations:

  • NullRowScales: zero-sized, Enabled is False, so it adds no kernel-argument bytes and no loads.
  • RealRowScales: one bfloat16 scale per output row in global memory, widened to float32 on load.

Structs​

  • ​NullRowScales: Zero-sized no-op row scales, the default for every caller.
  • ​RealRowScales: Row scales read from an array of one bfloat16 per output row.
  • ​TileRowScales: One output tile's scaled row factors, spread across a warp's lanes.

Traits​

  • ​RowScales: Per-row output scales applied in the epilogue.

Was this page helpful?