For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Mojo module
row_scales
Per-row output scales for the SM100 grouped matmul epilogues.
A grouped matmul scales each output row by its expert's scale. A RowScales
carrier adds a second, per-row factor: output row m is multiplied by
expert_scale * row_scales[m]. Activations quantized with a per-token tensor
scale need it to undo that scale in the epilogue.
Implementations:
NullRowScales: zero-sized,EnabledisFalse, so it adds no kernel-argument bytes and no loads.RealRowScales: onebfloat16scale per output row in global memory, widened tofloat32on load.
Structs
-
NullRowScales: Zero-sized no-op row scales, the default for every caller. -
RealRowScales: Row scales read from an array of onebfloat16per output row. -
TileRowScales: One output tile's scaled row factors, spread across a warp's lanes.
Traits
-
RowScales: Per-row output scales applied in the epilogue.