For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Mojo struct
Struct_smallm_streaming_matmul
struct Struct_smallm_streaming_matmul
MOGG wrapper for the MI355X small-M streaming matmul.
Computes c = a @ b^T where b is a bf16 weight ALREADY permuted
into the fragment-major layout of smallm_preshuffle_b (done on the
CPU at weight-load time). The layout is private to this op: reading a
row-major weight here is silently wrong, so nothing routes here through
generic dispatch — MiniMax-M3's MTP draft emits this op explicitly on
MI355X for its decode-band vocab projections.
N/K come from the weight's static [N, K] layout. The
streaming kernel wins up to M == 32 (measured on the MiniMax-M3
vocab-head shapes); larger runtime M falls back to generic matmul
dispatch over b, the row-major twin of the same weight, so a
batch-config change can never turn into a regression or a failure.
Implemented traits
Methods
execute
static def execute[target: StringSpan[ImmStaticOrigin]](c: ManagedTensorSlice[IOSpec[_, _].Output, static_spec=c.static_spec], a_scratch: ManagedTensorSlice[IOSpec[_, _].Output, static_spec=a_scratch.static_spec], a: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=a.static_spec], b_shuffled: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=b_shuffled.static_spec], b: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=b.static_spec], context: DeviceContext)
Executes the streaming matmul over a preshuffled weight.
Constraints:
b_shuffled must have a static [N, K] shape with
N % 16 == 0 and K % 256 == 0; b is the same weight
in row-major [N, K].
Args:
- c (
ManagedTensorSlice[IOSpec[_, _].Output, static_spec=c.static_spec]): Output[M, N]bf16 tensor. - a_scratch (
ManagedTensorSlice[IOSpec[_, _].Output, static_spec=a_scratch.static_spec]): Graph-managed[32, K]bf16 workspace for the activation shuffle. Graph memory keeps the captured launches' pointers valid across device-graph replays; a transient buffer here is silently wrong under capture. - a (
ManagedTensorSlice[IOSpec[_, _].Input, static_spec=a.static_spec]): Activation[M, K]bf16 tensor (row-major). - b_shuffled (
ManagedTensorSlice[IOSpec[_, _].Input, static_spec=b_shuffled.static_spec]): Weight insmallm_preshuffle_blayout. - b (
ManagedTensorSlice[IOSpec[_, _].Input, static_spec=b.static_spec]): The same weight, row-major (the above-band fallback operand). - context (
DeviceContext): The device context.