IMPORTANT: To view this page as Markdown, append `.md` to the URL (e.g. /get-started.md). For the complete documentation index, see llms.txt.
Skip to main content
For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).

Mojo struct

Struct_smallm_streaming_matmul

struct Struct_smallm_streaming_matmul

MOGG wrapper for the MI355X small-M streaming matmul.

Computes c = a @ b^T where b is a bf16 weight ALREADY permuted into the fragment-major layout of smallm_preshuffle_b (done on the CPU at weight-load time). The layout is private to this op: reading a row-major weight here is silently wrong, so nothing routes here through generic dispatch — MiniMax-M3's MTP draft emits this op explicitly on MI355X for its decode-band vocab projections.

N/K come from the weight's static [N, K] layout. The streaming kernel wins up to M == 32 (measured on the MiniMax-M3 vocab-head shapes); larger runtime M falls back to generic matmul dispatch over b, the row-major twin of the same weight, so a batch-config change can never turn into a regression or a failure.

Implemented traits​

AnyType, Deinitable, Movable

Methods​

execute​

static def execute[target: StringSpan[ImmStaticOrigin]](c: ManagedTensorSlice[IOSpec[_, _].Output, static_spec=c.static_spec], a_scratch: ManagedTensorSlice[IOSpec[_, _].Output, static_spec=a_scratch.static_spec], a: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=a.static_spec], b_shuffled: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=b_shuffled.static_spec], b: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=b.static_spec], context: DeviceContext)

Executes the streaming matmul over a preshuffled weight.

Constraints:

b_shuffled must have a static [N, K] shape with N % 16 == 0 and K % 256 == 0; b is the same weight in row-major [N, K].

Args:

Was this page helpful?