IMPORTANT: To view this page as Markdown, append `.md` to the URL (e.g. /get-started.md). For the complete documentation index, see llms.txt.
Skip to main content
For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).

Mojo struct

Struct_grouped_quantize_dynamic_block_scaled_with_row_indices

struct Struct_grouped_quantize_dynamic_block_scaled_with_row_indices

Registers the mo.grouped.quantize.dynamic.block.scaled.with_row_indices graph op with the graph compiler.

Row-gather variant of mo.grouped.quantize.dynamic.block.scaled: output row i quantizes input row row_indices[i], which fuses the MoE expert permutation gather into the quantize loads. A separate op carries the extra operand because graph ops have no optional inputs, and the fusion lives in the kernel because prologue fusion absorbs only elementwise and view producers, not gathers. Fold it back into the base op once the graph compiler fuses gathers generically.

Implemented traitsโ€‹

AnyType, Deinitable, Movable

Methodsโ€‹

executeโ€‹

static def execute[out_dtype: DType, scales_type: DType, in_dtype: DType, //, scales_rank: Int, target: StringSpan[ImmStaticOrigin]](output: ManagedTensorSlice[IOSpec[_, _].Output, static_spec=output.static_spec], scales: ManagedTensorSlice[IOSpec[_, _].Output, static_spec=scales.static_spec], input: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=input.static_spec], row_offsets: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=row_offsets.static_spec], scales_offsets: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=scales_offsets.static_spec], expert_ids: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=expert_ids.static_spec], sf_tensor: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=sf_tensor.static_spec], row_indices: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=row_indices.static_spec], context: DeviceContext)

Was this page helpful?