For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Mojo struct
Struct_grouped_quantize_dynamic_block_scaled_with_row_indices
struct Struct_grouped_quantize_dynamic_block_scaled_with_row_indices
Registers the mo.grouped.quantize.dynamic.block.scaled.with_row_indices graph op with the graph compiler.
Row-gather variant of mo.grouped.quantize.dynamic.block.scaled: output
row i quantizes input row row_indices[i], which fuses the MoE expert
permutation gather into the quantize loads. A separate op carries the
extra operand because graph ops have no optional inputs, and the fusion
lives in the kernel because prologue fusion absorbs only elementwise and
view producers, not gathers. Fold it back into the base op once the
graph compiler fuses gathers generically.
Implemented traitsโ
Methodsโ
executeโ
static def execute[out_dtype: DType, scales_type: DType, in_dtype: DType, //, scales_rank: Int, target: StringSpan[ImmStaticOrigin]](output: ManagedTensorSlice[IOSpec[_, _].Output, static_spec=output.static_spec], scales: ManagedTensorSlice[IOSpec[_, _].Output, static_spec=scales.static_spec], input: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=input.static_spec], row_offsets: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=row_offsets.static_spec], scales_offsets: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=scales_offsets.static_spec], expert_ids: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=expert_ids.static_spec], sf_tensor: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=sf_tensor.static_spec], row_indices: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=row_indices.static_spec], context: DeviceContext)