IMPORTANT: To view this page as Markdown, append `.md` to the URL (e.g. /get-started.md). For the complete documentation index, see llms.txt.
Skip to main content
For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).

Mojo struct

QuantizeDynamicScaledFloat8RowBounded

struct QuantizeDynamicScaledFloat8RowBounded

Registers the mo.quantize_dynamic_scaled_float8.row_bounded graph op.

Same numerics as mo.quantize_dynamic_scaled_float8, but the rows it touches stop at a count the GPU publishes rather than at the height of the input tensor. The EP MoE down projection needs this: its activation buffer is sized for the worst-case dispatch, the dispatch kernel writes the live row count into row_offsets, and the host never learns that count.

A separate symbol rather than an operand on the shared op, because the graph compiler's RMS-norm and all-reduce fusion patterns match the shared op by its two-operand signature.

Implemented traitsโ€‹

AnyType, Deinitable, Movable

Methodsโ€‹

executeโ€‹

static def execute[input_type: DType, scales_type: DType, output_type: DType, //, group_size_or_per_token: Int, target: StringSpan[ImmStaticOrigin]](output: ManagedTensorSlice[IOSpec[_, _].Output, static_spec=output.static_spec], scales: ManagedTensorSlice[IOSpec[_, _].Output, static_spec=scales.static_spec], input: ManagedTensorSlice[IOSpec[_, _].FusedInput, static_spec=input.static_spec], scale_ub: Float32, row_offsets: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=row_offsets.static_spec], ctx: DeviceContext)

Was this page helpful?