For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Mojo function
quantize_dynamic_scaled_fp8
def quantize_dynamic_scaled_fp8[out_dtype: DType, in_dtype: DType, scales_dtype: DType, InputFnType: def[width: Int, alignment: Int](row: Int, col: Int) -> SIMD[in_dtype, width] & RegisterPassable & ImplicitlyCopyable, //, group_size_or_per_token: Int, num_cols: Int, pdl_level: PDLLevel = PDLLevel.ON, row_bounded: Bool = False](input_fn: InputFnType, scaled_output: TileTensor[out_dtype, Engine=scaled_output.Engine, address_space=scaled_output.address_space, linear_idx_type=scaled_output.linear_idx_type], scales: TileTensor[scales_dtype, Engine=scales.Engine, address_space=scales.address_space, linear_idx_type=scales.linear_idx_type], scale_ub: Float32, ctx: DeviceContext, num_rows: Int, amax_floor: Float32 = 0, row_limit: Optional[Pointer[UInt32, ImmUntrackedOrigin]] = None)
TileTensor primary implementation of dynamic scaled FP8 quantization.
amax_floor lower-bounds a group's max-abs before the scale is derived,
which is how QAT recipes such as DeepSeek-V4's act_quant (floor 1e-4)
keep near-zero activation groups on the scale grid the checkpoint was
trained against. It is only honored on the float8_e8m0fnu scale path;
0 (the default) leaves the scale unchanged.
When row_bounded, row_limit points at a device-resident count of live
rows and the kernel grid-strides over [0, row_limit) instead of covering
all num_rows. Rows at or past that count keep whatever the output and
scales buffers already held.