For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Mojo function
rms_norm_gpu_warp_tiling
def rms_norm_gpu_warp_tiling[mut: Bool, LayoutType: TensorLayout, origin: Origin[mut=mut], dtype: DType, rank: Int, Engine: TensorEngine, shape_types: TypeList[shape_types.values], //, simd_width: Int, max_warps_per_block: Int, chunks_per_thread: Int, exact_fit: Bool, input_fn: def[width: Int](Coord[*?]) capturing thin -> SIMD[dtype, width], output_fn: def[width: SIMDLength, alignment: Int](Coord[*?], SIMD[dtype, width]) capturing thin -> None, multiply_before_cast: Bool, pdl_level: PDLLevel = PDLLevel.ON](row_spec: _RowSpec[shape_types], gamma: TileTensor[dtype, LayoutType, origin, Engine=Engine], epsilon: Float32, weight_offset: Float32)