For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Mojo function
group_norm_gpu_block
def group_norm_gpu_block[LayoutType: TensorLayout, origin: MutOrigin, //, dtype: DType, simd_width: Int, InputFnType: def[width: Int](row: Int, col: Int) -> SIMD[dtype, width] & RegisterPassable & ImplicitlyCopyable, GammaFnType: def[width: Int](Coord[*?]) -> SIMD[dtype, width] & RegisterPassable & ImplicitlyCopyable, BetaFnType: def[width: Int](Coord[*?]) -> SIMD[dtype, width] & RegisterPassable & ImplicitlyCopyable](output: TileTensor[dtype, LayoutType, origin], epsilon: Float32, num_groups: Int32, channels_per_group: Int32, spatial: Int32, input_fn: InputFnType, gamma_fn: GammaFnType, beta_fn: BetaFnType)
Block-per-row group_norm kernel.
input_fn, gamma_fn, and beta_fn are trailing host-layout arguments
so enqueue does not DevicePassable-encode those capturing closures.