IMPORTANT: To view this page as Markdown, append `.md` to the URL (e.g. /get-started.md). For the complete documentation index, see llms.txt.
Skip to main content
For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).

Mojo function

gated_group_rmsnorm_gpu

def gated_group_rmsnorm_gpu[dtype: DType, gate_dtype: DType, group_size: Int](output: TileTensor[dtype, Engine=output.Engine, address_space=output.address_space, linear_idx_type=output.linear_idx_type], y: TileTensor[dtype, Engine=y.Engine, address_space=y.address_space, linear_idx_type=y.linear_idx_type], gate: TileTensor[gate_dtype, Engine=gate.Engine, address_space=gate.address_space, linear_idx_type=gate.linear_idx_type], weight: TileTensor[.float32, Engine=weight.Engine, address_space=weight.address_space, linear_idx_type=weight.linear_idx_type], n_rows: Int, num_groups: Int, eps: Float32, ctx: DeviceContext)

Enqueues the fused gated group-RMSNorm; one warp per (row, group).

Parameters:

  • ​dtype (DType): Element type of y and output.
  • ​gate_dtype (DType): Element type of gate.
  • ​group_size (Int): Width of each independently normalized group.

Args:

Was this page helpful?