For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Mojo module
reducescatter
Multi-GPU reducescatter implementation for distributed tensor reduction across GPUs.
comptime values
elementwise_epilogue_type
comptime elementwise_epilogue_type = def[dtype: DType, width: SIMDLength, *, alignment: Int](Coord[*?], SIMD[dtype, width]) capturing thin -> None
reducescatter_relay_residual_tuning_table
comptime reducescatter_relay_residual_tuning_table = Table(List(RelayTuningConfig(Int(-1), Int(-1), Int(20), Int(12), Int(36)), RelayTuningConfig(Int(4), Int(131072), Int(0), Int(0), Int(0)), RelayTuningConfig(Int(2), Int(131072), Int(0), Int(0), Int(0)), RelayTuningConfig(Int(2), Int(2147483648), Int(20), Int(12), Int(32)), __list_literal__=NoneType(None)), String("reducescatter_relay_residual_table"))
reducescatter_relay_tuning_table
comptime reducescatter_relay_tuning_table = Table(List(RelayTuningConfig(Int(-1), Int(-1), Int(20), Int(12), Int(44)), RelayTuningConfig(Int(4), Int(262144), Int(0), Int(0), Int(0)), RelayTuningConfig(Int(2), Int(131072), Int(0), Int(0), Int(0)), __list_literal__=NoneType(None)), String("reducescatter_relay_table"))
Structs
-
ReduceScatterConfig: Configuration for axis-aware reduce-scatter partitioning.
Functions
-
reducescatter: Per-device reducescatter operation with axis-aware scatter.