For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Mojo function
foreach
def foreach[dtype: DType, rank: Int, //, func: def[width: Int](Coord[*?]) capturing thin -> SIMD[dtype, width], *, target: StringSpan[ImmStaticOrigin] = StringSpan("cpu"), simd_width: Int = get_kernel_simd_width[dtype, target](), _trace_name: StringSpan[ImmStaticOrigin] = StringSpan("mogg.for_each")](tensor: ManagedTensorSlice[static_spec=tensor.static_spec], ctx: DeviceContext)
Apply the function func to each element of the tensor slice.
The func body receives the element index as a Coord. Use
coord_to_index_list to convert it to an IndexList if integer index
arithmetic is needed.
Parameters:
- dtype (
DType): The data type of the elements in the tensor slice. - rank (
Int): The rank of the tensor slice. - func (
def[width: Int](Coord[*?]) capturing thin -> SIMD[dtype, width]): The function to apply to each element of the tensor slice. - target (
StringSpan[ImmStaticOrigin]): Indicates the type of the target device (e.g. "cpu", "gpu"). - simd_width (
Int): The SIMD width for the target (usually leave this as its default value). - _trace_name (
StringSpan[ImmStaticOrigin]): Name of the executed operation displayed in the trace_description.
Args:
- tensor (
ManagedTensorSlice[static_spec=tensor.static_spec]): The output tensor slice which receives the return values fromfunc. - ctx (
DeviceContext): The call context (forward this from the custom operation).
def foreach[dtype: DType, rank: Int, //, FuncType: def[width: Int](Coord[*?]) -> SIMD[dtype, width] & RegisterPassable & ImplicitlyCopyable, *, target: StringSpan[ImmStaticOrigin] = StringSpan("cpu"), simd_width: Int = get_kernel_simd_width[dtype, target](), _trace_name: StringSpan[ImmStaticOrigin] = StringSpan("mogg.for_each")](var func: FuncType, tensor: ManagedTensorSlice[static_spec=tensor.static_spec], ctx: DeviceContext)
Apply a RegisterPassable body to each element of the tensor slice.
Value-argument twin of the foreach overload above: the body is a runtime
closure passed by value rather than a comptime parameter, so callers write
a unified closure instead of a capturing one. The body receives the
element index as a Coord; use coord_to_index_list to convert it to an
IndexList if integer index arithmetic is needed.
The wrapper captures both func and tensor and routes the store through
tensor._fused_store, which preserves the fused store path that
FusedOutputTensor depends on.
Parameters:
- dtype (
DType): The data type of the elements in the tensor slice. - rank (
Int): The rank of the tensor slice. - FuncType (
def[width: Int](Coord[*?]) -> SIMD[dtype, width]&RegisterPassable&ImplicitlyCopyable): The type of the per-element body closure. - target (
StringSpan[ImmStaticOrigin]): Indicates the type of the target device (e.g. "cpu", "gpu"). - simd_width (
Int): The SIMD width for the target (usually leave this as its default value). - _trace_name (
StringSpan[ImmStaticOrigin]): Name of the executed operation displayed in the trace_description.
Args:
- func (
FuncType): The function to apply to each element of the tensor slice. - tensor (
ManagedTensorSlice[static_spec=tensor.static_spec]): The output tensor slice which receives the return values fromfunc. - ctx (
DeviceContext): The call context (forward this from the custom operation).