For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Mojo function
dual_elementwise
def dual_elementwise[simd_width: Int, *, target: StringSpan[ImmStaticOrigin] = StringSpan("gpu"), _trace_description: StringSpan[ImmStaticOrigin] = StringSpan("dual_elementwise")](shape_0: Coord, shape_1: Coord, context: DeviceContext, func_0: T, func_1: T)
Executes two elementwise functions over their respective shapes in a single GPU kernel launch. Each thread processes elements from both shapes, fusing two independent elementwise passes into one.
Parameters:
- simd_width (
Int): The SIMD vector width to use. - target (
StringSpan[ImmStaticOrigin]): The target to run on (must be GPU). - _trace_description (
StringSpan[ImmStaticOrigin]): Description of the trace.
Args:
- shape_0 (
Coord): The shape for the first function. - shape_1 (
Coord): The shape for the second function. - context (
DeviceContext): The device context to use. - func_0 (
T): The first body function. - func_1 (
T): The second body function.
Raises:
If the operation fails.