IMPORTANT: To view this page as Markdown, append `.md` to the URL (e.g. /get-started.md). For the complete documentation index, see llms.txt.
Skip to main content
For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).

Mojo function

dual_elementwise

def dual_elementwise[simd_width: Int, *, target: StringSpan[ImmStaticOrigin] = StringSpan("gpu"), _trace_description: StringSpan[ImmStaticOrigin] = StringSpan("dual_elementwise")](shape_0: Coord, shape_1: Coord, context: DeviceContext, func_0: T, func_1: T)

Executes two elementwise functions over their respective shapes in a single GPU kernel launch. Each thread processes elements from both shapes, fusing two independent elementwise passes into one.

Parameters:

  • ​simd_width (Int): The SIMD vector width to use.
  • ​target (StringSpan[ImmStaticOrigin]): The target to run on (must be GPU).
  • ​_trace_description (StringSpan[ImmStaticOrigin]): Description of the trace.

Args:

  • ​shape_0 (Coord): The shape for the first function.
  • ​shape_1 (Coord): The shape for the second function.
  • ​context (DeviceContext): The device context to use.
  • ​func_0 (T): The first body function.
  • ​func_1 (T): The second body function.

Raises:

If the operation fails.

Was this page helpful?