IMPORTANT: To view this page as Markdown, append `.md` to the URL (e.g. /get-started.md). For the complete documentation index, see llms.txt.
Skip to main content
For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).

Mojo function

relu_cvt_bf16x2

def relu_cvt_bf16x2(hi: Float32, lo: Float32) -> SIMD[.bfloat16, 2]

Returns {relu(lo), relu(hi)} as a packed bf16x2, in one instruction.

cvt.rn.relu.bf16x2.f32 fuses the clamp-at-zero into the narrowing convert and lowers to a single F2FP.RELU.BF16.F32.PACK_AB on SM100, so a relu that costs one FMNMX per lane in f32 (PTX has no max.f32x2 at any ISA version) becomes free. Pairs with a bf16x2 .fma() to fold two columns in two instructions instead of three.

Inline asm rather than max(...).cast[bfloat16]() because LLVM emits the plain cvt.rn.bf16x2.f32 and leaves the two max instructions standing.

Args:

  • ​hi (Float32): Value converted into the HIGH half of the result -- lane 1. PTX names the packed halves high-first, the reverse of the SIMD index order, so swapping these silently misaligns every lane downstream.
  • ​lo (Float32): Value converted into the LOW half of the result -- lane 0.

Returns:

SIMD[.bfloat16, 2]

Was this page helpful?