For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Mojo function
relu_cvt_bf16x2
def relu_cvt_bf16x2(hi: Float32, lo: Float32) -> SIMD[.bfloat16, 2]
Returns {relu(lo), relu(hi)} as a packed bf16x2, in one instruction.
cvt.rn.relu.bf16x2.f32 fuses the clamp-at-zero into the narrowing convert
and lowers to a single F2FP.RELU.BF16.F32.PACK_AB on SM100, so a relu that
costs one FMNMX per lane in f32 (PTX has no max.f32x2 at any ISA
version) becomes free. Pairs with a bf16x2 .fma() to fold two columns in
two instructions instead of three.
Inline asm rather than max(...).cast[bfloat16]() because LLVM emits the
plain cvt.rn.bf16x2.f32 and leaves the two max instructions standing.
Args:
- hi (
Float32): Value converted into the HIGH half of the result -- lane 1. PTX names the packed halves high-first, the reverse of the SIMD index order, so swapping these silently misaligns every lane downstream. - lo (
Float32): Value converted into the LOW half of the result -- lane 0.
Returns: