IMPORTANT: To view this page as Markdown, append `.md` to the URL (e.g. /get-started.md). For the complete documentation index, see llms.txt.
Skip to main content
For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).

Mojo module

concat

Implements tensor concatenation along a specified axis for CPU and GPU targets.

comptime values​

elementwise_epilogue_type​

comptime elementwise_epilogue_type = def[c_type: DType, rank: Int, width: SIMDLength = 1, *, alignment: Int = Int(1)](IndexList[rank], SIMD[c_type, width]) capturing thin -> None

ElementwiseEpilogueFn​

comptime ElementwiseEpilogueFn = def[c_type: DType, rank: Int, width: SIMDLength, *, alignment: Int](IndexList[rank], SIMD[c_type, width]) -> None & RegisterPassable & ImplicitlyCopyable

Value-taking counterpart of elementwise_epilogue_type.

A unified closure passed as a runtime value carries the origins of its captures, so the output it writes stays alive until the kernel that calls it has launched. width and alignment have no defaults; pass both where the legacy type relied on its defaults.

Functions​

  • ​concat: Concatenates inputs along axis into output for the given target.
  • ​concat_shape: Compute the output shape of a pad operation, and assert the inputs are compatible.
  • ​fused_concat: Concatenates inputs produced by input_fn along axis into output, applying output_0_fn to each element.
  • ​memcpy_or_fuse: Copies n bytes from src_data into dest_data at out_byte_offset, applying epilogue_fn elementwise when has_epilogue is set.
  • ​preferred_simd_width: SIMD scalar count for fused GPU concat vectorization.

Was this page helpful?