For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Mojo function
copy_dram_to_local
def copy_dram_to_local[src_thread_layout: Layout, num_threads: Int = src_thread_layout.size(), thread_scope: ThreadScope = ThreadScope.BLOCK, block_dim_count: Int = Int(1), cache_policy: CacheOperation = CacheOperation.ALWAYS](dst: LayoutTensor[address_space=dst.address_space, element_layout=dst.element_layout, layout_int_type=dst.layout_int_type, linear_idx_type=dst.linear_idx_type, masked=dst.masked, alignment=dst.alignment], src: LayoutTensor[address_space=src.address_space, element_layout=src.element_layout, layout_int_type=src.layout_int_type, linear_idx_type=src.linear_idx_type, masked=src.masked, alignment=src.alignment], src_base: LayoutTensor[address_space=src_base.address_space, element_layout=src_base.element_layout, layout_int_type=src_base.layout_int_type, linear_idx_type=src_base.linear_idx_type, masked=src_base.masked, alignment=src_base.alignment], offset: Optional[Int] = None)
Efficiently copy data from global memory (DRAM) to registers for AMD GPUs.
This function implements an optimized memory transfer operation specifically for AMD GPU architectures. It utilizes the hardware's buffer_load intrinsic to efficiently transfer data from global memory to registers while handling bounds checking. The function distributes the copy operation across multiple threads for maximum throughput.
Notes:
- The offset calculation method significantly impacts performance. Current implementation optimizes for throughput over flexibility.
- This function is particularly useful for prefetching data into registers before performing computations, reducing memory access latency.
Constraints:
- Only supported on AMD GPUs.
- The destination element layout size must match the SIMD width.
- Source fragments must be rank 2 with known dimensions.
Parameters:
- src_thread_layout (
Layout): The layout used to distribute the source tensor across threads. This determines how the workload is divided among participating threads. - num_threads (
Int): Total number of threads in the thread block. Threads beyondsrc_thread_layout.size()will be disabled and not participate in the copy operation. - thread_scope (
ThreadScope): Defines whether operations are performed atBLOCKorWARPlevel.BLOCKscope involves all threads in a thread block, whileWARPscope restricts operations to threads within the same warp. Defaults toThreadScope.BLOCK. - block_dim_count (
Int): The number of dimensions in the thread block. - cache_policy (
CacheOperation): The cache policy to use for the copy operation. Defaults toCacheOperation.ALWAYS.
Args:
- dst (
LayoutTensor[address_space=dst.address_space, element_layout=dst.element_layout, layout_int_type=dst.layout_int_type, linear_idx_type=dst.linear_idx_type, masked=dst.masked, alignment=dst.alignment]): The destination tensor in register memory (LOCAL address space). - src (
LayoutTensor[address_space=src.address_space, element_layout=src.element_layout, layout_int_type=src.layout_int_type, linear_idx_type=src.linear_idx_type, masked=src.masked, alignment=src.alignment]): The source tensor in global memory (DRAM) to be copied. - src_base (
LayoutTensor[address_space=src_base.address_space, element_layout=src_base.element_layout, layout_int_type=src_base.layout_int_type, linear_idx_type=src_base.linear_idx_type, masked=src_base.masked, alignment=src_base.alignment]): The original global memory tensor from which src is derived. This is used to construct the buffer struct required by AMD'sbuffer_loadintrinsic. - offset (
Optional[Int]): The offset in the global memory.