IMPORTANT: To view this page as Markdown, append `.md` to the URL (e.g. /get-started.md). For the complete documentation index, see llms.txt.
Skip to main content
For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).

Mojo struct

TileScheduler

struct TileScheduler[group_offsets_origin: ImmOrigin, offsets_layout: Layout, //, *, static_MN: Int, tile_shape: IndexList[Int(3)], cluster: IndexList[Int(3)] = Index[Int, Int, Int](Int(1), Int(1), Int(1)), cta_group: Int = Int(1), swizzle: Bool = False, swapAB: Bool = True]

Schedules output tiles across CTAs for a persistent grouped matmul kernel.

Parameters

  • group_offsets_origin (ImmOrigin): Memory origin of group_offsets (inferred).
  • offsets_layout (Layout): Memory layout of group_offsets (inferred).
  • static_MN (Int): Size of the static (non-reducing) output dimension. When swapAB is true this is M, otherwise N.
  • tile_shape (IndexList[Int(3)]): Per-tile shape (M, N, K) of output tiles in the original non-swapped AB orientation.
  • cluster (IndexList[Int(3)]): CTA cluster multicast shape (M, N, K). Only the M dimension is supported; cluster[1] and cluster[2] must be 1.
  • cta_group (Int): CTAs cooperating per tile group along M. Must equal cluster[0].
  • swizzle (Bool): Whether to swizzle block indices for improved L2 reuse.
  • swapAB (Bool): Whether to swap A and B operands. When true, the static dimension is M; when false, it is N.

Fields

  • num_active_experts (Int):
  • group_offsets (LayoutTensor[DType.uint32, offsets_layout, group_offsets_origin]):
  • current_iter (Int32):
  • current_group_idx (UInt32):
  • current_dynamic_dim_cumsum (UInt32):
  • block_idx_start (UInt32):

Implemented traits

AnyType, Copyable, Deinitable, ImplicitlyCopyable, Movable, RegisterPassable, TrivialRegisterPassable

comptime members

cta_group_tile_shape

comptime cta_group_tile_shape = Index[Int, Int](Int((mul tile_shape[Int(0)], cta_group)), Int((mul tile_shape[Int(1)], cta_group)))

div_dynamic_block

comptime div_dynamic_block = FastDiv(Index[Int, Int](Int((mul tile_shape[Int(0)], cta_group)), Int((mul tile_shape[Int(1)], cta_group)))[Int(1) if swapAB else Int(0)])

dynamic_dim

comptime dynamic_dim = Int(1) if swapAB else Int(0)

kNum1DBlocksPerGroup

comptime kNum1DBlocksPerGroup = UInt32(16)

num_static_dim_blocks

comptime num_static_dim_blocks = SIMD(ceildiv(static_MN, tile_shape[Int(0) if swapAB else Int(1)]))

static_dim

comptime static_dim = Int(0) if swapAB else Int(1)

Methods

__init__

def __init__(num_active_experts: Int, group_offsets: LayoutTensor[DType.uint32, offsets_layout, group_offsets_origin]) -> Self

fetch_next_work

def fetch_next_work(mut self) -> WorkInfo

Returns:

WorkInfo