For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Mojo module
moe
Implements Mixture-of-Experts (MoE) routing, token dispatch, and expert computation kernels.
Functions
-
eplb_remap: Launch the fused EPLB log->phy remap on GPU. -
eplb_remap_kernel: Fused EPLB per tile_token rows of router idx; one thread per (n,k) element. Each block cooperatively caces the current layer's logcnt and log2phy slices in SMEM, then every thread does: HBM-load logical id -> SMEM-looup cnt -> int mod -> SMEM-Lookup phy id -> HBM-store. -
group_limited_router_kernel: A manually fused MoE router with the group-limited strategy. It divides all the experts inton_groupsgroups and then finds the toptopk_groupgroups with the highest scores. The final experts for each token are selected from the experts in the selected groups. The bias will be applied to the scores during the selection process, but the final weights will not include the bias. -
moe_create_indices: Launches the MoE index creation kernel on GPU. -
moe_create_indices_kernel: Builds MoE routing indices in one CTA using a block-wide scan. -
moe_finalize: Fuses the MoE unpermute gather with the top-k weighted row sum. -
router_group_limited: A manually fused MoE router with the group-limited strategy. -
single_group_router: Launch the single-group MoE router on GPU. -
single_group_router_eplb: Launches the single-group MoE router with EPLB log->phy remap on GPU. -
single_group_router_eplb_kernel: Single-group MoE router fused with EPLB log->phy remap. -
single_group_router_kernel: Single-group MoE router kernel. One block per token, one thread per expert. -
sink_gate_router: Launch the fused sink-gate MoE router on GPU. -
sink_gate_router_kernel: Fused sigmoid-gate MoE router with always-on sink (shared-expert) lanes.