IMPORTANT: To view this page as Markdown, append `.md` to the URL (e.g. /get-started.md). For the complete documentation index, see llms.txt.
Skip to main content
For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).

Mojo module

moe

Implements Mixture-of-Experts (MoE) routing, token dispatch, and expert computation kernels.

Functions​

  • ​eplb_remap: Launch the fused EPLB log->phy remap on GPU.
  • ​eplb_remap_kernel: Fused EPLB per tile_token rows of router idx; one thread per (n,k) element. Each block cooperatively caces the current layer's logcnt and log2phy slices in SMEM, then every thread does: HBM-load logical id -> SMEM-looup cnt -> int mod -> SMEM-Lookup phy id -> HBM-store.
  • ​group_limited_router_kernel: A manually fused MoE router with the group-limited strategy. It divides all the experts into n_groups groups and then finds the top topk_group groups with the highest scores. The final experts for each token are selected from the experts in the selected groups. The bias will be applied to the scores during the selection process, but the final weights will not include the bias.
  • ​moe_create_indices: Launches the MoE index creation kernel on GPU.
  • ​moe_create_indices_kernel: Builds MoE routing indices in one CTA using a block-wide scan.
  • ​router_group_limited: A manually fused MoE router with the group-limited strategy.
  • ​single_group_router: Launch the single-group MoE router on GPU.
  • ​single_group_router_eplb: Launches the single-group MoE router with EPLB log->phy remap on GPU.
  • ​single_group_router_eplb_kernel: Single-group MoE router fused with EPLB log->phy remap.
  • ​single_group_router_kernel: Single-group MoE router kernel. One block per token, one thread per expert.
  • ​sink_gate_router: Launch the fused sink-gate MoE router on GPU.
  • ​sink_gate_router_kernel: Fused sigmoid-gate MoE router with always-on sink (shared-expert) lanes.

Was this page helpful?