For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Python module
max.experimental.nn.common_layers.functional_kernels
Functional wrappers for MAX kernel operations used in attention layers.
Functions
flash_attention_gpu | Computes flash attention using a GPU-optimized kernel. |
|---|---|
flash_attention_ragged | Computes flash (self) attention provided the !mo.opaque KV Cache. |
flash_attention_ragged_gpu | Computes flash attention for ragged inputs using a GPU-optimized kernel, without a KV cache. |
fused_silu | Performs the fused SILU operation for all the MLPs in the EP MoE module. |
grouped_matmul_ragged | Performs the grouped matmul used in the MoE layer. |
moe_create_indices | Creates indices for the MoE layer. |
moe_router_group_limited | Routes tokens with the group-limited MoE router. |
rms_norm_key_cache | Applies RMSNorm to the new entries in the KV cache. |
rope_split_store_ragged | Applies RoPE to Q and K from a flat QKV buffer and stores K/V to the cache. |
stack_device_shards | Reassembles a per-device weight-shard bundle into one Sharded tensor. |