IMPORTANT: To view this page as Markdown, append `.md` to the URL (e.g. /get-started.md). For the complete documentation index, see llms.txt.
Skip to main content
For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).

Mojo struct

AttnResMix

struct AttnResMix

Kimi K3 attention-residual softmax mixture, in one or two fused kernels.

Replaces the reference's ops.stack + RMS-normalize + score-reduce + softmax + weighted-reduce chain (6-7 separate kernel launches; see Kernels/lib/attn_res/mix.mojo's module docstring for the profile and the reassociation this fuses on), starting AFTER the stack (the caller still builds candidates with its own ops.stack). At hidden <= ATTN_RES_SPLIT_CHUNK that is one kernel, one CTA per token. Above it, attn_res_score_partials_gpu and attn_res_mix_from_partials_gpu split the hidden axis across CTAs, so a batch-1 decode fills more than one workgroup.

Tensor shapes: - output : [tokens, hidden] (OUT) - candidates : [tokens, num_candidates, hidden] - proj_weight : [1, hidden] - norm_weight : [hidden]

Implemented traitsโ€‹

AnyType, Deinitable, Movable

Methodsโ€‹

executeโ€‹

static def execute[dtype: DType, target: StringSpan[ImmStaticOrigin], eps: StringSpan[ImmStaticOrigin] = StringSpan("1e-6")](output: ManagedTensorSlice[IOSpec[_, _].Output, static_spec=output.static_spec], candidates: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=candidates.static_spec], proj_weight: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=proj_weight.static_spec], norm_weight: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=norm_weight.static_spec], ctx: DeviceContext)

Was this page helpful?