For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Mojo struct
GatedDeltaConv1dVerifyFwd
struct GatedDeltaConv1dVerifyFwd[rollback: Bool]
gated_delta_conv1d_fwd in a speculative verify graph.
The verify width K is the length of verify_width, whose contents are
never read. At K == 0 the window has no draft to reject, so the forward
launch writes it and the rollback launch does nothing. At K > 0 the
forward launch only reads it and the rollback launch places it at the
accepted length.
A rollback launch at K == 0 leaves conv_output_ragged unwritten. At
K > 0 the forward launch raises unless every row is one token and its
K drafts.
Tensor Shapes:
As gated_delta_conv1d_fwd, plus
- verify_width : [K] int64
Parameters
- rollback (
Bool): Whether this is the rollback's launch.
Implemented traits
Methods
execute
static def execute[work_dtype: DType, state_dtype: DType, target: StringSpan[ImmStaticOrigin]](conv_output_ragged: ManagedTensorSlice[IOSpec[_, _].Output, static_spec=conv_output_ragged.static_spec], qkv_input_ragged: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=qkv_input_ragged.static_spec], conv_weight: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=conv_weight.static_spec], conv_state: ManagedTensorSlice[IOSpec[_, _].MutableInput, static_spec=conv_state.static_spec], slot_idx: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=slot_idx.static_spec], input_row_offsets: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=input_row_offsets.static_spec], verify_width: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=verify_width.static_spec], ctx: DeviceContext)