For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Mojo struct
GatedDeltaRecurrenceVerifyRingFwd
struct GatedDeltaRecurrenceVerifyRingFwd
Gated DeltaNet recurrence over a speculative verify window.
Produces the same recurrence_output as gated_delta_recurrence_fwd
without writing recurrent_state. Each token writes its decay, raw key
and delta factor to its request's ring row, which
gated_delta_state_fold applies to the pool once acceptance is known.
The verify width K is the length of verify_width, whose contents are
never read. At K == 0 there is no draft to reject, so this runs
gated_delta_recurrence_fwd on recurrent_state instead and writes no
ring record. The fold skips at the same width. At K > 0 every row's
window of K + 1 tokens must fit RING_LEN.
Tensor Shapes: - recurrence_output : [total_seq_len, value_dim] (OUT) - qkv_conv_output : [total_seq_len, conv_dim] - decay_per_token : [total_seq_len, num_value_heads] - beta_per_token : [total_seq_len, num_value_heads] - recurrent_state : [max_slots, num_value_heads, KD, VD] (MUT) - ring : [ring_rows, num_key_heads, RING_LEN, record_stride] (MUT) - slot_idx : [batch_size] uint32 - ring_slot_idx : [batch_size] uint32 - input_row_offsets : [batch_size + 1] uint32 - verify_width : [K] int64
Implemented traitsโ
Methodsโ
executeโ
static def execute[work_dtype: DType, state_dtype: DType, ring_dtype: DType, target: StringSpan[ImmStaticOrigin]](recurrence_output: ManagedTensorSlice[IOSpec[_, _].Output, static_spec=recurrence_output.static_spec], qkv_conv_output: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=qkv_conv_output.static_spec], decay_per_token: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=decay_per_token.static_spec], beta_per_token: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=beta_per_token.static_spec], recurrent_state: ManagedTensorSlice[IOSpec[_, _].MutableInput, static_spec=recurrent_state.static_spec], ring: ManagedTensorSlice[IOSpec[_, _].MutableInput, static_spec=ring.static_spec], slot_idx: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=slot_idx.static_spec], ring_slot_idx: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=ring_slot_idx.static_spec], input_row_offsets: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=input_row_offsets.static_spec], verify_width: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=verify_width.static_spec], ctx: DeviceContext)