IMPORTANT: To view this page as Markdown, append `.md` to the URL (e.g. /get-started.md). For the complete documentation index, see llms.txt.
Skip to main content
For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).

Mojo struct

GatedDeltaRecurrenceVerifyRingFwd

struct GatedDeltaRecurrenceVerifyRingFwd

Gated DeltaNet recurrence over a speculative verify window.

Produces the same recurrence_output as gated_delta_recurrence_fwd without writing recurrent_state. Each token writes its decay, raw key and delta factor to its request's ring row, which gated_delta_state_fold applies to the pool once acceptance is known.

The verify width K is the length of verify_width, whose contents are never read. At K == 0 there is no draft to reject, so this runs gated_delta_recurrence_fwd on recurrent_state instead and writes no ring record. The fold skips at the same width. At K > 0 every row's window of K + 1 tokens must fit RING_LEN.

Tensor Shapes: - recurrence_output : [total_seq_len, value_dim] (OUT) - qkv_conv_output : [total_seq_len, conv_dim] - decay_per_token : [total_seq_len, num_value_heads] - beta_per_token : [total_seq_len, num_value_heads] - recurrent_state : [max_slots, num_value_heads, KD, VD] (MUT) - ring : [ring_rows, num_key_heads, RING_LEN, record_stride] (MUT) - slot_idx : [batch_size] uint32 - ring_slot_idx : [batch_size] uint32 - input_row_offsets : [batch_size + 1] uint32 - verify_width : [K] int64

Implemented traitsโ€‹

AnyType, Deinitable, Movable

Methodsโ€‹

executeโ€‹

static def execute[work_dtype: DType, state_dtype: DType, ring_dtype: DType, target: StringSpan[ImmStaticOrigin]](recurrence_output: ManagedTensorSlice[IOSpec[_, _].Output, static_spec=recurrence_output.static_spec], qkv_conv_output: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=qkv_conv_output.static_spec], decay_per_token: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=decay_per_token.static_spec], beta_per_token: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=beta_per_token.static_spec], recurrent_state: ManagedTensorSlice[IOSpec[_, _].MutableInput, static_spec=recurrent_state.static_spec], ring: ManagedTensorSlice[IOSpec[_, _].MutableInput, static_spec=ring.static_spec], slot_idx: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=slot_idx.static_spec], ring_slot_idx: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=ring_slot_idx.static_spec], input_row_offsets: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=input_row_offsets.static_spec], verify_width: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=verify_width.static_spec], ctx: DeviceContext)

Was this page helpful?