For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Python class
Proposed
Proposed
class max.pipelines.speculative.driver.Proposed(logits, hidden, reuse=<factory>, carry=<factory>, token=None)
Bases: object
One draft invocation’s contribution.
Carries the step’s logits alongside the hidden state the next step consumes.
-
Parameters:
-
- logits (TensorValue | None)
- hidden (list[TensorValue])
- reuse (list[TensorValue])
- carry (list[TensorValue])
- token (TensorValue | None)
carry
carry: list[TensorValue]
Step 0’s hidden state, already gathered at each accepted position.
Empty for a draft that hands back its whole verified window and lets the driver do the gather
hidden
hidden: list[TensorValue]
logits
logits: TensorValue | None
Logits over this call’s query, None only alongside token.
reuse
reuse: list[TensorValue]
Per-device work step 0 did that every later step reuses unchanged.
Read from the prefill only. A later step returns nothing here, because reusing step 0’s result is the point.
token
token: TensorValue | None = None
Step 0’s proposed token, for a draft that produced carry.