For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Python function
greedy_acceptance_sampler
greedy_acceptance_sampler()
max.nn.sampling.greedy_acceptance_sampler(draft_tokens, target_logits, token_bitmasks=None)
Target-only rejection sampler for speculative decoding.
Accepts a draft token only when it matches the argmax of the target logits. Recovered tokens are the target argmax at every draft position; the bonus token is the argmax at the final (+1) position.
When token_bitmasks is provided, grammar constraints mask the target
logits (fill -inf) before the argmax, so a grammar-invalid draft is
always rejected and recovered and bonus tokens always satisfy
structured-output constraints — same contract as the stochastic path.
Returns (first_rejected_idx, recovered_tokens, bonus_tokens)
-
Parameters:
-
- draft_tokens (TensorValue)
- target_logits (TensorValue)
- token_bitmasks (TensorValue | None)
-
Return type: