IMPORTANT: To view this page as Markdown, append `.md` to the URL (e.g. /get-started.md). For the complete documentation index, see llms.txt.
Skip to main content
For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).

Python function

greedy_acceptance_sampler

greedy_acceptance_sampler()​

max.nn.sampling.greedy_acceptance_sampler(draft_tokens, target_logits, token_bitmasks=None)

source

Target-only rejection sampler for speculative decoding.

Accepts a draft token only when it matches the argmax of the target logits. Recovered tokens are the target argmax at every draft position; the bonus token is the argmax at the final (+1) position.

When token_bitmasks is provided, grammar constraints mask the target logits (fill -inf) before the argmax, so a grammar-invalid draft is always rejected and recovered and bonus tokens always satisfy structured-output constraints — same contract as the stochastic path.

Returns (first_rejected_idx, recovered_tokens, bonus_tokens)

Parameters:

Return type:

tuple[TensorValue, TensorValue, TensorValue]