IMPORTANT: To view this page as Markdown, append `.md` to the URL (e.g. /get-started.md). For the complete documentation index, see llms.txt.
Skip to main content
For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).

Python class

AcceptanceSampler

AcceptanceSampler​

class max.nn.sampling.AcceptanceSampler(synthetic_acceptance_rate=None, num_draft_steps=1, use_stochastic=False, draft_proposal='argmax', vocab_size=None, relaxed_topk=None, relaxed_delta=None)

source

Bases: object

Dispatches between greedy, synthetic, and stochastic acceptance.

  • synthetic_acceptance_rate set → synthetic (benchmarking) mode. The per-position acceptance probability is calibrated so that the mean joint acceptance across num_draft_steps matches the configured rate, via compute_synthetic_acceptance_base_rate().
  • use_stochastic=True → stochastic rejection sampling. The caller must then pass per-row sampling params (temperature, top_k, max_k, top_p, min_top_p) at call time.
  • Otherwise → greedy (accept iff draft token == target argmax).

Synthetic mode takes priority over stochastic when both are configured; the stochastic params are ignored in that case.

relaxed_topk / relaxed_delta require draft_proposal="argmax"; the relaxed rule assumes the drafted token is the draft’s own argmax, so it does not carry over to a sampled proposal.

Parameters:

  • synthetic_acceptance_rate (float | None)
  • num_draft_steps (int)
  • use_stochastic (bool)
  • draft_proposal (Literal['argmax', 'sampled'])
  • vocab_size (int | None)
  • relaxed_topk (int | None)
  • relaxed_delta (float | None)

acceptance_rule​

property acceptance_rule: Literal['synthetic', 'stochastic', 'greedy']

source

The rule __call__() dispatches to, in its precedence order.

This is the only place the acceptance rule in effect is decided: SpeculativeConfig.rejection_sampling_strategy is read by nothing.