Top-p sampling, also called nucleus sampling, is a text generation decoding strategy that samples the next token only from the smallest set of tokens whose cumulative probability mass exceeds a threshold p.
What is Top-p Sampling (Nucleus Sampling)?
In an autoregressive language model, each next-token step produces a probability distribution over the vocabulary. Greedy decoding always picks the most likely token, while pure sampling can pick from the full distribution and may produce low probability, low quality tokens. Top-p sampling constrains randomness by dynamically selecting a “nucleus” of tokens: sort tokens by probability, then keep adding tokens until their cumulative probability reaches p, for example p = 0.9. The model then samples from this truncated set, usually after renormalizing probabilities. Because the size of the nucleus changes per step, top-p adapts to the model’s uncertainty: when the distribution is sharp, the nucleus is small, and outputs become more deterministic. When the distribution is flat, the nucleus grows, preserving diversity without allowing extremely unlikely tokens.
Where top-p sampling is used and why it matters
Top-p sampling is widely used in generative AI applications such as chat assistants, story generation, and code completion to balance creativity and coherence. It is often paired with temperature, repetition penalties, and max token limits. Compared with top-k sampling, which keeps a fixed number of candidates, top-p tends to be more stable across contexts because it uses probability mass rather than a fixed cutoff. Choosing p too low can make outputs repetitive or overly safe, while choosing p too high can increase hallucinations and stylistic drift.
Examples
1) With p = 0.9, the model may sample among several plausible continuations for an open-ended prompt, while still avoiding rare tokens.
2) For factual Q and A, using a lower p, like 0.7 to 0.85, can reduce variability and improve consistency.
3) In creative writing, a higher p, like 0.95, can introduce more diverse phrasing while remaining readable.
FAQs
1. How is top-p different from top-k sampling?
Top-k keeps a fixed number of tokens, while top-p keeps a variable number whose probabilities sum to at least p.
2. Should I use top-p or temperature?
They control different effects. Top-p limits the candidate set, while temperature reshapes probabilities. Many systems use both.
3. Does top-p reduce hallucinations?
It can reduce extreme low probability tokens, but hallucinations can still occur. Grounding methods like RAG matter more.
4. What is a common default value for p?
Many APIs default around 0.9 to 0.95, but the best value depends on task and desired creativity.