Paramètres de génération LLM

Generation parameters are the knobs that shape WHAT a model outputs once the prompt is fixed, not what you ask but how it produces it. Five of them show up everywhere. "max-tokens": the length ceiling of the answer, it drives cost and latency directly. "temperature": how strongly the model follows the highest probabilities, low gives predictable and repeatable, high gives varied and creative. "top-p" (nucleus): instead of tuning boldness, it cuts the tail of unlikely tokens and keeps only the plausible core. "repetition penalty": reduces words and phrasings that loop back. "stop-sequence": halts generation the moment a chosen word or symbol appears. Under the hood, two words explain why these knobs work. At each step the model gives every possible token a raw score, the **logits**, then a function called **softmax** turns those scores into probabilities. Temperature steps in right there: lowering it sharpens the distribution, so tokens that were already likely become even more likely and unlikely ones fade out. And random draws are not the only way to pick: **beam search**, offered by libraries such as Hugging Face Transformers, keeps several candidate sequences alive and retains the one that is most likely overall. It works well when output length is predictable (translation, summarization) and quickly turns repetitive on an open-ended answer. Mechanism of top-p: at 0.9, the model keeps the most likely next-token options until their combined probabilities reach at least 90%, then draws among them only. The rest of the tail is never sampled. Source: [Instagram reel DeMqr3Ijx22](https://www.instagram.com/reel/DeMqr3Ijx22/).

Strengths

Limitations

Best for

Official site

View on Coeurdar