AI & ML Tech Glossary
Clear definitions of 500+ AI, ML, and systems terms, built for professionals.
P
Prefix Caching
Prefix caching reuses precomputed model states for a shared prompt prefix (system prompts, tool schemas), reducing repeated prefill compute and improving time-to-first-token and serving capacity...
Prompt Ensembling
Prompt ensembling runs multiple prompt variants or sampling runs for the same task and aggregates results, such as voting or judge-model selection, to improve LLM...
Parameter-Efficient Fine-Tuning (PEFT)
PEFT adapts a pretrained model by training only a small number of additional or selected parameters, which reduces the compute, memory, and storage required compared...
Prompt Leakage
Prompt leakage is when an AI system unintentionally reveals hidden prompts or sensitive context—like system instructions, tool schemas, or private RAG documents—often due to prompt...
Prompt Caching
HotPrompt caching reuses computed states for repeated prompt prefixes (like system prompts or tool schemas), cutting prompt-processing time and improving time-to-first-token and serving cost for...