Definition

Inference (AI)

Inference is the process of applying a trained AI model to new input to produce an output such as a prediction or generated text — as opposed to training.

Every request to an AI system triggers inference and drives its ongoing running costs. Its speed and resource demands determine whether a model runs sensibly in the cloud or locally.

← Back to glossary