Inference is the process of running a trained model to generate an output.
Every time you send a prompt to a model and receive a response, the model is performing inference. For an agent, inference may happen many times during a single task as the model plans, uses tools, observes results, and decides what to do next.
Inference settings can affect quality, speed, and cost. Depending on the model, teams may be able to adjust factors such as reasoning effort, output length, sampling settings, or the amount of computation used for each request.
These settings should be evaluated alongside the model itself because the same model can behave very differently under different inference configurations.