Generate a response
Generates a response for the provided prompt
Body
Model name
Text for the model to generate a response from
Used for fill-in-the-middle models, text that appears after the user prompt and before the model response
Base64-encoded images for models that support image input
Structured output format for the model to generate a response from. Supports either the string "json" or a JSON schema object.
System prompt for the model to generate a response from
When true, returns a stream of partial responses
Controls a model's thinking output. Use /api/show to discover the supported values and default for the selected model. true requests thinking, false requests no thinking output, and null uses the model default. String values are model-defined; supported names must match /api/show exactly. Numbers are not supported.
When true, returns the raw response from the model without any prompt templating
Model keep-alive duration (for example 5m or 0 to unload immediately)
Runtime options that control text generation
Whether to return log probabilities of the output tokens
Number of most likely tokens to return at each token position when logprobs are enabled
Response
Generation responses
Model name
ISO 8601 timestamp of response creation
The model's generated text response
The model's generated thinking output
Indicates whether generation has finished
Reason the generation stopped
Time spent generating the response in nanoseconds
Time spent loading the model in nanoseconds
Number of input tokens in the prompt
Number of prompt tokens read from the cache
Time spent evaluating uncached prompt tokens in nanoseconds
Number of output tokens generated in the response
Time spent generating tokens in nanoseconds
Log probability information for the generated tokens when logprobs are enabled

