Generate a chat message
Generate the next chat message in a conversation between a user and an assistant.
Body
Model name
Chat history as an array of message objects (each with a role and content)
Optional list of function tools the model may call during the chat
Format to return a response in. Can be json or a JSON schema
json Runtime options that control text generation
Controls a model's thinking output. Use /api/show to discover the supported values and default for the selected model. true requests thinking, false requests no thinking output, and null uses the model default. String values are model-defined; supported names must match /api/show exactly. Numbers are not supported.
Model keep-alive duration (for example 5m or 0 to unload immediately)
Whether to return log probabilities of the output tokens
Number of most likely tokens to return at each token position when logprobs are enabled
Response
Chat response
Model name used to generate this message
Timestamp of response creation (ISO 8601)
Indicates whether the chat response has finished
Reason the response finished
Total time spent generating in nanoseconds
Time spent loading the model in nanoseconds
Number of tokens in the prompt
Number of prompt tokens read from the cache
Time spent evaluating uncached prompt tokens in nanoseconds
Number of tokens generated in the response
Time spent generating tokens in nanoseconds
Log probability information for the generated tokens when logprobs are enabled

