> ## Documentation Index
> Fetch the complete documentation index at: https://docs.ollama.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Anthropic compatibility

Connect Anthropic clients and tools such as Claude Code to Ollama. Ollama supports a subset of the [Anthropic Messages API](https://platform.claude.com/docs/en/api/http/messages/create).

## Direct cloud access

Set your [API key](https://ollama.com/settings/keys) in `OLLAMA_API_KEY`. No Ollama installation required.

```shell theme={"system"}
export OLLAMA_API_KEY="your_api_key"
```

Send a request with bearer authentication:

```shell theme={"system"}
curl https://ollama.com/v1/messages \
  -H "Authorization: Bearer $OLLAMA_API_KEY" \
  -H "Content-Type: application/json" \
  -H "anthropic-version: 2023-06-01" \
  -d '{
    "model": "gemma4:31b",
    "max_tokens": 1024,
    "messages": [
      {
        "role": "user",
        "content": "Say hello in one sentence."
      }
    ]
  }'
```

Read text blocks from `content`. The cloud endpoint requires `Authorization: Bearer`; it does not accept `x-api-key` alone.

Compatible models support messages, streaming, and function calling. Tool-choice controls, deferred tools, and hosted web search are not fully supported.

## Local server usage

These examples connect to your local Ollama server. Use `ollama` as the placeholder API key. To use cloud models through this server, [sign in to Ollama](/api/authentication#signing-in).

### Environment variables

To use Ollama with tools that expect the Anthropic API (like Claude Code), set these environment variables:

```shell theme={"system"}
export ANTHROPIC_AUTH_TOKEN=ollama  # required but ignored
export ANTHROPIC_BASE_URL=http://localhost:11434
```

### Simple `/v1/messages` example

<CodeGroup dropdown>
  ```python basic.py theme={"system"}
  import anthropic

  client = anthropic.Anthropic(
      base_url='http://localhost:11434',
      api_key='ollama',  # required but ignored
  )

  message = client.messages.create(
      model='qwen3-coder',
      max_tokens=1024,
      messages=[
          {'role': 'user', 'content': 'Hello, how are you?'}
      ]
  )
  print(message.content[0].text)
  ```

  ```javascript basic.js theme={"system"}
  import Anthropic from "@anthropic-ai/sdk";

  const anthropic = new Anthropic({
    baseURL: "http://localhost:11434",
    apiKey: "ollama", // required but ignored
  });

  const message = await anthropic.messages.create({
    model: "qwen3-coder",
    max_tokens: 1024,
    messages: [{ role: "user", content: "Hello, how are you?" }],
  });

  console.log(message.content[0].text);
  ```

  ```shell basic.sh theme={"system"}
  curl -X POST http://localhost:11434/v1/messages \
  -H "Content-Type: application/json" \
  -H "x-api-key: ollama" \
  -H "anthropic-version: 2023-06-01" \
  -d '{
    "model": "qwen3-coder",
    "max_tokens": 1024,
    "messages": [{ "role": "user", "content": "Hello, how are you?" }]
  }'
  ```
</CodeGroup>

### Streaming example

<CodeGroup dropdown>
  ```python streaming.py theme={"system"}
  import anthropic

  client = anthropic.Anthropic(
      base_url='http://localhost:11434',
      api_key='ollama',
  )

  with client.messages.stream(
      model='qwen3-coder',
      max_tokens=1024,
      messages=[{'role': 'user', 'content': 'Count from 1 to 10'}]
  ) as stream:
      for text in stream.text_stream:
          print(text, end='', flush=True)
  ```

  ```javascript streaming.js theme={"system"}
  import Anthropic from "@anthropic-ai/sdk";

  const anthropic = new Anthropic({
    baseURL: "http://localhost:11434",
    apiKey: "ollama",
  });

  const stream = await anthropic.messages.stream({
    model: "qwen3-coder",
    max_tokens: 1024,
    messages: [{ role: "user", content: "Count from 1 to 10" }],
  });

  for await (const event of stream) {
    if (
      event.type === "content_block_delta" &&
      event.delta.type === "text_delta"
    ) {
      process.stdout.write(event.delta.text);
    }
  }
  ```

  ```shell streaming.sh theme={"system"}
  curl -X POST http://localhost:11434/v1/messages \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen3-coder",
    "max_tokens": 1024,
    "stream": true,
    "messages": [{ "role": "user", "content": "Count from 1 to 10" }]
  }'
  ```
</CodeGroup>

### Tool calling example

<CodeGroup dropdown>
  ```python tools.py theme={"system"}
  import anthropic

  client = anthropic.Anthropic(
      base_url='http://localhost:11434',
      api_key='ollama',
  )

  message = client.messages.create(
      model='qwen3-coder',
      max_tokens=1024,
      tools=[
          {
              'name': 'get_weather',
              'description': 'Get the current weather in a location',
              'input_schema': {
                  'type': 'object',
                  'properties': {
                      'location': {
                          'type': 'string',
                          'description': 'The city and state, e.g. San Francisco, CA'
                      }
                  },
                  'required': ['location']
              }
          }
      ],
      messages=[{'role': 'user', 'content': "What's the weather in San Francisco?"}]
  )

  for block in message.content:
      if block.type == 'tool_use':
          print(f'Tool: {block.name}')
          print(f'Input: {block.input}')
  ```

  ```javascript tools.js theme={"system"}
  import Anthropic from "@anthropic-ai/sdk";

  const anthropic = new Anthropic({
    baseURL: "http://localhost:11434",
    apiKey: "ollama",
  });

  const message = await anthropic.messages.create({
    model: "qwen3-coder",
    max_tokens: 1024,
    tools: [
      {
        name: "get_weather",
        description: "Get the current weather in a location",
        input_schema: {
          type: "object",
          properties: {
            location: {
              type: "string",
              description: "The city and state, e.g. San Francisco, CA",
            },
          },
          required: ["location"],
        },
      },
    ],
    messages: [{ role: "user", content: "What's the weather in San Francisco?" }],
  });

  for (const block of message.content) {
    if (block.type === "tool_use") {
      console.log("Tool:", block.name);
      console.log("Input:", block.input);
    }
  }
  ```

  ```shell tools.sh theme={"system"}
  curl -X POST http://localhost:11434/v1/messages \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen3-coder",
    "max_tokens": 1024,
    "tools": [
      {
        "name": "get_weather",
        "description": "Get the current weather in a location",
        "input_schema": {
          "type": "object",
          "properties": {
            "location": {
              "type": "string",
              "description": "The city and state"
            }
          },
          "required": ["location"]
        }
      }
    ],
    "messages": [{ "role": "user", "content": "What is the weather in San Francisco?" }]
  }'
  ```
</CodeGroup>

## Using with Claude Code

See the [Claude Code guide](/integrations/claude-code) for Ollama Launch and [direct cloud setup](/integrations/claude-code#connect-directly-to-ollama-cloud).

## Endpoints

### `/v1/messages`

#### Supported features

* [x] Messages
* [x] Streaming
* [x] System prompts
* [x] Multi-turn conversations
* [x] Vision (images)
* [x] Tools (function calling)
* [x] Tool results
* [x] Thinking/extended thinking

#### Supported request fields

* [x] `model`
* [x] `max_tokens`
* [x] `messages`
  * [x] Text `content`
  * [x] Image `content` (base64)
  * [x] Array of content blocks
  * [x] `tool_use` blocks
  * [x] `tool_result` blocks
  * [x] `thinking` blocks
* [x] `system` (string or array)
* [x] `stream`
* [x] `temperature`
* [x] `top_p`
* [x] `top_k`
* [x] `stop_sequences`
* [x] `tools`
* [x] `thinking`
* [x] `output_config`
  * [x] `effort` (model-defined names)
* [ ] `tool_choice`
* [ ] `metadata`

#### Supported response fields

* [x] `id`
* [x] `type`
* [x] `role`
* [x] `model`
* [x] `content` (text, tool\_use, thinking blocks)
* [x] `stop_reason` (end\_turn, max\_tokens, tool\_use)
* [x] `usage` (input\_tokens, output\_tokens)

#### Streaming events

* [x] `message_start`
* [x] `content_block_start`
* [x] `content_block_delta` (text\_delta, input\_json\_delta, thinking\_delta)
* [x] `content_block_stop`
* [x] `message_delta`
* [x] `message_stop`
* [x] `ping`
* [x] `error`

Use `/api/show` to discover a model's supported thinking values and default. `thinking.type: "enabled"` and `"disabled"` act as explicit on/off controls. When a model uses its metadata to resolve named levels, supported `output_config.effort` names are applied exactly and unsupported names resolve to the model default. Models without metadata keep the existing compatibility behavior, which trims and lowercases effort names and maps `"xhigh"` to `"high"`.

## Models

Ollama supports both local and cloud models.

### Local models

Pull a local model before use:

```shell theme={"system"}
ollama pull qwen3-coder
```

Recommended local models:

* `qwen3-coder` - Excellent for coding tasks
* `gpt-oss:20b` - Strong general-purpose model

### Cloud models

Browse the [cloud model catalog](https://ollama.com/search?c=cloud). Through a signed-in Ollama server, use a cloud name such as `gemma4:cloud` without a separate pull. For direct cloud requests, use the hosted identifier, such as `gemma4:31b`.

### Default model names on a local server

For tooling that relies on default Anthropic model names such as `claude-3-5-sonnet`, use `ollama cp` to copy an existing model name:

```shell theme={"system"}
ollama cp qwen3-coder claude-3-5-sonnet
```

Afterwards, this new model name can be specified in the `model` field:

```shell theme={"system"}
curl http://localhost:11434/v1/messages \
    -H "Content-Type: application/json" \
    -d '{
        "model": "claude-3-5-sonnet",
        "max_tokens": 1024,
        "messages": [
            {
                "role": "user",
                "content": "Hello!"
            }
        ]
    }'
```

## Differences from the Anthropic API

### Behavior differences

* The local server does not validate API keys. Direct cloud inference requires a valid bearer token.
* The hosted endpoint rejects `anthropic-version` values older than `2023-06-01`. Use `2023-06-01` in requests.
* Token counts are approximations based on the underlying model's tokenizer

### Not supported

The following Anthropic API features are not currently supported:

| Feature | Description |
| - | - |
| `/v1/messages/count_tokens` | Token counting endpoint |
| `tool_choice` | Forcing specific tool use or disabling tools |
| `metadata` | Request metadata (user\_id) |
| Prompt caching | `cache_control` blocks for caching prefixes |
| Batches API | `/v1/messages/batches` for async batch processing |
| Citations | `citations` content blocks |
| PDF support | `document` content blocks with PDF files |
| Server-sent errors | `error` events during streaming (errors return HTTP status) |

### Partial support

| Feature | Status |
| - | - |
| Image content | Base64 images supported; URL images not supported |
| Extended thinking | Basic support; `budget_tokens` accepted but not enforced |
