Skip to main content
Ollama allows you to run open-weight large language models (LLMs), such as gpt-oss, locally. Ollama bundles model weights, configuration, and data into a single package, defined by a Modelfile. It optimizes setup and configuration details, including GPU usage. For a complete list of supported models and model variants, see the Ollama model library.
API ReferenceFor detailed documentation of all features and configuration options, head to the ChatOllama API reference.

Overview

Integration details

Model features

Setup

Download and install Ollama, then follow the Ollama setup instructions to start a local instance. Pull the model used in the chat and tool-calling examples:
Use the same model tag in ollama pull and ChatOllama(model=...). The image, log probability, and custom-role examples below include separate pull commands for their models.
  • List downloaded models with ollama list.
  • Chat from the terminal with ollama run gpt-oss:20b.
  • Run ollama help for available commands.
Optionally, enable tracing of your model calls with LangSmith by setting the following environment variables:

Installation

The LangChain Ollama integration lives in the langchain-ollama package:

Instantiation

Instantiate the model downloaded during setup:

Invocation

Read the response content blocks:

Tool calling

Use bind_tools() with an Ollama model that supports tool calling, such as gpt-oss:20b from setup. Create a tool with the @tool decorator. For details, see Customize tool properties.

Multimodal

Use an image-capable model, such as Gemma 3. Pull the 4b variant before running this example:

Log probabilities

ChatOllama supports token-level log probabilities via the logprobs and top_logprobs parameters. Log probabilities indicate how likely each token was at each generation step. Pull the model used in these examples:

Basic usage

Top-K alternatives per token

Use top_logprobs to return the most likely alternative tokens at each position:

Reasoning models and custom message roles

Some models, such as IBM’s Granite 3.2, support custom message roles to enable thinking processes. To access Granite 3.2’s thinking features, pass a message with a "control" role with content set to "thinking". Because "control" is a non-standard message role, use a ChatMessage object to implement it:
The model exposes its thought process in addition to its final response.

API reference

For detailed documentation of all ChatOllama features and configuration options, see the API reference.