Skip to main content
LiteLLM is a library that simplifies calling Anthropic, Azure, Huggingface, Replicate, etc. This page covers how to get started using LangChain with the LiteLLM I/O library. This integration provides two chat model classes:
  • ChatLiteLLM: The main LangChain chat wrapper for LiteLLM.
  • ChatLiteLLMRouter: A ChatLiteLLM wrapper that leverages LiteLLM’s Router for load balancing and fallbacks.
The package also ships LiteLLMEmbeddings, LiteLLMEmbeddingsRouter, and LiteLLMOCRLoader. See the providers page for details.

Overview

Integration details

Model features

Setup

To access ChatLiteLLM and ChatLiteLLMRouter models, you’ll need to install the langchain-litellm package and create an OpenAI, Anthropic, Azure, Replicate, OpenRouter, Hugging Face, Together AI, or Cohere account. Then, you have to get an API key and export it as an environment variable.

Credentials

You have to choose the LLM provider you want and sign up with them to get their API key.

Example - Anthropic

Head to the Claude console to sign up and generate a Claude API key. Once you’ve done this set the ANTHROPIC_API_KEY environment variable:

Example - OpenAI

Head to platform.openai.com/api-keys to sign up for OpenAI and generate an API key. Once you’ve done this, set the OPENAI_API_KEY environment variable.

Installation

The LangChain LiteLLM integration is available in the langchain-litellm package:

Instantiation

ChatLiteLLM

You can instantiate a ChatLiteLLM model by providing a model name supported by LiteLLM.

ChatLiteLLMRouter

You can also leverage LiteLLM’s routing capabilities by defining your model list as specified in the LiteLLM routing documentation.

Invocation

Whether you’ve instantiated a ChatLiteLLM or a ChatLiteLLMRouter, you can now use the ChatModel through LangChain’s API.

Async and streaming functionality

ChatLiteLLM and ChatLiteLLMRouter also support async and streaming functionality:

Advanced features

Use Google Search grounding with Gemini Enterprise Agent Platform models (e.g., gemini-3.6-flash). Grounding metadata is returned in response_metadata, for both batch and streaming calls.
Reading streamed grounding metadata from response_metadata requires langchain-litellm>=0.11.0. Before that, ChatLiteLLM returns it in additional_kwargs, and ChatLiteLLMRouter does not return it when streaming.
Setup
Batch usage
Streaming usage

Responses API

use_responses_api requires langchain-litellm>=0.10.0.
Set use_responses_api=True to send ChatLiteLLM calls to the provider’s Responses API instead of Chat Completions. LiteLLM translates each request and reply, so messages, tools, streaming, and structured output work as usual. OpenAI’s built-in tools, such as web search, need this route:
  • A model LiteLLM cannot send to a Responses API raises ValueError before any request goes out.
  • LiteLLM drops Chat Completions-only parameters on this route, such as stop, n, and seed.
  • When use_responses_api is None (the default) or False, LiteLLM picks the API itself, and it already sends some models, such as gpt-5-pro, to the Responses API.

Reasoning items

Keeping reasoning items requires langchain-litellm>=0.11.0.
A reasoning model on the Responses API returns its reasoning as reasoning items. Sending a turn’s item back on the next request lets the model continue its reasoning through a tool loop rather than start again after each tool call. An item goes back only with its encrypted content, which OpenAI returns by default on stateless requests, sent with "store": False. OpenAI also accepts an explicit include for that content:
A reply keeps its reasoning items in additional_kwargs["reasoning_items"], and only the items that carry encrypted content. A turn the model answers without reasoning has none. Each kept item carries an origin key that names the request that issued it, so keep that key when you store a conversation. A later request sends a turn’s item back when the turn holds one item and the request is configured with the same model, base URL, and credentials as the request that issued it. A request that does not match sends no items and loses only that turn’s reasoning. When the request could reach another endpoint, such as through a fallback, LiteLLM’s response cache, or litellm.use_litellm_proxy, it neither sends nor keeps items. OpenAI rejects an item it cannot decrypt with a 400 invalid_encrypted_content error. The match covers only the settings ChatLiteLLM can read, so an item can still reach an account that cannot decrypt it through a gateway that routes to other accounts, such as a LiteLLM proxy, or through credentials LiteLLM finds on its own, such as OpenAI workload identity. Some turns send no items back:
  • Calls LiteLLM sends to the Responses API on its own, without use_responses_api or a responses/ model name, such as a call to gpt-5-pro. Their replies keep no items.
  • A streamed turn that holds more than one item. When LiteLLM does not stream a reply, it keeps only the last item before the reply’s text and the last item before its tool calls, so the same turn read with invoke can send its item back.
On ChatLiteLLMRouter, items go back only when no fallback is set and every deployment of the group shares one model, base URL, set of credentials, and set of tool and reasoning settings. Name the deployments <provider>/responses/<model>, as in Router deployments. A deployment LiteLLM sends to the Responses API on its own, such as openai/gpt-5-pro, keeps no items, with or without use_responses_api.

Router deployments

The Router picks a deployment on each call, so there is no single model name for use_responses_api to reroute. Name each deployment’s model <provider>/responses/<model> to send it to the Responses API:
From langchain-litellm 0.11.0, use_responses_api=True on ChatLiteLLMRouter checks the deployments instead of rerouting them. The call goes out unchanged when LiteLLM sends every deployment of the called group to a Responses API, and raises ValueError before any request otherwise, naming each deployment that does not go there. It also raises when the call may reach deployments outside the group, such as through a fallback or a model_group_alias. Earlier versions raise ValueError for the flag on ChatLiteLLMRouter.

API reference

For detailed documentation of all ChatLiteLLM and ChatLiteLLMRouter features and configurations, see the langchain-litellm API reference.