- ChatLiteLLM: The main LangChain chat wrapper for LiteLLM.
- ChatLiteLLMRouter: A
ChatLiteLLMwrapper that leverages LiteLLM’s Router for load balancing and fallbacks.
Overview
Integration details
Model features
Setup
To accessChatLiteLLM and ChatLiteLLMRouter models, you’ll need to install the langchain-litellm package and create an OpenAI, Anthropic, Azure, Replicate, OpenRouter, Hugging Face, Together AI, or Cohere account. Then, you have to get an API key and export it as an environment variable.
Credentials
You have to choose the LLM provider you want and sign up with them to get their API key.Example - Anthropic
Head to the Claude console to sign up and generate a Claude API key. Once you’ve done this set theANTHROPIC_API_KEY environment variable:
Example - OpenAI
Head to platform.openai.com/api-keys to sign up for OpenAI and generate an API key. Once you’ve done this, set the OPENAI_API_KEY environment variable.Installation
The LangChain LiteLLM integration is available in thelangchain-litellm package:
Instantiation
ChatLiteLLM
You can instantiate aChatLiteLLM model by providing a model name supported by LiteLLM.
ChatLiteLLMRouter
You can also leverage LiteLLM’s routing capabilities by defining your model list as specified in the LiteLLM routing documentation.Invocation
Whether you’ve instantiated aChatLiteLLM or a ChatLiteLLMRouter, you can now use the ChatModel through LangChain’s API.
Async and streaming functionality
ChatLiteLLM and ChatLiteLLMRouter also support async and streaming functionality:
Advanced features
Gemini Enterprise Agent Platform grounding (Google Search)
Use Google Search grounding with Gemini Enterprise Agent Platform models (e.g.,gemini-3.6-flash). Grounding metadata is returned in response_metadata, for both batch and streaming calls.
Reading streamed grounding metadata from
response_metadata requires langchain-litellm>=0.11.0. Before that, ChatLiteLLM returns it in additional_kwargs, and ChatLiteLLMRouter does not return it when streaming.Responses API
use_responses_api requires langchain-litellm>=0.10.0.use_responses_api=True to send ChatLiteLLM calls to the provider’s Responses API instead of Chat Completions. LiteLLM translates each request and reply, so messages, tools, streaming, and structured output work as usual. OpenAI’s built-in tools, such as web search, need this route:
- A model LiteLLM cannot send to a Responses API raises
ValueErrorbefore any request goes out. - LiteLLM drops Chat Completions-only parameters on this route, such as
stop,n, andseed. - When
use_responses_apiisNone(the default) orFalse, LiteLLM picks the API itself, and it already sends some models, such asgpt-5-pro, to the Responses API.
Reasoning items
Keeping reasoning items requires
langchain-litellm>=0.11.0."store": False. OpenAI also accepts an explicit include for that content:
additional_kwargs["reasoning_items"], and only the items that carry encrypted content. A turn the model answers without reasoning has none. Each kept item carries an origin key that names the request that issued it, so keep that key when you store a conversation.
A later request sends a turn’s item back when the turn holds one item and the request is configured with the same model, base URL, and credentials as the request that issued it. A request that does not match sends no items and loses only that turn’s reasoning. When the request could reach another endpoint, such as through a fallback, LiteLLM’s response cache, or litellm.use_litellm_proxy, it neither sends nor keeps items.
OpenAI rejects an item it cannot decrypt with a 400 invalid_encrypted_content error. The match covers only the settings ChatLiteLLM can read, so an item can still reach an account that cannot decrypt it through a gateway that routes to other accounts, such as a LiteLLM proxy, or through credentials LiteLLM finds on its own, such as OpenAI workload identity.
Some turns send no items back:
- Calls LiteLLM sends to the Responses API on its own, without
use_responses_apior aresponses/model name, such as a call togpt-5-pro. Their replies keep no items. - A streamed turn that holds more than one item. When LiteLLM does not stream a reply, it keeps only the last item before the reply’s text and the last item before its tool calls, so the same turn read with
invokecan send its item back.
ChatLiteLLMRouter, items go back only when no fallback is set and every deployment of the group shares one model, base URL, set of credentials, and set of tool and reasoning settings. Name the deployments <provider>/responses/<model>, as in Router deployments. A deployment LiteLLM sends to the Responses API on its own, such as openai/gpt-5-pro, keeps no items, with or without use_responses_api.
Router deployments
The Router picks a deployment on each call, so there is no single model name foruse_responses_api to reroute. Name each deployment’s model <provider>/responses/<model> to send it to the Responses API:
langchain-litellm 0.11.0, use_responses_api=True on ChatLiteLLMRouter checks the deployments instead of rerouting them. The call goes out unchanged when LiteLLM sends every deployment of the called group to a Responses API, and raises ValueError before any request otherwise, naming each deployment that does not go there. It also raises when the call may reach deployments outside the group, such as through a fallback or a model_group_alias. Earlier versions raise ValueError for the flag on ChatLiteLLMRouter.
API reference
For detailed documentation of allChatLiteLLM and ChatLiteLLMRouter features and configurations, see the langchain-litellm API reference.
Connect these docs to your agent of choice via MCP for real-time answers.

