Skip to main content
For failures such as pods stuck Pending, object-store access denied, or metastore connection errors, see Troubleshoot SmithDB.
SmithDB adds three infrastructure dependencies to a self-hosted LangSmith deployment: a PostgreSQL metastore, object storage, and cache storage for query, ingestion, and compaction worker. These are integration requirements and practical recommendations, not a prescribed cloud architecture. For AWS (EKS), GCP (GKE), and Azure (AKS) setup, see Provider setup. The minimum LangSmith version depends on your cloud. See Cloud support.

Requirements

Before enabling SmithDB, provide:
  • A dedicated, empty PostgreSQL database for the SmithDB metastore. Do not use the PostgreSQL database that stores the rest of LangSmith operational data.
  • A dedicated object-storage bucket for SmithDB durable data.
  • Cache storage for query, ingestion, and compaction worker: network-attached disks, or local SSD for the best cache performance.
  • Private network connectivity to the database and object store, plus credentials or workload identity for both.
  • If LangSmith uses an HTTP proxy, NO_PROXY entries for the IP range assigned to SmithDB pods and the cluster’s internal service domain.

PostgreSQL metastore

The metastore holds SmithDB catalog and coordination data. Create an empty database for the metastore. SmithDB initializes the schema during installation. Use PostgreSQL 18 or later and allow database connections from the Kubernetes cluster. For database options and connectivity, see Provider setup.

Metastore Secret

Create a Kubernetes Secret in the LangSmith release namespace containing the database host, name, username, and password. Map its keys through smithdb.config.metastore.
Configure the chart to use the corresponding keys:
DB_NAME can be any name for the dedicated, empty database, such as smithdb. The chart does not require these exact Secret key names; the Helm values map whatever names you choose.

Object storage

Object storage is SmithDB’s durable data layer. Use a bucket reserved for SmithDB data, in the same region as the Kubernetes cluster, to minimize latency and transfer costs. This bucket is separate from optional LangSmith blob storage, which stores payloads and attachments for the broader LangSmith deployment. Do not add bucket lifecycle rules that expire objects on their own schedule. Deleting live objects can make data unavailable.

Use private object-storage connectivity

Use private connectivity to avoid unnecessary data-transfer and NAT gateway costs. Configure the endpoint for your cloud under Provider setup. Configure access so SmithDB components can list the bucket and read, write, and delete objects.

Select a ServiceAccount

SmithDB workloads share smithdb.serviceAccount.
  • Default: The chart creates <HELM_RELEASE>-smithdb.
  • Custom: Set name to create a differently named account.
  • Existing: Set create: false and name, then configure workload identity externally.
Workload identity must target the selected namespace and name. With create: false, name is required. Otherwise pods use the default ServiceAccount.

Migration source-bucket access

If your installation uses LangSmith blob storage, grant smithdb.serviceAccount read access to that bucket before migrating historical data.

Cache storage

Query, ingestion, and compaction worker cache trace data on a volume mounted at /data. Object storage retains the durable copy; replacing a pod discards its cache.
smithdb.cache requires Helm chart 0.17.0 or later. The LangSmith 0.16 tabs show the equivalent configuration for earlier charts.
  • Network-attached disk: The default on LangSmith 0.17. Kubernetes provisions a volume per pod from a StorageClass, so no dedicated node pool is needed.
  • Local SSD: The default on LangSmith 0.16 and an explicit override on 0.17. Recommended for production, where it gives the best cache performance. Requires a node pool whose local disks back Kubernetes ephemeral storage.
Upgrading to 0.17 moves the default cache from emptyDir to a per-pod PersistentVolumeClaim. To stay on local SSD, add the 0.17 local SSD values to the same upgrade. If you carry custom cache volumes forward, rename them from local-ssd-storage to cache.

Network-attached disk

The chart requests a generic ephemeral volume for each pod. Kubernetes creates the PersistentVolumeClaim with the pod and deletes it with the pod. Give the StorageClass reclaimPolicy: Delete so the backing disk goes with it. Cluster default StorageClasses are usually too slow for the cache. Provision at least 7000 IOPS and 1000 MiB/s per volume, using the StorageClass for your provider under Provider setup. The examples below use the small tier. Substitute the sizes for your tier.
Set storageClassName, or leave it empty to use the cluster default. The tier sets the volume size, CPU, and memory. The chart sets fsGroup: 1001 and the query disk cache limit.
To give one component a different StorageClass or size, set its deployment.volumes to a generic ephemeral volume named cache with the class and size you want. Repeat for each component you change. Do not point volumes at an existing claim through persistentVolumeClaim.claimName; each replica needs its own volume.
No node selectors or tolerations are needed; SmithDB runs on your general node pool. The migration Job is the exception. It does not use the cache volume and still requests 100 GiB of node ephemeral-storage by default, so schedule it on nodes with that much allocatable. Confirm each pod has a bound claim and that /data uses it:

Local SSD

Configure the node’s local SSDs to back Kubernetes ephemeral storage, with usable capacity reported as allocatable ephemeral-storage. SmithDB uses this storage for its emptyDir cache volumes. Use scheduling controls to keep SmithDB cache workloads on SSD-backed nodes. Size nodes with headroom above pod requests, images, logs, and Kubernetes reservations. The examples below use the small tier. Substitute the sizes for your tier.
On each disk-using component, set an emptyDir named cache and matching ephemeral-storage requests and limits. The chart derives the query disk cache limit from that limit.

Schedule SmithDB workloads

Place query, ingestion, compaction worker, and migration on the local SSD pool. Place compaction and cluster manager on the general compute pool. Node-pool labels and taints must match the Helm selectors and tolerations. metastoreMigration can run on any node and does not need a selector. The label and taint values in these examples match the provider samples on this page. SmithDB does not require these exact values; any labels and taints that keep cache workloads on SSD-backed nodes work.
On LangSmith 0.16, the migration Job’s pod settings live under smithdb.migration.deployment rather than smithdb.migration.job. All other keys are the same.
Confirm the intended labels, taints, scheduler-visible capacity, pod placement, and cache filesystem:
For pods stuck Pending, slow cache I/O, or evictions, see Troubleshoot SmithDB.

Provider setup

A common AWS mapping is EKS, RDS for PostgreSQL, S3, IRSA, and EC2 instance store or EBS gp3 for the cache.

Configure the metastore

Use RDS for PostgreSQL or Aurora PostgreSQL that meets the metastore requirements.

Configure S3 access

Add an S3 Gateway VPC endpoint to the cluster’s private route tables so bucket traffic stays off the public internet.Choose IRSA or EKS Pod Identity. IRSA requires the role annotation and trust for system:serviceaccount:<NAMESPACE>:<HELM_RELEASE>-smithdb. Pod Identity uses an external association and no annotation.Scope the role’s S3 access to listing the bucket and reading its location, plus object read, write, delete, and multipart operations. For migration source reads, also grant s3:ListBucket and s3:GetObject on the LangSmith blob-storage bucket.

Create a cache StorageClass

For the network-attached disk option, create a gp3 StorageClass with provisioned IOPS and throughput. The EBS CSI driver must be installed.

Provision EKS nodes with Karpenter

The node examples that follow apply to the local SSD option. Karpenter is the recommended way to provision SmithDB capacity on EKS, but it is not required. Other node provisioners must produce the same labels, taints, and Kubernetes-visible ephemeral-storage capacity.Install Karpenter v1 and its CRDs by following the Karpenter EKS guide. Before applying the example below:
  • Replace CLUSTER_NAME and KarpenterNodeRole-CLUSTER_NAME.
  • Tag the selected subnets and security group with karpenter.sh/discovery: CLUSTER_NAME, or replace the selectors with tags or IDs used by your environment.
  • Confirm the node IAM role and EKS access entry are configured for Karpenter-provisioned nodes.
Applying these manifests creates provisioning configuration. EC2 nodes launch when matching SmithDB pods require capacity.
In this example, instanceStorePolicy: RAID0 makes local NVMe available as node ephemeral storage. The smithdb-instance-store NodePool requires at least 800 GiB and consolidates only when empty, avoiding unnecessary churn of nodes with warm caches.That 800 GiB floor suits the medium tier. Size the requirement against the largest per-replica ephemeral-storage request in your chosen tier, plus node headroom. See Configure SmithDB for scale.
With custom AMIs, bootstrap must format and mount instance-store devices for kubelet and container-runtime storage.

Provision EKS managed node groups

If Karpenter is unavailable, use EKS Managed Node Groups instead. These examples configure instance-store NVMe as RAID0 and apply the label and taint used by the Helm scheduling values.
Replace the uppercase placeholders. Choose the instance type and capacity from your sizing baseline and regional availability. Match AMI_TYPE to the instance architecture. For example, i8g.4xlarge uses AL2023_ARM_64_STANDARD.
AWS CLI requires an EC2 launch template for the nodeadm configuration.