Subscribe: This changelog includes an RSS feed that can integrate with Slack, email, Discord bots like Readybot or RSS Feeds to Discord Bot, and other subscription tools.
If you use self-hosted LangSmith, see the self-hosted changelog for updates.
- LangSmith Cloud
- LangSmith Fleet
Observability and evaluations
Datasets and experiments
- LangSmith exposes annotation queue item endpoints for adding, listing, updating, deleting, counting, positioning, and reviewing run or thread queue items through the public API and SDK generation flow.
- Uploading a .csv or .jsonl dataset now works regardless of the Content-Type the browser reports. Windows browsers label .csv files as an Excel type, which previously caused valid uploads to fail. Uppercase filenames such as DATASET.CSV are also accepted.
- Evaluator lists on a dataset or tracing project now show an evaluator’s current name instead of the name it had when it was attached. Feedback keys are unchanged by a rename.
- Metadata columns in the experiment comparison grid, including
example.metadata.<key>, now render their values instead of staying empty. - LangSmith marks every legacy endpoint replaced by the SmithDB SDK migration guide as deprecated: the v1 runs query and retrieve endpoints, the v1 run sharing and public-run read endpoints,
POST /api/v1/datasets/{dataset_id}/runs, and the annotation queue run endpoints. All of them now respond withDeprecation: true, aSunsetdate of January 31, 2027, and aLinkheader pointing at the migration guide and, where a single replacement exists, the successor endpoint. Learn more. - Dataset experiment tables can sort by feedback score when SmithDB queries are enabled and ClickHouse queries are disabled.
- Annotation queue item APIs use project_id for the tracing project. Request bodies also accept session_id as an alias.
- Dataset example views restore clear spacing between the example details and tab navigation.
- Pairwise annotation queue runs again include the tracing project id needed to create feedback when ClickHouse query support is disabled.
- Correcting an evaluator score from the experiment results grid now updates the cell and its popover right away instead of requiring a page refresh.
- The Configure Evaluator pane’s header and templates navigation again paint the same background as the pane itself in dark mode.
- Dataset and run attachments now resolve relative signed download URLs before previewing, opening, or downloading them in self-hosted deployments.
Monitoring and alerting
- Adds a reusable ChartCard component to the LangSmith design system, standardizing chart titles, move, expand, and overflow actions, responsive full-width layouts, and chart and legend spacing.
Engine
- A resolved issue on an Engine Issue Board now returns to Open as soon as Engine links a new matching trace to it. Previously the trace was filed as evidence but the issue stayed closed, so a problem that came back never resurfaced on the board. Dismissed issues stay dismissed.
- The issue detail header now renders category and tag badges on the same line as the title instead of stacking them underneath, tightening the header and reducing wasted vertical space.
- The “Engine failed to complete a run” trigger is no longer offered when configuring a Slack channel or webhook destination on an issue board. Destinations already subscribed to it keep receiving those notifications.
- Engine run webhooks and sandbox links now resolve the externally reachable API base from LANGSMITH_PUBLIC_API_ENDPOINT, falling back to LANGCHAIN_PLATFORM_ENDPOINT and then LANGCHAIN_ENDPOINT. Installs whose chart set only LANGSMITH_PUBLIC_API_ENDPOINT were building relative URLs, which made every non-shadow Engine run fail because the run webhook was rejected as a loopback address.
Tracing
- Trace detail panes once again use an elevated background that matches their section headers.
- Bulk exports accept a new opt-in
feedbackscolumn that carries each run’s individual feedback entries, with their key and comment, as a JSON array. Addfeedbackstoexport_fieldswhen creating an export; exports that omitexport_fieldskeep their existing columns. - Delete an entire trace from the run details actions menu after confirming the destructive action.
- Negative feedback-key filters now return matching traces correctly when ClickHouse uses optimized runs tables.
- When a project configures a custom output renderer, the trace Output section now offers it as a Custom option alongside Markdown, Plain, JSON, and YAML instead of replacing them. Custom stays the default, and your choice is remembered.
- Tracing project activity and sorting stay up to date for self-hosted deployments using Redis versions before 6.2.
POST /api/v1/runs/statsnow returns a 404 when the requested tracing project does not exist in your workspace, and reports other client errors, such as astart_timeolder than the supported lookback window, with their real status and message instead of a generic 500 internal server error.- Self-hosted deployments now catch up missing tracing project last-run timestamps so project sorting reflects recent historical activity.
- Run filters now support total, prompt, and completion token counts and costs consistently across routed query backends.
- Token Count / Cost filters now default to total tokens / cost instead of input tokens. Input and output token / cost breakdowns remain available in the selector.
- On a tracing project, resetting a view now moves the Threads/Traces/Runs switcher back in step with the rows being shown. Previously the switcher could stay on Runs while the table had already returned to traces.
Feedback
- The annotation queue item endpoints are now documented at their served path under /api/v1/platform, so the generated SDK methods for listing, adding, updating, counting, deleting, and placing queue items reach the API instead of returning 404.
- PDFs and other documents attached to a message now render in a full-width preview frame with a header control that opens them nearly full screen, instead of collapsing to a thumbnail-sized box.
Prompts and playground
- When a model provider rejects a playground run, such as a wrong API key or an exhausted quota, the playground now shows the provider’s own error message instead of a generic server error, so the cause is clear from the error itself.
- Playground batch and invoke endpoints now sanitize buffered run trees into JSON-safe payloads before responding, so online evaluations no longer fail with opaque 500s when a run graph cannot be serialized.
Automations
- Forking an evaluator attaches the copy to the project or dataset named in the fork dialog. Previously the copy could be created attached to nothing, leaving the original evaluator running the version you had just edited away from.
Deployment
- The Create New Deployment form now shows a clear, per-field message when a submitted value fails validation (e.g. an invalid image path) instead of the raw backend error payload.
- Deleting a tracing project whose LangGraph deployment is still within its post-deletion retention window now schedules the project to be removed along with the deployment, and explains that in the error message instead of asking you to delete a deployment you already deleted.
- The delete confirmation for a LangGraph deployment now says that cleanup of the underlying database runs after you confirm, and that a deleted deployment’s name stays reserved until that cleanup finishes.
Sandboxes
- Sandboxes accept a new streaming execute request that returns stdout and stderr as Server-Sent Events, for clients that cannot hold a WebSocket. Passing a command ID reuses a running command, and a separate resume request continues an interrupted stream instead of running the command again.
- Sandboxes now ship with the
langsmithCLI on PATH, so agents can query traces, runs, and datasets without installing it first. - Sandbox and snapshot rows now use single-column layouts with actions in context menus, snapshot sources remain available in their menus, and wrapped names stay left-aligned.
- Engine sandbox commands in self-hosted deployments now authenticate over WebSocket with the deployment service key, preventing 401 failures after successful sandbox creation.
- The
langsmithCLI in sandboxes is updated to v0.2.44. Its requests now resolve on self-hosted deployments that serve the API under/api, where commands such astrace messagesand the project issues commands previously failed.
Administration
- The LangSmith home page now provides quick access to copy the current organization and workspace IDs.
- On self-hosted installations authenticating with OAuth/SSO, the Remote MCP authorization endpoint returned a 400 because the SSO login route shadowed it. OAuth clients can now complete the authorize step and connect to the Remote MCP server.
- Self-hosted deployments with an online license key (beacon access) now see the monthly organization usage graph automatically, without needing the enable_monthly_usage_charts org config. Offline deployments are now pointed to the Granular usage tab for locally-recorded billable usage.
- Switching workspaces or organizations keeps you on the same page when the route has no workspace-specific resource IDs. Routes that reference a specific resource continue to open the destination workspace home page.
- LangSmith now refreshes expired browser sessions and retries interrupted API requests before asking you to log in again.
LLM Gateway
- Gateway policies with a blank or whitespace-only name now fall back to showing the policy ID instead of rendering an empty name cell, and the policy update endpoint now rejects blank names the same way creation already does.
- The Gateway usage spend chart now shows the top 12 individual spenders per bucket, rolling the rest into an “Other” series, and its hover tooltip lists every contributor at once with the bucket total pinned beneath them.
- The home page onboarding flow now includes a step for choosing how your agent reaches a model: bring your own provider API key, or use Gateway Credits. Choosing Gateway Credits lets an organization admin pre-purchase prepaid credit and shows the API key and gateway URL needed to start sending traffic, with no provider account required.
- The custom X-Gateway-* header condition on LLM Gateway spend-cap and rate-limit policies now works with organization-, workspace-, and user-scoped policies too, not just API-key-scoped ones, so you can split a single subject’s traffic into separate caps by header value regardless of how the policy is scoped.
- Default spend cap and rate limit policies in LLM Gateway now expand to show the per-user, per-workspace, and per-API-key policies materialized underneath them.
- The Gateway Credits purchase dialog now shows the credit balance your purchase will land you on, and states the fee-inclusive total directly above a single purchase button. Large amounts no longer push the credits readout and total outside the dialog.
- Enterprise organizations without Data Protection now see the tab in the LLM Gateway, greyed out with a link to request access, instead of the tab being hidden entirely.
- The connect card now appears above the Gateway Credits balance on the LLM Gateway Home page, and the generated code snippets list the API key before the base URL to match common convention.
- Generating an API key from the LLM Gateway Home connect card now shows the standard one-time key reveal dialog, and each provider’s “Configured” status now reflects the workspace’s actual secrets instead of a fixed list.
- The “Purchase Credits” button on LLM Gateway Home now opens the credit purchase dialog instead of showing a “coming soon” message, and the balance bar now shows spend against what’s actually purchased instead of against the plan’s purchase limit.
- The Cost Controls and Model Fallbacks shortcuts on LLM Gateway Home now read “View cost controls”/“View model fallbacks” for members who can’t manage the org, instead of “Manage”/“Configure”.
- Gateway Home code samples now use the gateway hostname constructed for each LangSmith region and default to the Responses API where supported. Learn more.
- LLM Gateway policy tabs now clarify that policies apply across the organization, while Usage clarifies that spend is scoped to the selected workspace.
- The home onboarding step now states your Gateway Credits balance in US dollars, matching the amount you purchase, instead of converting it to LCUs.
- The prompt you copy into your coding agent during onboarding now states how the agent should reach a model, based on the provider you picked: Gateway Credits, or your own provider API key.
- Homepage spacing and surface tinting now match design review feedback, and several small copy fixes clarify credit limits, provider status, and organization-level purchase limits.
- LLM Gateway Home again includes Google Gemini and generates valid model identifiers for Gemini and Baseten connect samples. Learn more.
- LLM Gateway Home now highlights the selected model and lets you switch connect samples between Chat Completions, Messages, and Responses formats.
- The LLM Gateway Usage tab now explains when usage queries aren’t available for a deployment instead of showing failed dashboard requests.
- When you select Gateway Credits during onboarding, the prompt copied into your coding agent now includes the correct Gateway URL for your deployment. Learn more.
- Gateway Credits checkout now remains on the active purchase step while a free workspace upgrades to the Developer plan, instead of briefly showing the saved-card view before closing.
- Requests that set prompt_cache_options (or the deprecated prompt_cache_retention) now enable Anthropic prompt caching when the LLM Gateway translates an OpenAI Chat Completions or Responses request to a Claude model, instead of ignoring the field. Anthropic’s default cache lifetime applies, and prompt_cache_key, prompt_cache_retention, and prompt_cache_options are all preserved when translating between the Chat Completions and Responses formats.
- Tooltips on disabled LLM Gateway policy controls now read “You need organization admin access to create policies” instead of referencing the raw organization:manage permission string.
- The Gateway Usage tab no longer shows an in-page workspace dropdown. The page is already scoped to your current workspace, and its subtitle now names that workspace directly.
Observability and evaluations
Datasets and experiments
GET /annotation-queues/{id}/itemsnow returns THREAD queue items alongside RUN items, with an optional item_type filter. THREAD list rows expose identity on the item (thread_id/session_id/start_time) and omit nested thread; RUN list still hydrates nested run for now (same omit planned for metadata-only list). Traces and messages load via the v2 threads APIs when reviewing.GET /annotation-queues/{id}/runsand related size endpoints exclude THREAD queue rows (run_id is null) so mixed queues no longer return 500. Use GET /items for thread listing.- Project settings now require a thread idle time of at least two minutes to ensure thread evaluations include all ingested runs.
- Annotation queue item add requests now consistently enforce a 100-item limit for runs and threads. Over-limit errors clearly show the configured limit.
GET /annotation-queues/{id}/itemsreturns Postgres membership metadata for RUN and THREAD items without hydrating nested run or thread payloads from ClickHouse or SmithDB. Orphan queue rows remain listed. The include_stats query parameter is removed. Annotate payloads load via get-one / review APIs.- When you bulk-add runs from a tracing project, its default dataset now appears first in the dataset picker.
- Annotation queue review lists now label unnamed run rows as
Run <ID>, matching thread rows and making mixed queues easier to scan. - Opening the pairwise experiment comparison from a pairwise annotation queue no longer crashes the page when both compared runs come from the same experiment.
Tracing
- The Insights reports pane can now be collapsed to give report details more space.
- Insights cluster summary columns can now expand to show more of each summary.
- LangSmith MCP’s
fetch_runstool now returnsfirst_token_timewhen that value is recorded for fetched runs, making TTFT analysis available without a separate SDK query. - Legacy run URLs now resolve the run metadata and redirect to the SmithDB trace view instead of showing an error.
- Trace usage limit banners now appear only for workspace members whose user-scoped limit has been exceeded.
- The Deployment button on tracing project pages now opens the deployment page within the app instead of reloading the UI.
- Clicking the already-selected run or trace in a thread’s trace tree now keeps it selected instead of deselecting it and scrolling back to the first trace.
- Fixed two frontend call sites that could reach POST /runs/stats with an empty or missing session, preventing spurious 422 errors in the Insights job config and session rules form.
- MCP run query tools now return gateway timeout responses without retrying, reducing duplicate load when a run query times out.
- Pressing Enter to confirm characters from an input method editor (e.g. Pinyin for Chinese, or Japanese/Korean IMEs) in LangSmith Chat now commits the composed text instead of prematurely sending the message.
- LangSmith Chat now surfaces a run’s system prompt inline when reading a traced LLM run, so it can explain why a model behaved a certain way without missing the system prompt. It also recognizes system prompts stored as provider-level fields (OpenAI Responses
instructions, Anthropicsystem). - LangSmith Chat traces now show the model you configured instead of mislabeling it as GPT-3.5-Turbo when a custom model endpoint is used.
- The Run in Studio button is now hidden on public (shared) run pages, where it previously pointed to an authenticated page that shared viewers cannot open.
- Navigating between traces now clears stale sharing state, so unshared traces no longer appear public.
- The run, trace, and thread detail panels now enforce a minimum width when resized, so their header no longer overflows and forces the view to scroll sideways.
- Restore
is_in_dataseton trace and run responses when the query is proxied to the V1 backend in ClickHouse-only mode. The V1 select now forwardsin_datasetso the Python backend computes it and the proxy renames it back tois_in_datasetfor V2 callers. - BYOC workspaces now avoid requesting trace table fields that older data planes do not support, preventing invalid request errors when viewing traces.
- Custom dashboard charts can now query summed latency and first-token time metrics through the runs analytics SmithDB path.
- Custom dashboard charts can now query minimum and maximum latency, time-to-first-token, token, and cost values through SmithDB-backed runs analytics.
- Custom dashboard charts can now query P90 and P95 for latency, first token time, tokens, and cost metrics, in addition to the existing P50 and P99.
- Custom dashboard charts can now aggregate feedback scores by sum, P50, P90, P95, and P99, in addition to the existing average, min, and max.
- The LangSmith homepage now provides clearer onboarding steps for coding agents and tracing, including a direct shortcut for creating a tracing project.
Prompts and playground
- Editing agent and skill metadata in the Context Hub now shows save progress, reports actionable errors without discarding changes, and displays successful updates immediately.
- Prompts with a repo readme now display it in a dedicated Readme section of the prompt view.
- Gemini 3.6 Flash and Gemini 3.5 Flash Lite are now available in Fleet, Agent Builder, and playground model selectors, with usage pricing support.
- Pasting content into rich-text editors, such as the prompt playground and agent chat, now works reliably again.
- Previewing a single dataset row after running a full experiment now resolves the row’s evaluator scores instead of showing a feedback cell that loads indefinitely, and the resulting feedback chips render with their proper colors.
- LangSmith no longer keeps system-added top_p values when switching OpenAI prompts to reasoning models, preventing invalid invocation parameters. Users who still need top_p can add it as an extra model parameter for supported non-reasoning models.
Engine
- The organization Engine usage page now lets you switch between a workspaces view and a projects view of month-to-date LCU spend, each ranked by spend. The Engine spend API returns the authoritative total independently of the breakdown.
- Engine no longer shows redundant hover tooltips on issue category badges or the default Fix action. Permission and PR status explanations still appear when they add context.
- Engine opens the Slack or webhook destination form immediately when no destinations are configured.
- Engine issue category labels now appear on a dedicated row below the title for more consistent spacing and readability.
- Engine now shows a warning beside a linked repository when it cannot access it, including failures caused by renamed or deleted repositories and broken GitHub connections.
- On an Engine issue, navigating to the next or previous linked trace now stays in the conversation view for traces that belong to a thread, instead of switching to the single-trace view.
- Engine now shows a clickable Paused status in the issues header when scheduled scanning is paused, so you can see its status and open settings directly.
- Engine now works in supported self-hosted deployments without Eppo rollout configuration, while organization enablement and existing permissions remain enforced.
- Opening the project spend limit from the pause confirmation now scrolls the settings pane to the limit editor.
- Engine Overview now displays the current Engine package version so you can see when the underlying experience changes.
Monitoring and alerting
- Self-hosted alert webhook delivery now honors
SSRF_ALLOW_K8S_INTERNAL, so internal Kubernetes service hostnames can be used when that setting is enabled. Metadata endpoints, localhost, and private IP protections remain controlled by their existing SSRF policy settings.
Automations
- Thread (grouped) evaluators now require a minimum idle time of 120 seconds. Setting a project’s thread idle time below 120s is rejected.
- Leaving feedback on a run in a thread now makes the thread eligible for re-evaluation, so thread-level evaluators re-run when new feedback arrives.
- Editing an online evaluator (LLM-as-judge or custom code) no longer intermittently fails when sandbox validation is slow.
Deployment
- Worker and API server CPU charts now plot a peak (max) series alongside the average, so a single replica with high CPU usage is no longer hidden by the fleet-wide average.
- Worker and API server memory charts now plot a peak (max) series alongside the average, so a single replica approaching its memory limit is no longer hidden by the fleet-wide average.
- Creating a deployment with a name that’s already in use within your workspace now returns a 409 Conflict instead of a 500 error.
- Hybrid deployments remain compatible with older listeners during control-plane upgrades, preventing new deployments from remaining queued.
Sandboxes
- A sandbox proxy configuration can now define environment variables that are applied to every command in the sandbox. This is handy when a tool refuses to run unless a credential env var is present (for example gh needs GH_TOKEN) even though the egress proxy injects the real credential on the wire. Set a placeholder value so the command starts.
- Attach free-form key/value labels when creating a sandbox or snapshot. Labels are stored and returned on reads; sandboxes inherit their snapshot’s labels, and snapshots built from a Docker image inherit the image’s labels.
- Sandbox network egress now tries every resolved IP for a destination instead of only the first, so requests to multi-homed hosts (for example apt package mirrors) no longer fail when the first address is unreachable on the requested port.
- Creating a sandbox with only mem_bytes set now derives a matching CPU allocation automatically, so requests for larger-memory sandboxes no longer need an explicit vcpus value to satisfy the CPU-to-memory ratio.
Administration
- Selecting All Workspaces on the Granular Billable Usage page now loads usage successfully for organizations with many workspaces.
- The Granular Usage page now shows a notice that long-lived trace usage isn’t tracked in self-hosted deployments, so the “Long-lived only” filter is expected to show zero results there.
LLM Gateway
- Gateway Monitoring now shows spend for the workspace you’re viewing rather than the whole organization, with a dropdown to switch workspaces from the page. Spend cards show N/A instead of a repeated error message when a workspace’s gateway project can’t be resolved.
- The Rate Limiting tab in Gateway Policies now supports creating, editing, deleting, and enabling/disabling request- and token-based rate-limit policies, alongside the existing cost-control and data-protection policy management.
- Editing a materialized LLM Gateway policy now turns it into a standalone override, preventing later default policy changes from overwriting its custom limits.
- You can now call LangChain-managed models through the LLM gateway without configuring your own provider credentials. Usage is metered at cost and bounded by a monthly spend cap based on your plan; once the cap is reached, further requests are blocked until the next month. The cap can be raised on request. Learn more.
- Selecting an entity filter on the LLM Gateway spend monitoring page no longer flips the breakdown to a different dimension.
- Selecting more than one entity in any Gateway Monitoring breakdown filter (model, user, or API key) now returns spend for all chosen entities instead of no data.
- The Gateway Monitoring spend chart now formats axis labels and tooltip ranges in UTC to match its UTC-anchored buckets, so viewers in non-UTC timezones no longer see off-by-one dates.
- The LLM Gateway spend chart and table now label spend from service keys as “Unaffiliated with any user” instead of showing a blank name, with a tooltip explaining that the spend came from a workspace/org-scoped key rather than an individual.
- API-key-scoped LLM Gateway spend-cap and rate-limit policies can now add a custom X-Gateway-* header condition, so a single API key can match different limits per header value. For example, a reseller can set separate caps per downstream customer without distributing multiple keys.
- Stat card and table headers in the LLM Gateway monitoring page’s Spend tab (e.g. “Total Spend”, “Daily Avg”, “API Key”) now capitalize every word, matching the header style used elsewhere in the product.
- When a specific start/end date is selected in the LLM Gateway monitoring page’s date range picker, the button now shows the dates in UTC and appends “(UTC)” so it’s clear the range doesn’t follow your local timezone. Relative ranges like “Last 7 days” are unaffected.
- The “Spend share” column on the LLM Gateway Monitoring spend dashboard no longer cuts off its header text.
- The LLM Gateway now lives in a dedicated top-level sidebar section instead of under Settings, with a new Home tab listing your custom model configurations and a ready-to-run code snippet for the gateway. Old Settings gateway links redirect automatically.
- A Home banner for LangSmith Cloud orgs with LLM Gateway enabled highlights how Gateway manages costs and improves runtime reliability. Learn more.
Observability and evaluations
Datasets and experiments
- The legacy feedback formula endpoints (
POST/GET /feedback/formulasandGET/PUT/DELETE /feedback/formulas/{feedback_formula_id}) that back composite scores are deprecated in favor of composite evaluators, which implement a composite score as a code evaluator plus a run rule, and are scheduled for removal on 2026-08-20. Migrate existing feedback formulas to the new composite model. - Model, prompt, and tool chips in the Experiments table config cells now lay out from real measurements for accurate truncation, and the +N overflow badge is a clickable dropdown whose entries expose the same actions (filter, group by, open in playground, and details) as a chip’s own menu.
- Expanding the run tree for repetition runs in experiment comparison views now works reliably when a repetition root has a project ID but no session ID.
- Evaluators linked to Hub prompts now load correctly for flat and playground-shaped prompt commits, fixing crashes when editing existing evaluators.
- Code evaluator upload now accepts Python entrypoints annotated with PEP 604 union return types (for example
-> dict | None). - POST /v2/datasets//experiment-runs is the supported public API for paginated experiment comparison. Legacy dataset comparison helpers are removed from the public OpenAPI spec and generated SDKs; existing HTTP routes continue to work for LangSmith UI clients.
- Each example’s dataset splits now render as chips in the dataset Examples table, laid out from real measurements with a clickable +N overflow menu when an example belongs to more splits than fit the column.
- Adds
langsmith evaluator create-llmto define structured LLM-as-judge evaluator rules from a prompt, schema, and model config file, targeting a project or dataset. - The experiment comparison view now offers an optional, reorderable “Splits (latest)” column that shows each example’s current dataset split assignments as chips, reflecting live membership rather than the as-of-run snapshot.
- Evaluator spend charts on project and dataset evaluator tabs keep their desktop layout on narrow screens and scroll horizontally instead of compressing the chart and stat cards.
- The experiment comparison and group-by views now show each example’s current dataset split rather than the split it had when the experiment ran, so you can tell whether failures already belong to a split without re-running the experiment.
- Comparison view now loads token and cost stats from SmithDB for root runs, so the stats columns populate again instead of staying blank
- LangSmith now caps reusable evaluators per workspace to prevent unbounded resource growth. Contact support if your workspace needs a higher limit.
- Creating dataset examples from source runs now correctly fetches run inputs and outputs backed by SmithDB, and no longer fails the whole request if one of several source runs can’t be found.
- Select multiple rows in an experiment (or select all matching the current filters) and add, replace, or remove their dataset splits in one action, or copy the selected examples to another dataset, instead of editing rows one at a time.
- The
/runs/rules/validateendpoint now supports thread evaluators. Passtest_thread_idandsession_idto test a multi-turn evaluator against a real conversation before saving. - Custom code evaluators that time out or fail on a run now record an error on that run instead of silently leaving it without feedback, so partial evaluation failures are visible on the experiment.
- The Open source run action on an example page now reads session and start time from dedicated example fields populated at creation, enabling reliable navigation to the source trace on SmithDB.
- The thread evaluator config preview now shows the thread message formats the evaluator actually maps, instead of listing every available format.
- Multi-turn evaluators now include a Test action that runs the evaluator against a sample thread before you save the rule.
- The evaluator config now shows a locked “Trace count ≥ 2” filter for managed thread evaluators, making it clear they only run on threads with multiple turns.
- Experiment comparison and individual experiment views now load run rows on self-hosted deployments that authenticate the UI via SSO/OAuth session cookies. Previously these views could show ‘No results found’ even though metrics and feedback loaded.
- Experiment statistics now refresh promptly for recently run experiments while keeping historical experiment scans bounded.
- The Assertions evaluator added via “Add evaluator” now reads assertions from the reference output like the auto-attached version, so it grades against the real assertions instead of always failing.
- Evaluator spend chart y-axes now abbreviate amounts of $1,000 or more, making high-spend values easier to scan.
- Exporting a dataset comparison view as CSV now returns a clear “file is too large to export” error instead of a generic server error when the export exceeds internal size limits.
- Each split chip in a row’s Splits cell is now interactive in the experiment results and comparison views, with an Edit splits action that opens the single-example split picker so you can reassign splits without leaving the table.
- Add RUN items to a single annotation queue with POST /annotation-queues//items. The server resolves runs via ClickHouse or SmithDB and returns a standards-shaped items envelope; THREAD support follows in a later release.
- The LangSmith CLI now updates existing code evaluator rules in place when
evaluator upload --replaceis used, avoiding a delete-before-create window if the replacement upload fails. - Split the read datasets into a new download datasets permission. Enforce this new permission in both the application and in APIs. The download button is disabled for those users without the download permission. Learn more.
- Public dataset experiment traces open correctly when experiment runs provide their project identifier through the v2 response shape.
- A run rule with a 0 sampling rate processes no runs, but the scheduler still enumerated it every tick. The scheduler query now skips rules with sampling_rate 0 (parity with the is_enabled check), so they are never dispatched.
- Dataset and experiment tables now truncate long input and reference-output text and show detected base64 images as small thumbnails with a delayed larger preview, avoiding oversized hidden DOM content.
- Experiment tables now defer full payload rendering and output diff preparation until those views are requested, improving responsiveness for runs with large agent trajectories.
- Public dataset share links now resolve the sessions list (with stats) from SmithDB when ClickHouse querying is disabled, so shared dataset pages no longer fail to load on SmithDB-only deployments.
- Add conversation threads to a single annotation queue with POST /annotation-queues//items using item_type THREAD (thread_id + session_id). Mixed RUN and THREAD batches are supported; the server resolves threads via ClickHouse or SmithDB.
- Code evaluators now get more time to run each batch, so evaluators that import heavy libraries like scikit-learn are less likely to time out.
- POST /annotation-queues//items now accepts at most 200 items per request and returns a clear validation error when the limit is exceeded. Requests at the limit continue to succeed.
- Applying an evaluator to an existing experiment could fail with “Failed to start evaluation” on large experiments. It now starts reliably even when the run count is temporarily unavailable.
- Linked runs load correctly from public dataset shares when LangSmith uses the ClickHouse compatibility path.
Tracing
- The batched-run ingestion log now emits run_verbs as a list of run_id and verbs objects instead of a map keyed by run UUID, preventing structured-log aggregators from exhausting dynamic field limits.
- LangSmith now enforces user-defined monthly trace limits scoped to individual projects and users. New traces that exceed a configured limit are rejected, while patches and feedback for already-accepted traces continue to flow through.
- The tracing and evaluation onboarding quickstarts now show the correct LANGSMITH_ENDPOINT for bring-your-own-cloud data plane workspaces instead of the shared multi-tenant endpoint.
- Sharing, viewing, or unsharing any run in a trace now operates on the trace root, so every run in a shared trace is publicly viewable, and public run links open the selected run within the shared trace.
- Projects with existing traces no longer incorrectly display the onboarding screen when filtered or scoped to a time window with no recent runs. The project run-count check now looks back 30 days instead of the previous one-hour window.
- Bulk export compression now defaults to zstandard (zstd) for improved performance. Self-hosted environments retain the gzip default via the FF_BULK_EXPORT_DEFAULT_COMPRESSION environment variable.
- Authenticated users viewing public runs now see sidebar navigation for their last selected workspace. Logged-out viewers continue to see the public run without authenticated workspace navigation.
- LangSmith now returns clearer 409 Conflict messages when duplicate run create or update payloads are submitted. The message indicates whether the duplicate was a run create or run update request when possible.
- LangSmith MCP tools that fetch runs or thread history now accept project UUIDs in addition to project names, making trace URL investigations faster and less error-prone.
- OpenTelemetry resource attributes (set via OTEL_RESOURCE_ATTRIBUTES) now appear on traces as metadata namespaced under otel.resource.*, so you can attach details like user IDs without changing how your tracer emits spans.
- Vercel AI SDK traces sent over raw OpenTelemetry now render in the Messages view. Previously these traces showed an empty Messages tab because no format adapter claimed them.
- Thread stats requests that opt into streaming now return the main stats first and add feedback stats when they are ready.
- Native OpenTelemetry child spans are no longer dropped when they arrive before an SDK-attributed parent span; they are buffered and correctly nested regardless of arrival order.
- When a runs query times out, the runs table now shows a timeout banner for better responsiveness.
- LLM spans in the trace view now show the model provider’s brand logo (OpenAI, Anthropic, Google/Gemini, Azure, Mistral, DeepSeek, xAI, and speech providers), resolved from the run’s ls_provider metadata.
- LangSmith now preserves traces in multipart ingestion batches when one run has oversized inputs or outputs. Oversized input and output fields are replaced with a placeholder instead of rejecting the entire batch.
- Thread pages now show an explicit access-control message when trace loading is denied by ABAC, instead of a generic retrieval error.
- All time filters in tracing views now query the full retention window instead of falling back to a shorter backend default. This keeps trace, thread, and run results consistent when expanding the time range.
- OpenTelemetry traces from VS Code Copilot Chat now render as one clean nested trace per user turn. Auxiliary title/summary calls and orphaned tool spans are suppressed, message roles are corrected, token counts are de-duplicated, and standardized metadata (integration, agent runtime, thread ID, repo/git details) is attached automatically.
- Insights cluster run stats (run count, latency, tokens, and feedback) now reflect only the runs in each cluster instead of showing the same project-wide totals for every cluster.
- LangSmith Chat now authenticates to Chat LangChain with guest tokens when searching documentation, so docs answers keep working as Chat LangChain tightens authentication.
- The Trace Messages viewer now identifies the “main” conversation for traces that include middleware guardrails or subagent side-conversations, so the message list shows only the primary interaction instead of interleaving middleware/subagent partitions. Correctness is verified by an expanded snapshot suite covering 11 integrations across LangChain, OpenAI Agents SDK, Vercel AI SDK, Claude Agent SDK, deepagents, and raw provider wrappers.
- Fixed a bug where non-primitive metadata values did not appear in run details.
- Custom dashboard charts can now query P50 and P99 for input and output costs without failing runs analytics requests.
- Run stats scoped to an explicit run-id list (for example Insights per-cluster stats) now compute on SmithDB, which scopes results to those runs instead of falling back to project-wide totals.
- The thread stats API now accepts a
filterquery parameter, letting you scope aggregated stats to traces matching a LangSmith filter expression (e.g. start time or trace ID). - Organization model settings now let you search pricing rules by model name, match rule, or provider. Paginated loading fetches additional rules as you scroll, making large numbers of model price maps manageable.
- LangSmith Chat now mints Managed Deep Agent guest tokens from the Chat LangChain LangGraph host (
POST /identity/guest) when searching documentation, instead of the legacy Chat LangChain frontend guest route. - Assistant messages carrying tool calls were rendered twice in the v2 messages view for traces produced by the @anthropic-ai/sdk JavaScript SDK. Dedup now normalizes content-block field order so the same message emitted as an LLM output and replayed as an input on the next turn collapses to a single row.
- Run errors whose stack trace arrived fully escaped (no real line breaks) now render as properly formatted multi-line text instead of one long wrapped line.
- LangSmith MCP’s
fetch_runstool now acceptsmin_start_timeandmax_start_timearguments, so agents can search traces outside the default recent window. - Adds a
GET /v2/runs/{run_id}/urlendpoint that returns the LangSmith UI URL for a specific run.
Engine
- When an Engine project reaches its monthly spend limit, the Next Run status chip and project spend card now show a clear “Monthly spend limit reached” state with a button that takes you straight to raising the limit.
- Upgrades the Redis client to improve recovery from Redis cluster topology changes, fixing cases where cluster reconnects could stall.
- Engine now lets the parent agent recover from model-actionable subtask failures and retries transient provider or network errors before failing a run. This helps issue scans continue through recoverable model errors while preserving hard failures for auth, configuration, and code exceptions.
- LangSmith exposes Engine issue listing and retrieval through hosted MCP tools and generated SDK methods. Agents and API clients can fetch issue details directly by issue ID or filter issues by project, status, severity, tag, and update time.
- A new Engine board callout points you to the trace-scope setting, where you can restrict Engine’s reviews to runs matching a run name or metadata value.
- Engine-generated examples with assertions now add the Assertions evaluator when saved to a dataset from an annotation queue, matching the direct Add offline examples flow.
- The Engine setup screen now shows an estimated monthly cost based on the project’s recent trace volume and size, so you know roughly what to expect before starting analysis.
- The Engine issue list now uses a single filter and sort menu with a compact, nested layout for Priority, Status, Tags, and Sort by, replacing the previous two separate popovers.
- The Engine issue list now shows the active sort order as a removable chip next to your filter chips whenever it differs from the default.
- Engine issues can now be marked Fixing or Watching, and you can get a Slack alert when new traces recur on a watched issue.
- The Engine issue list no longer shows scan-timing details (next scan countdown, last run time, or a Run now action); a Pause/Resume control remains available in its own section in board settings.
- Engine now verifies concrete claims in agent responses against trace evidence, improving detection of ungrounded artifacts, values, and claimed actions.
Prompts and playground
- Self-hosted Playground and evaluator outbound model calls now honor proxy environment variables while preserving SSRF validation on every request.
- When you save a prompt to an application from the playground, LangSmith keeps the workspace application filter on All Applications instead of switching the rest of the UI to that application.
- Typing a workspace member’s name or email in the Context Hub search box now also returns the prompts and resources they created.
- The playground now includes Claude Sonnet 5, Claude Fable 5, and Claude Opus 4.8 in the Anthropic, Bedrock, and Vertex AI model selectors. New Anthropic playground sessions default to Claude Sonnet 5.
- Playground and evaluator calls to Amazon Bedrock using IAM Trusted Entity now resolve the correct LangSmith AWS credentials before assuming customer roles in AWS-hosted LangSmith. This fixes failures that reported “Failed to assume role” before the customer role was assumed.
- Playground runs now retain evaluator scores and reasoning while backend feedback updates are polled, preventing completed results from appearing blank.
- Outbound model calls that route through a forward proxy now send the original hostname in the proxy CONNECT tunnel instead of a resolved IP, so proxies that allowlist tunnel targets by domain no longer reject them. This fixes self-hosted Playground and evaluator calls to internal OpenAI-compatible endpoints reachable only through such a proxy.
- Reviewing a prompt commit now displays every extra parameter (such as verbosity) set on the model, not just a fixed subset.
- LangSmith now waits for model preset defaults to finish loading before initializing the Playground, preventing OpenAI from replacing a custom default preset during page load.
- The model configuration default button now switches to a selected state when you make a preset your default.
- Playground model settings now apply typed custom model names when the selector closes, so you no longer need to click the typed option explicitly.
- Custom evaluator errors in the Playground results table now reliably show the failure message, instead of sometimes displaying a blank error indicator.
- Configure workspace-wide HTTPS webhooks for every Context Hub commit, with signed payloads, custom headers, and secret rotation controls.
Feedback
- Editing the score on evaluator-generated feedback (for example from the experiment comparison view) now saves correctly instead of failing with “Failed to add feedback correction”.
- POST requests to add runs to an annotation queue accept an optional
extend_trace_retentionquery parameter. When set to false, short-lived traces are not upgraded to extended retention. The default remains true for backward compatibility. - Adding feedback or reviewer notes from the LangSmith UI no longer upgrades short-lived traces to extended retention. Long-lived traces are unchanged.
- Feedback statistics queries now route through the official ClickHouse client, resolving query failures and improving compatibility with ClickHouse 25.x.
- Feedback creation resolves run metadata from SmithDB when the client provides session and start time, so SmithDB-only deployments no longer depend on ClickHouse for eager feedback writes.
- Adding runs to an annotation queue via the by-key endpoint now falls back to the ClickHouse run lookup when SmithDB queries are disabled, so the SDK’s annotation-queue additions work regardless of whether SmithDB is enabled.
- The POST /feedback/eager endpoint is deprecated in favor of POST /feedback and is scheduled for removal on 2026-08-10. Update any direct integrations calling /feedback/eager to use POST /feedback instead.
- Feedback creation now accepts a thread identifier, enabling feedback to be associated with a conversation thread instead of only an individual run or session.
- GET feedback requests can now filter by a thread ID within a project, making thread-level feedback retrievable without resolving a run first.
- Annotation queue rubric feedback now loads the thread-scoped feedback for thread queue items.
- Annotation queue rubric feedback now saves against the selected thread for thread queue items.
Monitoring and alerting
- Alert chart previews now handle relative date ranges consistently, preventing failures when loading 14-day or 30-day previews.
- Dashboard chart tooltips and axes now show up to eight fractional digits (previously two), so very small costs and rates no longer round down to zero.
- Time-series charts on custom dashboards now leave gaps for missing data points instead of plotting them as zero, and lines connect across those gaps so trends remain readable.
- When a custom dashboard chart has no data or would produce too many bins, the empty state now surfaces the active stride (e.g. 1M) and selected range (e.g. Last 12 hours) so it’s clear what to adjust.
- When hovering the +N chip in a dashboard chart’s legend, the expanded popover now paints above adjacent chart cards instead of being clipped behind them.
- Metadata grouping keys without returned values no longer show a misleading empty value tooltip in dashboards.
Automations
- Applying a prebuilt evaluator without a filter now defaults to running on root runs only, matching manually created evaluators. Previously it ran on every nested run in a trace.
- Turning an online evaluator or automation on or off now saves for any role that can edit rules, instead of silently reverting for members without the retention-configuration permission.
- Resolved an unbounded memory leak in the SAQ queue worker where croniter objects were rebuilt every second, accumulating cached entries that were never released. The croniter dependency is bumped to 6.2.2+ and croniter objects are now reused across schedule ticks.
Deployment
- Self-hosted deployments can now request CPU and memory above the previous Cloud limits of 8/16 cores and 32/16 GB, bounded only by your cluster capacity. Lower bounds, multiple-of-128 granularity, and Redis memory ordering are still enforced.
- Custom Slack app triggers can now opt in to let third-party bots trigger an agent. Enable the allow bot triggers toggle on a registration to accept events from external bots; echoes from your own and other LangSmith-registered bots are still dropped to prevent loops.
- Agents now skip unreachable or misconfigured non-default MCP servers immediately instead of retrying them, removing a slow round-trip from the tool-loading step and cutting time-to-first-token.
- Standby (uptime) minutes for LangGraph Platform deployments could be billed more than once when replicas reported overlapping intervals across separate usage-reporting runs. Reporting now deduplicates each minute across runs so it is billed at most once.
- The multi-select dropdown (e.g. Selected Tools) on the Studio assistants page now renders above the configuration dialog instead of behind it, so its options are visible and selectable.
- Redis connections using Microsoft Entra ID (Azure IAM) authentication now re-authenticate automatically before the access token expires, so long-lived connections no longer drop. Clustered Azure Redis is now supported for IAM auth as well.
- The deployment Crons tab now shows each schedule in your local timezone instead of raw UTC, matching the Next Run Date column.
- LangSmith Deployment now supports updating a deployment to a fixed resource tier through the control plane API. The update applies the selected tier’s resource configuration, resizes Cloud SQL or RDS, and rolls a new revision.
- You can now edit an existing cron’s schedule, input, and end time from a deployment’s Crons tab, instead of deleting and recreating it.
- You can now rename a deployment from its Settings: give it a friendly display name without recreating it. The deployment’s URLs and infrastructure are unchanged.
- LangSmith frontend images now install nginx 1.31 packages to pick up the latest Chainguard security fixes.
- Deployment creation now checks free deployment usage with the same backend quota count used during submission, preventing the form from offering a free Serverless or Development option when the organization quota is already used.
- LangSmith Deployment now lets you update compute and database resource tiers independently for supported hosted deployments. The scaling action applies the selected resources and rolls out a new revision.
- Hosted project deployment views now label scale-to-zero development deployments as Serverless, with free deployments shown as Serverless (free).
- The deployment form now shows the free serverless option immediately while checking an organization’s remaining deployment allowance.
- Refines error handling when attempting to create a deployment with no GitHub repository selected.
- Serverless deployments can now update compute tiers correctly without requiring an external database tier.
- Self-hosted deployments now authenticate correctly to node-based AWS ElastiCache with IAM in both single-node and cluster configurations.
Sandboxes
- Sandbox command output is now re-chunked into bounded single WebSocket frames, so clients that do not reassemble continuation frames (including the Go SDK) can read large streamed or replayed output without truncated JSON.
- S3 sandbox mounts now default endpoint_url to https://s3.amazonaws.com when it is not provided, so the field is no longer required when mounting standard AWS S3 buckets.
- Sandboxes can now burst CPU up to 2x their requested allocation when the host has spare capacity, and you can request fractional (sub-core) vCPU down to 0.05.
- When creating a sandbox, you can now configure Git, S3, and GCS filesystem mounts, including mount paths, Git remotes, bucket settings, and cache options. Configured mounts appear in the sandbox table and detail view.
- The LangSmith SDKs now support creating, listing, updating, and deleting sandbox registries for pulling private container images, alongside the existing sandbox and snapshot operations.
- Sandbox creation no longer fails intermittently with “sandbox not ready” errors when an underlying host is disrupted. Affected capacity now retries the contended resource lock and recovers automatically instead of leaving the pool degraded.
- Sandbox host startup now validates the full version directory before reuse, so a missing initrd no longer causes create-time failures after a partial or stale install.
- Creating a sandbox snapshot from a Docker image now records the image’s tag (e.g. ubuntu:24.04 becomes the 24.04 tag), and creating a sandbox from a snapshot name without a tag resolves the latest tag, mirroring Docker.
- Self-hosted LangSmith installations now show the Sandboxes navigation item and use the instance-level sandbox flag to open the Sandboxes page.
- Shells and tools inside a sandbox now report the sandbox’s name as the hostname instead of a generic default, and the name resolves from within the sandbox.
- Self-hosted LangSmith installations can open the Sandboxes page without enabling the Deployments frontend.
- Sandboxes now set common CA-bundle environment variables by default, so Python, Node, Deno, curl, and git tooling automatically trusts the sandbox’s egress proxy certificate and no longer fails with TLS certificate-verification errors when its traffic is proxied.
- Sandboxes can now opt into keeping their memory when they stop, so the next start resumes where it left off instead of cold-booting. Set preserve_memory_on_stop when creating a sandbox; it defaults to off.
Administration
- The roles table on the Organization Roles settings page now scrolls correctly when there are more roles than fit on screen.
- A new Project and user limits tab on the enterprise Usage configuration page lets you set monthly trace-count limits scoped to a specific project or user. Add, edit, and delete limits from the page.
- Anonymous organizations now show an “Anonymity mode is on” banner on the members page, and the usage breakdown hides the group-by-user option for non-internal viewers.
- New API keys now default to a finite expiration date instead of requiring a custom value. When an organization enforces a shorter maximum, the form defaults to that maximum instead.
- You can now fetch a single workspace directly via GET /api/v1/workspaces/ instead of listing all workspaces and filtering client-side.
- Org and workspace admins can now edit the role of a pending member invite directly from the Members settings page, without needing to cancel and re-send the invite.
- The Usage limits page now shows each workspace’s configured total and extended (long-lived) trace limits, including caps that were previously hidden while the spend limit displayed “Unlimited”.
- The batch workspace invite endpoint no longer returns a 409 error when inviting users who are already pending org invitees or active org members. Those users are added directly to the workspace without requiring a new org invite.
- The role selector in the edit pending member invite dialog now uses a scrollable select, matching the invite flow. This ensures all custom roles are accessible when many workspace roles are defined.
- Self-hosted deployments can now encode spaces in the OIDC authorization request as %20 instead of +, so single sign-on works with identity providers that reject the default + encoding of the scope list. Enable it by setting OAUTH_URL_ENCODE_SCOPE_SPACES=true.
- Billing upgrade dialogs now stay within the viewport and scroll when payment or business details make the form taller than the screen.
- Non-admin callers with manage-members permission can no longer assign restricted roles to workspace members or invite users with restricted roles to the workspace.
- Filter the organization’s service keys and personal access tokens by workspace on the API keys settings page.
- Users without workspaces:manage permission cannot use restricted roles for invites, role changes, or user deletions in the UI.
- Organization admins can disable model providers across every workspace from organization settings. Disabled providers are hidden in the playground, evaluators, Fleet, and other model pickers, and workspace admins cannot re-enable them.
- Adding existing active or pending organization members to a workspace no longer fails when organization-level invites are disabled. Disabled org invites continue to block new organization invitees.
- The Roles settings page now scrolls correctly when an organization has more roles than fit on screen.
- Organization admins can once again edit the role of and remove other organization admins from the Organization Members settings page. Organization Operators, who share the same admin-level permissions but should not manage other admins, are now correctly prevented from editing, removing, or promoting members to Organization Admin.
- The email confirmation page now shows only the Confirm account step in the sidebar instead of future onboarding steps you have not reached yet.
- Self-hosted deployments now apply explicit DEFAULT_ORG_FEATURE_* and DEFAULT_FEATURE_* environment variables over stored organization and tenant config values, so operators can enable or disable features and limits globally without editing Postgres.
- The navigation product switcher now shows the configured organization logo alongside the LangSmith or Fleet wordmark instead of repeating the organization logo.
- Organization admins can now toggle role restriction from the Roles settings page. Restricted roles can only be assigned by users with the workspaces:manage permission.
- The organization-wide public sharing toggle now lives on the General settings page alongside the other organization settings, replacing its standalone Configuration section.
- When a user is removed from all mapped SSO groups, the organization and workspace access granted through SSO group sync is revoked on their next sign-in. Access assigned by other means (SCIM, JIT, or manual invitation) is unaffected.
- Workspace invite batch requests are now rate limited per workspace to reduce bulk invitation abuse. Learn more.
- Workspace switcher labels now show the full workspace name on hover when the visible label is truncated. This makes similarly prefixed workspace names easier to distinguish.
- LangSmith Home now shows a banner promoting Interrupt, our agent conference in London and NYC this fall, with a link to get tickets.
- Some new users could get stuck on the last onboarding step, with a loading spinner that never finished. This is now fixed.
- Organization admins can now rename their organization directly from the organization switcher in settings.
- Organization admins can now generate, view, and delete SCIM bearer tokens directly from Settings > Access and Security, instead of using the API, to set up SCIM provisioning with their identity provider. Learn more.
LLM Gateway
- LLM gateway data protection policies can now configure whether a guard pipeline timeout allows the request through or blocks it. Existing policies default to allowing requests on timeout.
- The LLM gateway now supports POST /openai/v1/responses/compact (and the legacy /responses/compact), routing it through the chat-shape responses handler.
- Guard policies now let you choose which PII rule categories to detect, with separate faster rule-based and slower model-based detection options, instead of a single on/off PII toggle.
- Gateway guard secret redaction now detects additional token formats, including SendGrid API tokens, Google OAuth access tokens, JWTs, Slack webhook URLs, and legacy LangSmith keys.
- When a gateway spend-cap policy targets more than one user, workspace, or API key, the create/edit policy form now explains that the limit applies to the combined spend across the selected entities rather than per entity.
- The LLM gateway now forwards every documented OpenAI API route it does not handle directly (models, files, batches, images, and more) to the upstream provider, so clients can reach the full OpenAI surface through the gateway. Custom OpenAI-compatible providers inherit the same passthrough routes.
- The LLM Gateway policies page now lets you sort each section by spend limit or usage percentage, and filter down to a specific workspace, user, or API key.
- LLM gateway data protection redaction now prepends a short disclaimer to redacted message text so models know SAFE_TO_USE placeholders are safe to reuse verbatim.
- The LLM Gateway now proxies Anthropic’s Files and Managed Agents endpoints, so you can use them with your gateway-managed workspace key alongside Messages and Models.
- Creating an LLM Gateway spend or data protection policy now applies to the organization you are signed in to, replacing the organization dropdown with a read-only display of the current organization.
- Long selected values, like a user’s email in the Gateway Policies filter, now truncate with an ellipsis instead of overlapping the dropdown chevron.
- The LLM Gateway now accepts workspace-scoped LangSmith OAuth bearer tokens across its provider routes, so OAuth clients can invoke configured models without a LangSmith API key.
Other
- When you add runs to an annotation queue without specifying
extend_trace_retention, short-lived traces stay on short-lived retention. Passextend_trace_retention=trueto upgrade traces to extended retention.
Observability and evaluations
Datasets and experiments
- Model, prompt, and tool chips in the Experiments table config cells now lay out from real measurements for accurate truncation, and the +N overflow badge is a clickable dropdown whose entries expose the same actions (filter, group by, open in playground, and details) as a chip’s own menu.
- Expanding the run tree for repetition runs in experiment comparison views now works reliably when a repetition root has a
project IDbut nosession ID. - Evaluators linked to Hub prompts now load correctly for flat and playground-shaped prompt commits, fixing crashes when editing existing evaluators.
- Code evaluator upload now accepts Python entrypoints annotated with PEP 604 union return types (for example
-> dict | None). POST /v2/datasets/{dataset_id}/experiment-runsis the supported public API for paginated experiment comparison. Legacy dataset comparison helpers are removed from the public OpenAPI spec and generated SDKs; existing HTTP routes continue to work for LangSmith UI clients.- Each example’s dataset splits now render as chips in the dataset Examples table, laid out from real measurements with a clickable +N overflow menu when an example belongs to more splits than fit the column.
- The experiment comparison view now offers an optional, reorderable “Splits (latest)” column that shows each example’s current dataset split assignments as chips, reflecting live membership rather than the as-of-run snapshot.
- Evaluator spend charts on project and dataset evaluator tabs keep their desktop layout on narrow screens and scroll horizontally instead of compressing the chart and stat cards.
- The experiment comparison and group-by views now show each example’s current dataset split rather than the split it had when the experiment ran, so you can tell whether failures already belong to a split without re-running the experiment.
- Comparison view now loads token and cost stats from SmithDB for root runs, so the stats columns populate again instead of staying blank
- LangSmith now caps reusable evaluators per workspace to prevent unbounded resource growth. Contact support if your workspace needs a higher limit.
- Creating dataset examples from source runs now correctly fetches run inputs and outputs backed by SmithDB, and no longer fails the whole request if one of several source runs can’t be found.
- Select multiple rows in an experiment (or select all matching the current filters) and add, replace, or remove their dataset splits in one action, or copy the selected examples to another dataset, instead of editing rows one at a time.
- The
/runs/rules/validateendpoint now supports thread evaluators. Passtest_thread_idandsession_idto test a multi-turn evaluator against a real conversation before saving. - Custom code evaluators that time out or fail on a run now record an error on that run instead of silently leaving it without feedback, so partial evaluation failures are visible on the experiment.
- The Open source run action on an example page now reads session and start time from dedicated example fields populated at creation, enabling reliable navigation to the source trace on SmithDB.
- The thread evaluator config preview now shows the thread message formats the evaluator actually maps, instead of listing every available format.
- The evaluator config now shows a locked “Trace count ≥ 2” filter for managed thread evaluators, making it clear they only run on threads with multiple turns.
- Experiment comparison and individual experiment views now load run rows on self-hosted deployments that authenticate the UI via SSO/OAuth session cookies. Previously these views could show ‘No results found’ even though metrics and feedback loaded.
- Experiment statistics now refresh promptly for recently run experiments while keeping historical experiment scans bounded.
- The Assertions evaluator added via “Add evaluator” now reads assertions from the reference output like the auto-attached version, so it grades against the real assertions instead of always failing.
- Evaluator spend chart y-axes now abbreviate amounts of $1,000 or more, making high-spend values easier to scan.
- A run rule whose sampling rate was 0 (or unset) sent an out-of-range sample_rate to the SmithDB query service (which rejected it) and zeroed out ClickHouse thread grouping. Both the flat and grouped fetch paths now fall back to 1.0 (no sampling) so these rules query successfully.
- Exporting a dataset comparison view as CSV now returns a clear “file is too large to export” error instead of a generic server error when the export exceeds internal size limits.
- Each split chip in a row’s Splits cell is now interactive in the experiment results and comparison views, with an Edit splits action that opens the single-example split picker so you can reassign splits without leaving the table.
Tracing
- The batched-run ingestion log now emits run_verbs as a list of run_id and verbs objects instead of a map keyed by run UUID, preventing structured-log aggregators from exhausting dynamic field limits.
- LangSmith now enforces user-defined monthly trace limits scoped to individual projects and users. New traces that exceed a configured limit are rejected, while patches and feedback for already-accepted traces continue to flow through.
- The tracing and evaluation onboarding quickstarts now show the correct
LANGSMITH_ENDPOINTfor bring-your-own-cloud data plane workspaces instead of the shared multi-tenant endpoint. - Sharing, viewing, or unsharing any run in a trace now operates on the trace root, so every run in a shared trace is publicly viewable, and public run links open the selected run within the shared trace.
- Projects with existing traces no longer incorrectly display the onboarding screen when filtered or scoped to a time window with no recent runs. The project run-count check now looks back 30 days instead of the previous one-hour window.
- Bulk export compression now defaults to zstandard (zstd) for improved performance. Self-hosted environments retain the gzip default via the
FF_BULK_EXPORT_DEFAULT_COMPRESSIONenvironment variable. - Authenticated users viewing public runs now see sidebar navigation for their last selected workspace. Logged-out viewers continue to see the public run without authenticated workspace navigation.
- LangSmith now returns clearer 409 Conflict messages when duplicate run create or update payloads are submitted. The message indicates whether the duplicate was a run create or run update request when possible.
- LangSmith MCP tools that fetch runs or thread history now accept project UUIDs in addition to project names, making trace URL investigations faster and less error-prone.
- OpenTelemetry resource attributes (set via
OTEL_RESOURCE_ATTRIBUTES) now appear on traces as metadata namespaced underotel.resource.*, so you can attach details like user IDs without changing how your tracer emits spans. - Vercel AI SDK traces sent over raw OpenTelemetry now render in the Messages view. Previously these traces showed an empty Messages tab because no format adapter claimed them.
- Thread stats requests that opt into streaming now return the main stats first and add feedback stats when they are ready.
- Native OpenTelemetry child spans are no longer dropped when they arrive before an SDK-attributed parent span; they are buffered and correctly nested regardless of arrival order.
- When a runs query times out, the runs table now shows a timeout banner for better responsiveness.
- LangSmith now preserves traces in multipart ingestion batches when one run has oversized inputs or outputs. Oversized input and output fields are replaced with a placeholder instead of rejecting the entire batch.
- Thread pages now show an explicit access-control message when trace loading is denied by ABAC, instead of a generic retrieval error.
- All time filters in tracing views now query the full retention window instead of falling back to a shorter backend default. This keeps trace, thread, and run results consistent when expanding the time range.
- OpenTelemetry traces from VS Code Copilot Chat now render as one clean nested trace per user turn. Auxiliary title/summary calls and orphaned tool spans are suppressed, message roles are corrected, token counts are de-duplicated, and standardized metadata (integration, agent runtime, thread ID, repo/git details) is attached automatically.
- Insights cluster run stats (run count, latency, tokens, and feedback) now reflect only the runs in each cluster instead of showing the same project-wide totals for every cluster.
- LangSmith Chat now authenticates to Chat LangChain with guest tokens when searching documentation, so docs answers keep working as Chat LangChain tightens authentication.
- The Trace Messages viewer now identifies the “main” conversation for traces that include middleware guardrails or subagent side-conversations, so the message list shows only the primary interaction instead of interleaving middleware/subagent partitions. Correctness is verified by an expanded snapshot suite covering 11 integrations across LangChain, OpenAI Agents SDK, Vercel AI SDK, Claude Agent SDK, deepagents, and raw provider wrappers.
- Custom dashboard charts can now query P50 and P99 for input and output costs without failing runs analytics requests.
- The thread stats API now accepts a
filterquery parameter, letting you scope aggregated stats to traces matching a LangSmith filter expression (e.g. start time or trace ID). - LangSmith Chat now mints Managed Deep Agent guest tokens from the Chat LangChain LangGraph host (
POST /identity/guest) when searching documentation, instead of the legacy Chat LangChain frontend guest route.
Engine
- When an Engine project reaches its monthly spend limit, the Next Run status chip and project spend card now show a clear “Monthly spend limit reached” state with a button that takes you straight to raising the limit.
- Upgrades the Redis client to improve recovery from Redis cluster topology changes, fixing cases where cluster reconnects could stall.
- Engine now lets the parent agent recover from model-actionable subtask failures and retries transient provider or network errors before failing a run. This helps issue scans continue through recoverable model errors while preserving hard failures for auth, configuration, and code exceptions.
- LangSmith exposes Engine issue listing and retrieval through hosted MCP tools and generated SDK methods. Agents and API clients can fetch issue details directly by issue ID or filter issues by project, status, severity, tag, and update time.
- A new Engine board callout points you to the trace-scope setting, where you can restrict Engine’s reviews to runs matching a run name or metadata value.
- Engine-generated examples with assertions now add the Assertions evaluator when saved to a dataset from an annotation queue, matching the direct Add offline examples flow.
- The Engine issue list now uses a single filter and sort menu with a compact, collapsible layout for Priority, Status, Tags, and Sort by, replacing the previous two separate popovers.
- The Engine issue list now shows the active sort order as a removable chip next to your filter chips whenever it differs from the default.
- The Engine issue list no longer shows scan-timing details (next scan countdown, last run time, or a Run now action); a Pause/Resume control remains available in its own section in board settings.
Prompts and playground
- Self-hosted Playground and evaluator outbound model calls now honor proxy environment variables while preserving SSRF validation on every request.
- When you save a prompt to an application from the playground, LangSmith keeps the workspace application filter on All Applications instead of switching the rest of the UI to that application.
- Typing a workspace member’s name or email in the Context Hub search box now also returns the prompts and resources they created.
- The playground now includes Claude Sonnet 5, Claude Fable 5, and Claude Opus 4.8 in the Anthropic, Bedrock, and Vertex AI model selectors. New Anthropic playground sessions default to Claude Sonnet 5.
- Playground and evaluator calls to Amazon Bedrock using IAM Trusted Entity now resolve the correct LangSmith AWS credentials before assuming customer roles in AWS-hosted LangSmith. This fixes failures that reported “Failed to assume role” before the customer role was assumed.
- Outbound model calls that route through a forward proxy now send the original hostname in the proxy CONNECT tunnel instead of a resolved IP, so proxies that allowlist tunnel targets by domain no longer reject them. This fixes self-hosted Playground and evaluator calls to internal OpenAI-compatible endpoints reachable only through such a proxy.
- Reviewing a prompt commit now displays every extra parameter (such as verbosity) set on the model, not just a fixed subset.
Feedback
- Editing the score on evaluator-generated feedback (for example from the experiment comparison view) now saves correctly instead of failing with “Failed to add feedback correction”.
- POST requests to add runs to an annotation queue accept an optional
extend_trace_retentionquery parameter. When set to false, short-lived traces are not upgraded to extended retention. The default remains true for backward compatibility. - Adding feedback or reviewer notes from the LangSmith UI no longer upgrades short-lived traces to extended retention. Long-lived traces are unchanged.
- Feedback statistics queries now route through the official ClickHouse client, resolving query failures and improving compatibility with ClickHouse 25.x.
- Feedback creation resolves run metadata from SmithDB when the client provides session and start time, so SmithDB-only deployments no longer depend on ClickHouse for eager feedback writes.
- The
POST /feedback/eagerendpoint is deprecated in favor ofPOST /feedbackand is scheduled for removal on 2026-08-10. Update any direct integrations calling/feedback/eagerto usePOST /feedbackinstead.
Monitoring and alerting
- Alert chart previews now handle relative date ranges consistently, preventing failures when loading 14-day or 30-day previews.
- Dashboard chart tooltips and axes now show up to eight fractional digits (previously two), so very small costs and rates no longer round down to zero.
- Time-series charts on custom dashboards now leave gaps for missing data points instead of plotting them as zero, and lines connect across those gaps so trends remain readable.
Automations
- Applying a prebuilt evaluator without a filter now defaults to running on root runs only, matching manually created evaluators. Previously it ran on every nested run in a trace.
- Turning an online evaluator or automation on or off now saves for any role that can edit rules, instead of silently reverting for members without the retention-configuration permission.
Deployment
- Self-hosted deployments can now request CPU and memory above the previous Cloud limits of 8/16 cores and 32/16 GB, bounded only by your cluster capacity. Lower bounds, multiple-of-128 granularity, and Redis memory ordering are still enforced.
- Custom Slack app triggers can now opt in to let third-party bots trigger an agent. Enable the allow bot triggers toggle on a registration to accept events from external bots; echoes from your own and other LangSmith-registered bots are still dropped to prevent loops.
- Agents now skip unreachable or misconfigured non-default MCP servers immediately instead of retrying them, removing a slow round-trip from the tool-loading step and cutting time-to-first-token.
- Standby (uptime) minutes for LangGraph Platform deployments could be billed more than once when replicas reported overlapping intervals across separate usage-reporting runs. Reporting now deduplicates each minute across runs so it is billed at most once.
- The multi-select dropdown (e.g. Selected Tools) on the Studio assistants page now renders above the configuration dialog instead of behind it, so its options are visible and selectable.
- Redis connections using Microsoft Entra ID (Azure IAM) authentication now re-authenticate automatically before the access token expires, so long-lived connections no longer drop. Clustered Azure Redis is now supported for IAM auth as well.
- The deployment Crons tab now shows each schedule in your local timezone instead of raw UTC, matching the Next Run Date column.
- LangSmith Deployment now supports updating a deployment to a fixed resource tier through the control plane API. The update applies the selected tier’s resource configuration, resizes Cloud SQL or RDS, and rolls a new revision.
- You can now rename a deployment from its Settings: give it a friendly display name without recreating it. The deployment’s URLs and infrastructure are unchanged.
Sandboxes
- Sandbox command output is now re-chunked into bounded single WebSocket frames, so clients that do not reassemble continuation frames (including the Go SDK) can read large streamed or replayed output without truncated JSON.
- S3 sandbox mounts now default endpoint_url to https://s3.amazonaws.com when it is not provided, so the field is no longer required when mounting standard AWS S3 buckets.
- Sandboxes can now burst CPU up to 2x their requested allocation when the host has spare capacity, and you can request fractional (sub-core) vCPU down to 0.05.
- When creating a sandbox, you can now configure Git, S3, and GCS filesystem mounts, including mount paths, Git remotes, bucket settings, and cache options. Configured mounts appear in the sandbox table and detail view.
- The LangSmith SDKs now support creating, listing, updating, and deleting sandbox registries for pulling private container images, alongside the existing sandbox and snapshot operations.
- Sandbox creation no longer fails intermittently with “sandbox not ready” errors when an underlying host is disrupted. Affected capacity now retries the contended resource lock and recovers automatically instead of leaving the pool degraded.
- Sandbox host startup now validates the full version directory before reuse, so a missing initrd no longer causes create-time failures after a partial or stale install.
- Creating a sandbox snapshot from a Docker image now records the image’s tag (e.g. ubuntu:24.04 becomes the 24.04 tag), and creating a sandbox from a snapshot name without a tag resolves the latest tag, mirroring Docker.
- Self-hosted LangSmith installations now show the Sandboxes navigation item and use the instance-level sandbox flag to open the Sandboxes page.
- Shells and tools inside a sandbox now report the sandbox’s name as the hostname instead of a generic default, and the name resolves from within the sandbox.
- Self-hosted LangSmith installations can open the Sandboxes page without enabling the Deployments frontend.
- Sandboxes now set common CA-bundle environment variables by default, so Python, Node, Deno, curl, and git tooling automatically trusts the sandbox’s egress proxy certificate and no longer fails with TLS certificate-verification errors when its traffic is proxied.
Administration
- The roles table on the Organization Roles settings page now scrolls correctly when there are more roles than fit on screen.
- A new Project and user limits tab on the enterprise Usage configuration page lets you set monthly trace-count limits scoped to a specific project or user. Add, edit, and delete limits from the page.
- Anonymous organizations now show an “Anonymity mode is on” banner on the members page, and the usage breakdown hides the group-by-user option for non-internal viewers.
- New API keys now default to a finite expiration date instead of requiring a custom value. When an organization enforces a shorter maximum, the form defaults to that maximum instead.
- You can now fetch a single workspace directly via
GET /api/v1/workspaces/{workspace_id}instead of listing all workspaces and filtering client-side. - Org and workspace admins can now edit the role of a pending member invite directly from the Members settings page, without needing to cancel and re-send the invite.
- The Usage limits page now shows each workspace’s configured total and extended (long-lived) trace limits, including caps that were previously hidden while the spend limit displayed “Unlimited”.
- The batch workspace invite endpoint no longer returns a 409 error when inviting users who are already pending org invitees or active org members. Those users are added directly to the workspace without requiring a new org invite.
- The role selector in the edit pending member invite dialog now uses a scrollable select, matching the invite flow. This ensures all custom roles are accessible when many workspace roles are defined.
- Self-hosted deployments can now encode spaces in the OIDC authorization request as %20 instead of +, so single sign-on works with identity providers that reject the default + encoding of the scope list. Enable it by setting OAUTH_URL_ENCODE_SCOPE_SPACES=true.
- Billing upgrade dialogs now stay within the viewport and scroll when payment or business details make the form taller than the screen.
- Non-admin callers with manage-members permission can no longer assign restricted roles to workspace members or invite users with restricted roles to the workspace.
- Filter the organization’s service keys and personal access tokens by workspace on the API keys settings page.
- Users without workspaces:manage permission cannot use restricted roles for invites, role changes, or user deletions in the UI.
- Organization admins can disable model providers across every workspace from organization settings. Disabled providers are hidden in the playground, evaluators, Fleet, and other model pickers, and workspace admins cannot re-enable them.
- Adding existing active or pending organization members to a workspace no longer fails when organization-level invites are disabled. Disabled org invites continue to block new organization invitees.
- The Roles settings page now scrolls correctly when an organization has more roles than fit on screen.
- Organization admins can once again edit the role of and remove other organization admins from the Organization Members settings page. Organization Operators, who share the same admin-level permissions but should not manage other admins, are now correctly prevented from editing, removing, or promoting members to Organization Admin.
- The email confirmation page now shows only the Confirm account step in the sidebar instead of future onboarding steps you have not reached yet.
- Self-hosted deployments now apply explicit DEFAULT_ORG_FEATURE_* and DEFAULT_FEATURE_* environment variables over stored organization and tenant config values, so operators can enable or disable features and limits globally without editing Postgres.
- The navigation product switcher now shows the configured organization logo alongside the LangSmith or Fleet wordmark instead of repeating the organization logo.
- Organization admins can now toggle role restriction from the Roles settings page. Restricted roles can only be assigned by users with the workspaces:manage permission.
- The organization-wide public sharing toggle now lives on the General settings page alongside the other organization settings, replacing its standalone Configuration section.
- When a user is removed from all mapped SSO groups, the organization and workspace access granted through SSO group sync is revoked on their next sign-in. Access assigned by other means (SCIM, JIT, or manual invitation) is unaffected.
- Workspace invite batch requests are now rate limited per workspace to reduce bulk invitation abuse. Learn more.
- LangSmith Home now shows a banner promoting Interrupt, our agent conference in London and NYC this fall, with a link to get tickets.
LLM Gateway
- LLM gateway data protection policies can now configure whether a guard pipeline timeout allows the request through or blocks it. Existing policies default to allowing requests on timeout.
- The LLM gateway now supports
POST /openai/v1/responses/compact(and the legacy/responses/compact), routing it through the chat-shape responses handler. - Guard policies now let you choose which PII rule categories to detect, with separate faster rule-based and slower model-based detection options, instead of a single on/off PII toggle.
- Gateway guard secret redaction now detects additional token formats, including SendGrid API tokens, Google OAuth access tokens, JWTs, Slack webhook URLs, and legacy LangSmith keys.
- When a gateway spend-cap policy targets more than one user, workspace, or API key, the create/edit policy form now explains that the limit applies to the combined spend across the selected entities rather than per entity.
- The LLM gateway now forwards every documented OpenAI API route it does not handle directly (models, files, batches, images, and more) to the upstream provider, so clients can reach the full OpenAI surface through the gateway. Custom OpenAI-compatible providers inherit the same passthrough routes.
- The LLM Gateway policies page now lets you sort each section by spend limit or usage percentage, and filter down to a specific workspace, user, or API key.
- LLM gateway data protection redaction now prepends a short disclaimer to redacted message text so models know SAFE_TO_USE placeholders are safe to reuse verbatim.
- The LLM Gateway now proxies Anthropic’s Files and Managed Agents endpoints, so you can use them with your gateway-managed workspace key alongside Messages and Models.
- Creating an LLM Gateway spend or data protection policy now applies to the organization you are signed in to, replacing the organization dropdown with a read-only display of the current organization.
- Long selected values, like a user’s email in the Gateway Policies filter, now truncate with an ellipsis instead of overlapping the dropdown chevron.
- The LLM Gateway now accepts workspace-scoped LangSmith OAuth bearer tokens across its provider routes, so OAuth clients can invoke configured models without a LangSmith API key.
Other
- When you add runs to an annotation queue without specifying
extend_trace_retention, short-lived traces stay on short-lived retention. Passextend_trace_retention=trueto upgrade traces to extended retention.
Observability and evaluations
Datasets and experiments
- Model, prompt, and tool chips in the Experiments table config cells now lay out from real measurements for accurate truncation, and the +N overflow badge is a clickable dropdown whose entries expose the same actions (filter, group by, open in playground, and details) as a chip’s own menu.
- Expanding the run tree for repetition runs in experiment comparison views now works reliably when a repetition root has a
project IDbut nosession ID. - Evaluators linked to Hub prompts now load correctly for flat and playground-shaped prompt commits, fixing crashes when editing existing evaluators.
- Code evaluator upload now accepts Python entrypoints annotated with PEP 604 union return types (for example
-> dict | None). POST /v2/datasets/{dataset_id}/experiment-runsis the supported public API for paginated experiment comparison. Legacy dataset comparison helpers are removed from the public OpenAPI spec and generated SDKs; existing HTTP routes continue to work for LangSmith UI clients.- Each example’s dataset splits now render as chips in the dataset Examples table, laid out from real measurements with a clickable +N overflow menu when an example belongs to more splits than fit the column.
- The experiment comparison view now offers an optional, reorderable “Splits (latest)” column that shows each example’s current dataset split assignments as chips, reflecting live membership rather than the as-of-run snapshot.
- Evaluators spend charts on project and dataset evaluator tabs keep their desktop layout on narrow screens and scroll horizontally instead of compressing the chart and stat cards.
- The experiment comparison and group-by views now show each example’s current dataset split rather than the split it had when the experiment ran, so you can tell whether failures already belong to a split without re-running the experiment.
- LangSmith now caps reusable evaluators per workspace to prevent unbounded resource growth. Contact support if your workspace needs a higher limit.
- Creating dataset examples from source runs now correctly fetches run inputs and outputs backed by SmithDB, and no longer fails the whole request if one of several source runs cannot be found.
- Custom code evaluators that time out or fail on a run now record an error on that run instead of silently leaving it without feedback, so partial evaluation failures are visible on the experiment.
- The Open source run action on an example page now reads session and start time from dedicated example fields populated at creation, enabling reliable navigation to the source trace on SmithDB.
Tracing
- LangSmith now enforces user-defined monthly trace limits scoped to individual projects and users. New traces that exceed a configured limit are rejected, while patches and feedback for already-accepted traces continue to flow through.
- The tracing and evaluation onboarding quickstarts now show the correct
LANGSMITH_ENDPOINTfor bring-your-own-cloud data plane workspaces instead of the shared multi-tenant endpoint. - Sharing, viewing, or unsharing any run in a trace now operates on the trace root, so every run in a shared trace is publicly viewable, and public run links open the selected run within the shared trace.
- Projects with existing traces no longer incorrectly display the onboarding screen when filtered or scoped to a time window with no recent runs. The project run-count check now looks back 30 days instead of the previous one-hour window.
- Bulk export compression now defaults to zstandard (zstd) for improved performance. Self-hosted environments retain the gzip default via the
FF_BULK_EXPORT_DEFAULT_COMPRESSIONenvironment variable. - LangSmith now returns clearer 409 Conflict messages when duplicate run create or update payloads are submitted. The message indicates whether the duplicate was a run create or run update request when possible.
- LangSmith MCP tools that fetch runs or thread history now accept
project UUIDsin addition to project names, making trace URL investigations faster and less error-prone. - OpenTelemetry resource attributes (set via
OTEL_RESOURCE_ATTRIBUTES) now appear on traces as metadata namespaced under otel.resource.*, so you can attach details like user IDs without changing how your tracer emits spans. - Vercel AI SDK traces sent over raw OpenTelemetry now render in the Messages view. Previously these traces showed an empty Messages tab because no format adapter claimed them.
- Thread stats requests that opt into streaming now return the main stats first and add feedback stats when they are ready.
- When a runs query times out, the runs table now shows a timeout banner for better responsiveness.
- LangSmith now preserves traces in multipart ingestion batches when one run has oversized inputs or outputs. Oversized input and output fields are replaced with a placeholder instead of rejecting the entire batch.
Engine
- When an Engine project reaches its monthly spend limit, the Next Run status chip and project spend card now show a clear “Monthly spend limit reached” state with a button that takes you straight to raising the limit.
- LangSmith exposes Engine issue listing and retrieval through hosted MCP tools and generated SDK methods. Agents and API clients can fetch issue details directly by
issue IDor filter issues by project, status, severity, tag, and update time. - A new Engine board callout points you to the trace-scope setting, where you can restrict Engine’s reviews to runs matching a run name or metadata value.
Prompts and playground
- Self-hosted Playground and evaluator outbound model calls now honor proxy environment variables while preserving SSRF validation on every request.
- When you save a prompt to an application from the playground, LangSmith keeps the workspace application filter on All Applications instead of switching the rest of the UI to that application.
- Typing a workspace member’s name or email in the Context Hub search box now also returns the prompts and resources they created.
- The playground now includes Claude Sonnet 5, Claude Fable 5, and Claude Opus 4.8 in the Anthropic, Bedrock, and Vertex AI model selectors. New Anthropic playground sessions default to Claude Sonnet 5.
Feedback
- Editing the score on evaluator-generated feedback (for example from the experiment comparison view) now saves correctly instead of failing with “Failed to add feedback correction”.
- Adding feedback or reviewer notes from the LangSmith UI no longer upgrades short-lived traces to extended retention. Long-lived traces are unchanged.
- Feedback statistics queries now route through the official ClickHouse client, resolving query failures and improving compatibility with ClickHouse 25.x.
Monitoring and alerting
- Alert chart previews now handle relative date ranges consistently, preventing failures when loading 14-day or 30-day previews.
- Dashboards chart tooltips and axes now show up to eight fractional digits (previously two), so very small costs and rates no longer round down to zero.
- Time-series charts on custom dashboards now leave gaps for missing data points instead of plotting them as zero, and lines connect across those gaps so trends remain readable.
Automations
- Applying a prebuilt evaluator without a filter now defaults to running on root runs only, matching manually created evaluators. Previously it ran on every nested run in a trace.
- Turning an online evaluator or automation on or off now saves for any role that can edit rules, instead of silently reverting for members without the retention-configuration permission.
Deployment
- Self-hosted deployments can now request CPU and memory above the previous Cloud limits of 8/16 cores and 32/16 GB, bounded only by your cluster capacity. Lower bounds, multiple-of-128 granularity, and Redis memory ordering are still enforced.
- Custom Slack app triggers can now opt in to let third-party bots trigger an agent. Enable the allow bot triggers toggle on a registration to accept events from external bots; echoes from your own and other LangSmith-registered bots are still dropped to prevent loops.
- Agents now skip unreachable or misconfigured non-default MCP servers immediately instead of retrying them, removing a slow round-trip from the tool-loading step and cutting time-to-first-token.
- Standby (uptime) minutes for LangGraph Platform deployments could be billed more than once when replicas reported overlapping intervals across separate usage-reporting runs. Reporting now deduplicates each minute across runs so it is billed at most once.
Sandboxes
- Sandbox command output is now re-chunked into bounded single WebSocket frames, so clients that do not reassemble continuation frames (including the Go SDK) can read large streamed or replayed output without truncated JSON.
- S3 sandbox mounts now default
endpoint_urlto https://s3.amazonaws.com when it is not provided, so the field is no longer required when mounting standard AWS S3 buckets. - Sandboxes can now burst CPU up to 2x their requested allocation when the host has spare capacity, and you can request fractional (sub-core) vCPU down to 0.05.
- When creating a sandbox, you can now configure Git, S3, and GCS filesystem mounts, including mount paths, Git remotes, bucket settings, and cache options. Configured mounts appear in the sandbox table and detail view.
- The LangSmith SDKs now support creating, listing, updating, and deleting sandbox registries for pulling private container images, alongside the existing sandbox and snapshot operations.
- Sandbox snapshot builds can now request an XFS root filesystem for sandbox-host based environments.
- Sandbox creation no longer fails intermittently with “sandbox not ready” errors when an underlying host is disrupted. Affected capacity now retries the contended resource lock and recovers automatically instead of leaving the pool degraded.
- Sandbox host startup now validates the full version directory before reuse, so a missing initrd no longer causes create-time failures after a partial or stale install.
- Creating a sandbox snapshot from a Docker image now records the image’s tag (e.g. ubuntu:24.04 becomes the 24.04 tag), and creating a sandbox from a snapshot name without a tag resolves the latest tag, mirroring Docker.
- Self-hosted LangSmith installations now show the Sandboxes navigation item and use the instance-level sandbox flag to open the Sandboxes page.
Administration
- A new Project and user limits tab on the enterprise Usage configuration page lets you set monthly trace-count limits scoped to a specific project or user. Add, edit, and delete limits from the page.
- Anonymous organizations now show an “Anonymity mode is on” banner on the members page, and the usage breakdown hides the group-by-user option for non-internal viewers.
- New API keys now default to a finite expiration date instead of requiring a custom value. When an organization enforces a shorter maximum, the form defaults to that maximum instead.
- You can now fetch a single workspace directly via GET /api/v1/workspaces/ instead of listing all workspaces and filtering client-side.
- Org and workspace admins can now edit the role of a pending member invite directly from the Members settings page, without needing to cancel and re-send the invite.
- The Usage limits page now shows each workspace’s configured total and extended (long-lived) trace limits, including caps that were previously hidden while the spend limit displayed “Unlimited”.
- The batch workspace invite endpoint no longer returns a 409 error when inviting users who are already pending org invitees or active org members. Those users are added directly to the workspace without requiring a new org invite.
- The role selector in the edit pending member invite dialog now uses a scrollable select, matching the invite flow. This ensures all custom roles are accessible when many workspace roles are defined.
- Self-hosted deployments can now encode spaces in the OIDC authorization request as %20 instead of +, so single sign-on works with identity providers that reject the default + encoding of the scope list. Enable it by setting OAUTH_URL_ENCODE_SCOPE_SPACES=true.
- Billing upgrade dialogs now stay within the viewport and scroll when payment or business details make the form taller than the screen.
- Non-admin callers with manage-members permission can no longer assign restricted roles to workspace members or invite users with restricted roles to the workspace.
- Filter the organization’s service keys and personal access tokens by workspace on the API keys settings page.
- Users without workspaces:manage permission cannot use restricted roles for invites, role changes, or user deletions in the UI.
- Adding existing active or pending organization members to a workspace no longer fails when organization-level invites are disabled. Disabled org invites continue to block new organization invitees.
- The Roles settings page now scrolls correctly when an organization has more roles than fit on screen.
- Organization admins can once again edit the role of and remove other organization admins from the Organization Members settings page. Organization Operators, who share the same admin-level permissions but should not manage other admins, are now correctly prevented from editing, removing, or promoting members to Organization Admin.
LLM Gateway
- LLM gateway data protection policies can now configure whether a guard pipeline timeout allows the request through or blocks it. Existing policies default to allowing requests on timeout.
- The LLM gateway now supports POST /openai/v1/responses/compact (and the legacy /responses/compact), routing it through the chat-shape responses handler.
- Guard policies now let you choose which PII rule categories to detect, with separate faster rule-based and slower model-based detection options, instead of a single on/off PII toggle.
- Gateway guard secret redaction now detects additional token formats, including SendGrid API tokens, Google OAuth access tokens, JWTs, Slack webhook URLs, and legacy LangSmith keys.
- The LLM Gateway policies page now lets you sort each section by spend limit or usage percentage, and filter down to a specific workspace, user, or API key.
Observability and evaluations
Automations
- Automations now let you control trace retention per action, so traces matched by a rule can stay at base retention instead of being upgraded.
Engine
- The Engine issue board now shows a Connect GitHub action when GitHub is not connected, so you can set up pull request creation without leaving the board.
- Engine now has a unified enablement screen with access requests, and organization settings consolidate Engine usage and limits in one place.
- Organization admins now receive Engine spend emails when spend crosses each configured threshold, and pausing or disabling Engine now asks for confirmation.
Datasets and experiments
- Experiments now show live loading progress in the header and the Progress column, so you can track completed and evaluated runs in real time.
- Evaluators now include a trace-retention toggle in the advanced options, so scored traces can stay at base retention when that fits your workflow.
- Evaluator prompt editing now offers an advanced mode for editing Mustache templates directly with separate variable mappings.
- You can now apply resource tags when creating a dataset, including from scratch, file upload, or a clone.
- Auto-attached Assertions evaluators now read assertions from the reference output, so experiment scores reflect actual pass and fail results.
Prompts and playground
- OAuth client credentials now support per-workspace setup on model configurations, so workspace admins can self-serve OAuth on saved prompts and models.
- The Playground now exposes a Reasoning Summary option for OpenAI reasoning models on the Responses API.
- The model dropdown no longer suggests OpenAI models for an OpenAI Compatible Endpoint, so you can enter your own custom model name.
Tracing
- Trace query syntax now has a full operator reference, field table, and quick examples, so API filtering is easier to discover.
- The OpenTelemetry guide now explains how to link spans to an existing LangSmith SDK trace and what happens when a parent span never arrives, so cross-process traces are easier to debug.
Monitoring and alerting
- Dashboards now include a chart builder with chart templates, a create and edit pane, and brush and series controls on time series charts.
- You can now send alerts to Slack as a native notification target and connect or disconnect the Slack app from the UI.
Deployment
- Preview deployments now build the image for the preview commit instead of reusing the parent deployment’s image.
Sandboxes
- Sandbox auth proxy now documents GCP rules and service-account handling, so Google API access through the proxy is clearer.
- Sandboxes now marks AWS US SaaS availability as generally available, so the region table reflects the current rollout.
- Sandboxes now support Git mounts and Google Cloud Storage bucket mounts.
Admin and billing
Administration
- Organization settings now clarify that SSO/SCIM group names can omit spaces, so enterprise IdPs that disallow spaces still work cleanly.
- The Vanta MCP integration is now generally available to all workspaces.
- Applying tags when creating datasets, prompts, and projects is now governed by dedicated tag-on-create permissions.
LLM Gateway
- The LLM gateway now supports native Gemini routes for Vertex AI and the OpenAI embeddings endpoint.
- Gateway guard policies now accept a granular PII configuration and a configurable timeout action.
Usage and billing
- Granular billable usage now clarifies org scoping, so you can interpret usage totals more accurately.
Observability and evaluations
Engine
- Engine now shows only project-level spend in project view, so org-wide spend stays in the org settings surface.
- Engine now keeps the Slack issue-alert deck pinned above the scrolling issues list, so the callout stays visible as you browse.
Datasets and experiments
The experiments table now displays loading progress bars showing the number of runs completed and evaluated, and experiments that predate this feature show a placeholder progress bar.- Dashboards now support time series bar and line charts backed by the v2 chart API, so monitored metrics can use the newer chart type.
- Categorical feedback now shows derived percentages in experiment tables, so pass/fail metrics are easier to scan.
Prompts and playground
- Playground now mints OAuth bearers end to end for OAuth-enabled presets, so long-running batches and streams keep working.
Sandboxes
- Sandbox auth proxy now supports GCP auth flows, so sandbox workloads can reach Google APIs through the proxy.
Fixes
- The Engine trial modal no longer shows the rough-math LCU bullet, so the pricing copy is less misleading.
Observability and evaluations
Automations
- Run rule webhook payloads now include a trace deep link for each run, so downstream systems can jump straight back to the trace.
Engine
- Per-workspace Engine spend is now generally available: you can view LCU and USD spend directly on the Engine settings page, including session-level spend.
- The Engine settings page now surfaces additional Engine details in one place.
- You can rotate Engine issue-board webhook signing secrets from both the API and the webhook settings UI.
- The Engine issues list adds a sort option by trace count.
Datasets and experiments
- A new out-of-the-box Assertions evaluator scores outputs against an explicit list of criteria specified in the reference output, and an Assertions rule is auto-attached when you add assertion-style examples to a dataset.
- Evaluator metrics are improved in the experiment detail, comparison, and global experiments tables.
Prompts and playground
- The Playground supports Amazon Bedrock API key authentication, letting you authenticate with a bearer token instead of AWS credentials.
Tracing
- The trace view now shows an unread indicator on a run’s actions menu when the run has reviewer notes you have not seen yet.
- The waterfall view is now full-height with sticky turn headers, so you keep your place while scrolling through long traces.
- Global search now includes context and sandboxes
Deployment
- You can now trigger a LangSmith Deployment from the Studio page.
- LangSmith Deployment now supports deploying Google Agent Development Kit (ADK) agents.
Sandboxes
- Sandbox proxy rules now support configuring AWS authentication, so sandboxes can reach AWS services through the proxy with signed requests.
- Sandboxes can create snapshots from a Dockerfile build source.
Admin and billing
Administration
- Organization admins can now disable personal access token creation from the organization settings page.
Usage and billing
- Granular billable usage now supports filtering and grouping by retention tier, separating long-lived from short-lived traces.
- The Granular Billable Usage page now surfaces LangSmith Deployment usage, including nodes executed, agent runs, and agent uptime, alongside trace usage.
Fixes
- Performance improvements for the loading of large traces.
- Filter values for metadata are now preserved when you reopen a filter dropdown to edit it.
- Dataset creation now uses a multi-select dropdown for choosing CSV fields.
Observability and evaluations
Insights
- The Insights Agent now supports scheduled reports on daily, weekly, or custom cron intervals, so report generation runs without manual triggering. Time ranges compute dynamically, so a “last 24 hours” report always reflects the most recent window when it runs, not when you configured it.
Datasets and experiments
- You can now pin any experiment as a baseline. The pinned experiment stays at the top of the Experiments view and serves as the automatic comparison point for later runs, surfacing performance deltas across every column so improvements and regressions are immediately clear.
Observability and evaluations
Cost tracking
- Cost tracking now extends beyond LLM calls. Submit custom cost metadata for any run, such as an expensive tool call, a third-party API, or a retrieval step, to monitor, debug, and optimize spend across your entire agent stack from a single dashboard.
Tracing
- You can now configure which parts of a trace’s inputs and outputs appear in the tracing table, so teams working with custom trace formats can surface the most relevant fields, reduce clutter, and identify traces that need a closer look faster.
Observability and evaluations
Annotation and human feedback
- New pairwise annotation queues let reviewers compare two runs side by side and choose whether option A is better, option B is better, or the two are equal across rubric items. LangSmith automatically pairs runs between two experiments and manages queues, reviewer assignments, and trace access, so you can run A/B evaluations across agents, prompts, and models, including for subjective dimensions like tone, correctness, usefulness, or style.
Observability and evaluations
Tracing
- LangSmith Fetch, a new command-line tool, brings LangSmith traces directly into your terminal, coding environment, or IDE. Install it with
pip install langsmith-fetch, then retrieve traces with filters such as--limit,--after, and--last-n-minutes, or bulk-export traces and threads to files for analysis, scripting, or dataset creation.
Observability and evaluations
Cost tracking
- Cost tracking now automatically records token usage and derived costs for major model providers, and you can submit custom cost data for tools, retrieval steps, and other operations. Costs appear across trace trees, project stats, and dashboards, with an editable price map for non-standard pricing.
Admin and billing
Administration
- LangSmith is now on the Okta Integration Network, so enterprise teams can provision and deprovision users with SCIM and configure SSO through Okta’s guided setup. See the administration overview for access control options.
Observability and evaluations
Insights
- The Insights Agent is now generally available for Plus and Enterprise plans. It analyzes production traces to surface usage patterns, agent behaviors, and failure modes, with usage-pattern clustering, poor-interaction analysis, and custom grouping and filtering.
Datasets and experiments
- Multi-turn evals measure end-to-end agent conversations across multiple exchanges, scoring semantic intent, semantic outcomes, and agent trajectory, including tool calls and decisions.
Observability and evaluations
Datasets and experiments
- Dataset creation now infers schema automatically from uploaded CSV and JSONL files, supports adding metadata fields during upload, supports column mapping and renaming, and supports bulk additions to existing datasets from new uploads.
Deployment
- LangGraph Platform is now LangSmith Deployment and LangGraph Studio is now LangSmith Studio. LangSmith now spans three services: Observability, Evaluation, and Deployment. Existing deployments, APIs, workflows, pricing, and contracts are unchanged, and no action is required.
Observability and evaluations
Datasets and experiments
- You can now write custom code evaluators in JavaScript in addition to Python, so TypeScript teams can stay in their ecosystem end to end.
Observability and evaluations
Datasets and experiments
- Composite evaluators combine multiple evaluator scores into a single metric using a weighted average or weighted sum, with customizable weights.
Admin and billing
Administration
- You can now create service keys at the organization level, scoped to multiple workspaces or the entire organization, and assign roles, including custom roles, for granular permissions.
Deployment
- LangSmith Deployment now queues revisions automatically, processing each new revision only after the current one finishes to prevent overlapping deployments and conflicts.
Observability and evaluations
Datasets and experiments
- Align Evals provides a playground-like interface for iterating on evaluator prompts and comparing human-graded scores side by side with LLM-generated scores to surface misaligned cases.
Deployment
- LangSmith now links traces to the server logs in LangSmith Deployment, so you can open user and system logs directly from a trace.
Observability and evaluations
Tracing
- Data export now supports scheduled exports of traces, so external systems such as data warehouses, monitoring platforms, and dashboards stay in sync without custom infrastructure.
Deployment
- A new Monitoring tab shows deployment metrics, including CPU and memory usage, API request latency, and active run counts, over a customizable time range.
Observability and evaluations
Datasets and experiments
- You can now create custom views of evaluation results by breaking fields from inputs, outputs, and reference outputs into their own columns, hiding or reordering columns, and adjusting decimal precision on feedback scores.
Admin and billing
Administration
- LangSmith API keys now support expiration dates, so you can scope access for temporary tasks or team members.
Observability and evaluations
Prompts and playground
- The Playground now supports calling built-in tools from OpenAI and Anthropic, such as web search and MCP, so you can verify tool selection and argument passing.
Deployment
- Studio now lets you run agent evaluations in the UI without code, comparing against reference outputs and grading responses with custom criteria.
Observability and evaluations
Cost tracking
- Cost tracking now accounts for cached tokens, multiple token modalities such as text and image, and reasoning tokens, and supports tracking costs for arbitrary token types.
Deployment
- Every agent deployed on LangSmith now exposes its own Model Context Protocol (MCP) endpoint, so the agent can be used as a tool in any client that supports streamable HTTP for MCP, with no custom code or infrastructure.
Admin and billing
Usage and billing
- SaaS customers can now view monthly usage charts that track all billable metrics in one place.
Observability and evaluations
Monitoring and alerting
- Agent observability surfaces tool calls and run stats, including the most-used tools and runs, their latency, and which generate the most errors.
Deployment
- LangGraph Platform, now LangSmith Deployment, reached general availability for deploying and managing long-running, stateful agents at scale, with one-click GitHub-to-production deployment, integrated memory and persistence, scalable APIs, and an agent registry across cloud, hybrid, self-hosted, and developer deployment options.
- Studio v2 runs locally without the desktop app, supports editing prompts and configuration in the UI, integrates with the Playground, and lets you download production traces to debug them locally.
Observability and evaluations
Tracing
- LangSmith now supports multimodal content for images, PDFs, and audio across the playground, annotation queues, and datasets, including attaching files to dataset examples without base64 encoding and visualizing the content in the app.
Observability and evaluations
Prompts and playground
- The Playground now lets you create datasets inline and add examples to existing datasets without leaving the Playground.
Observability and evaluations
Tracing
- LangSmith now has end-to-end native OpenTelemetry support for LangChain and LangGraph applications, including distributed tracing across microservices.
Datasets and experiments
- You can now define evaluators for datasets and tracing projects directly in the UI with no code, including LLM-as-a-judge evaluators with prebuilt templates, customizable prompts, variable mapping, scoring, and few-shot support.
Observability and evaluations
Tracing
- LangSmith now supports tracing OpenAI Agents SDK applications with two lines of code, for step-by-step observability of agent execution and reasoning.
Datasets and experiments
- You can now rename an experiment in the UI, either from the Playground table header after a run or with the pencil icon in the Experiments view.
Observability and evaluations
Datasets and experiments
- You can now group experiment results by metadata to analyze evaluation performance across segments such as user groups or subject areas.
Fixes
- A new ingest-backend service separates trace ingestion from frontend request handling, improving average request processing and high-traffic response times.
Observability and evaluations
Prompts and playground
- The Playground can now use workspace secrets saved in LangSmith, for consistent credential management across environments.
Observability and evaluations
Datasets and experiments
- A new experiment view gives each feedback key its own column and adds filtering, sorting, and a heat map to spot patterns and performance areas.
Deployment
- You can now open LLM runs from Studio in the LangSmith Playground for debugging, visualization, and prompt experimentation within threads.
Observability and evaluations
Prompts and playground
- The Playground adds a streamlined prompt settings UI, a default model configuration, an enhanced tool management modal, and improved side-by-side comparison.
Datasets and experiments
- New Pytest and Vitest integrations let you run evaluations using familiar testing frameworks, with debugging, metrics tracking, and built-in evaluation functions.
Connect these docs to Claude, VSCode, and more via MCP for real-time answers.

