Skip to main content
SAP HANA Cloud Vector Engine is a vector store fully integrated into the SAP HANA Cloud database.

Setup

Install the @sap/hana-langchain external integration package, as well as the other packages used throughout this notebook.

Credentials

Ensure your SAP HANA instance is running. Load credentials from environment variables and create a connection using your preferred HANA client.
Learn more about SAP HANA in What is SAP HANA?.

Initialization

To initialize a HanaDB vector store, you need a database connection and an embedding instance. SAP HANA Cloud Vector Engine supports both external and internal embeddings.

Using external embeddings

Using internal embeddings

Alternatively, you can compute embeddings directly in SAP HANA using its native VECTOR_EMBEDDING() function. If you have internal embedding support available in your TypeScript environment, initialize and pass it to HanaDB similarly. For more information about internal embedding, see the SAP HANA VECTOR_EMBEDDING Function.
Caution: Ensure NLP is enabled in your SAP HANA Cloud instance.
Once you have your connection and embedding instance, create the vector store by passing them to HanaDB along with a table name for storing vectors:

Manage vector store

Once you have created your vector store, we can interact with it by adding and deleting different items.

Add items to vector store

We can add items to our vector store by using the addDocuments function.
Add documents with metadata

Delete items from vector store

Query vector store

Query directly

Performing a simple similarity search with filtering on metadata can be done as follows:
Performing a Maximal Marginal Relevance (MMR) with filtering on metadata search can be done as follows:

Query by turning into retriever

You can also transform the vector store into a retriever for easier usage in your chains.

Distance similarity algorithm

HanaDB supports the following distance similarity algorithms:
  • Cosine Similarity (default)
  • Euclidean Distance (L2)
You can specify the distance strategy when initializing the HanaDB instance by using the distanceStrategy parameter.

Creating a HNSW index

A vector index can significantly speed up top-k nearest neighbor queries for vectors. Users can create a Hierarchical Navigable Small World (HNSW) vector index using the createHnswIndex function. For more information about creating an index at the database level, please refer to the official documentation.
If no other parameters are specified, the default values will be used Default values: m=64, ef_construction=128, ef_search=200 The default index name will be: “<TABLE_NAME>_idx”

Advanced filtering

In addition to the basic value-based filtering capabilities, it is possible to use more advanced filtering. The table below shows the available filter operators.
Filtering with $ne, $gt, $gte, $lt, $lte
Filtering with $between, $in, $nin
Text filtering with $like
Text filtering with $contains
Combined filtering with $and, $or

Usage for retrieval-augmented generation

For guides on how to use this vector store for retrieval-augmented generation (RAG), see the following sections:

Standard tables vs. custom tables with vector data

By default, the embeddings table contains three columns:
  • VEC_TEXT: the document text
  • VEC_META: the document metadata
  • VEC_VECTOR: the embedding vector
Custom tables must have at least three columns that match the semantics of a standard table:
  • An NCLOB/NVARCHAR column for the text/context
  • An NCLOB/NVARCHAR column for the metadata
  • A REAL_VECTOR column for the embedding vector
Additional columns are allowed; ensure they accept NULLs for new document inserts.
Show the columns in table “LANGCHAIN_DEMO_NEW_TABLE”
Show the value of the inserted document in the three columns Since, HANA’s dbapi driver outputs the vector columns in Buffer objects by default, we will create a helper function to convert the function into a list of numbers.
Custom tables must have at least three columns that match the semantics of a standard table
  • A column with type NCLOB or NVARCHAR for the text/context of the embeddings
  • A column with type NCLOB or NVARCHAR for the metadata
  • A column with type REAL_VECTOR or HALF_VECTOR for the embedding vector
The table can contain additional columns. When new Documents are inserted into the table, these additional columns must allow NULL values.
Add another document and perform a similarity search on the custom table.

Filter performance optimization with custom columns

To allow flexible metadata values, all metadata is stored as JSON in the metadata column by default. If some of the used metadata keys and value types are known, they can be stored in additional columns instead by creating the target table with the key names as column names and passing them to the HanaDB constructor via the specificMetadataColumns list. Metadata keys that match those values are copied into the special column during insert. Filters use the special columns instead of the metadata JSON column for keys in the specificMetadataColumns list.
The special columns are completely transparent to the rest of the langchain interface. Everything works as it did before, just more performant.

A simple example

Load the sample document “state_of_the_union.txt” and create chunks from it. First, install the @langchain/textsplitters package:
Add the loaded document chunks to the table. For this example, we delete any previous content from the table which might exist from previous runs.
Perform a query to get the two best-matching document chunks from the ones that were added in the previous step. By default “Cosine Similarity” is used for the search.
Query the same content with “Euclidean Distance”. The results should be the same as with “Cosine Similarity”.

Maximal marginal relevance search (MMR)

Maximal marginal relevance optimizes for similarity to query AND diversity among selected documents. The first 20 (fetch_k) items will be retrieved from the DB. The MMR algorithm will then find the best 2 (k) matches.

Creating a HNSW vector index

A vector index can significantly speed up top-k nearest neighbor queries for vectors. Users can create a Hierarchical Navigable Small World (HNSW) vector index using the createHnswIndex function. For more information about creating an index at the database level, please refer to the official documentation.
Key Points:
  • Similarity Function: The similarity function for the index is cosine similarity by default. If you want to use a different similarity function (e.g., L2 distance), you need to specify it when initializing the HanaDB instance.
  • Default Parameters: In the createHnswIndex function, if the user does not provide custom values for parameters like m, efConstruction, or efSearch, the default values (e.g., m=64, efConstruction=128, efSearch=200) will be used automatically. These values ensure the index is created with reasonable performance without requiring user intervention.