Cookie settings

We use cookies to deliver and improve our services, analyze site usage, and if you agree, to customize or personalize your experience and market our services to you. You can read our Cookie Policy here.

Claude Platform Docs
MessagesModel capabilities

Embeddings

Text embeddings are numerical representations of text that enable measuring semantic similarity. This guide introduces embeddings, their applications, and how to use embedding models for tasks like search, recommendations, and anomaly detection.

Before implementing embeddings

When selecting an embeddings provider, there are several factors you can consider depending on your needs and preferences:

  • Dataset size & domain specificity: size of the model training dataset and its relevance to the domain you want to embed. Larger or more domain-specific data generally produces better in-domain embeddings
  • Inference performance: embedding lookup speed and end-to-end latency. This is a particularly important consideration for large scale production deployments
  • Customization: options for continued training on private data, or specialization of models for very specific domains. This can improve performance on unique vocabularies

How to get embeddings with Anthropic

Anthropic does not offer its own embedding model. One embeddings provider with a wide variety of models and capabilities is Voyage AI by MongoDB.

Voyage AI makes embedding models and rerankers. Its embedding models include general-purpose, multimodal, contextualized, and domain-specific models.

The rest of this guide is for Voyage AI, but you should assess a variety of embeddings vendors to find the best fit for your specific use case.

Available models

Voyage AI offers the following text embedding models:

Latest generation

ModelContext lengthEmbedding dimensionDescription
voyage-4-large32,0001024 (default), 256, 512, 2048The best general-purpose and multilingual retrieval quality. See the Voyage 4 blog post for details.
voyage-432,0001024 (default), 256, 512, 2048Optimized for general-purpose and multilingual retrieval quality. Balances quality and efficiency. See the Voyage 4 blog post for details.
voyage-4-lite32,0001024 (default), 256, 512, 2048Optimized for latency and cost. See the Voyage 4 blog post for details.
voyage-code-432,0001024 (default), 256, 512, 2048Optimized for code retrieval and agentic coding applications. See the voyage-code-4 blog post for details.
voyage-4-nano32,0002048 (default), 256, 512, 1024Open-weight model (Apache 2.0 license) that you download from Hugging Face and run yourself. Not available through the Atlas Embedding and Reranking API. See the Voyage 4 blog post for details.

Previous generation

For each model's lifecycle status and recommended replacement, see Model deprecations, lifecycle states, and support in the MongoDB documentation.

ModelContext lengthEmbedding dimensionDescription
voyage-3-large32,0001024 (default), 256, 512, 2048Previous generation of voyage-4-large. See the voyage-3-large blog post for details.
voyage-3.532,0001024 (default), 256, 512, 2048Previous generation of voyage-4. See the voyage-3.5 blog post for details.
voyage-3.5-lite32,0001024 (default), 256, 512, 2048Previous generation of voyage-4-lite. See the voyage-3.5 blog post for details.
voyage-code-332,0001024 (default), 256, 512, 2048Previous generation of voyage-code-4. See the voyage-code-3 blog post for details.
voyage-finance-232,0001024Optimized for finance retrieval and RAG. See the voyage-finance-2 blog post for details.
voyage-law-216,0001024Optimized for legal retrieval and RAG. See the voyage-law-2 blog post for details.

Additionally, Voyage AI offers the following multimodal embedding models. Call these models with multimodal_embed() instead of embed():

ModelContext lengthEmbedding dimensionDescription
voyage-multimodal-3.532,0001024 (default), 256, 512, 2048Rich multimodal embedding model that can vectorize interleaved text, images, and videos. Includes video support as the first production-grade video embedding model. See the voyage-multimodal-3.5 blog post for details.
voyage-multimodal-332,0001024Previous generation of voyage-multimodal-3.5. Vectorizes interleaved text and content-rich images, such as screenshots of PDFs, slides, tables, figures, and more. See the voyage-multimodal-3 blog post for details.

The following contextualized chunk embedding models produce chunk-level vectors that capture full document context without manual metadata augmentation. Call these models with contextualized_embed() instead of embed():

ModelContext lengthEmbedding dimensionDescription
voyage-context-4120,0001024 (default), 256, 512, 2048Contextualized chunk embeddings optimized for general-purpose and multilingual retrieval quality. See the voyage-context-4 blog post for details.
voyage-context-3120,0001024 (default), 256, 512, 2048Previous generation of voyage-context-4. See the voyage-context-3 blog post for details.

The 120,000-token limit applies when you set enable_auto_chunking to true. Otherwise, the total number of tokens across all inputs can't exceed 32,000.

Voyage AI also offers rerankers, which take a query and a list of documents and return them ranked by relevance to the query. Call these models with rerank():

ModelContext lengthDescription
rerank-332,000Highest accuracy. Recommended for most applications. See the rerank-3 blog post for details.
rerank-3-lite32,000Optimized for latency and cost. See the rerank-3 blog post for details.
rerank-2.532,000Previous generation of rerank-3. See the rerank-2.5 blog post for details.
rerank-2.5-lite32,000Previous generation of rerank-3-lite. See the rerank-2.5 blog post for details.

Need help deciding which model to use? See Voyage AI embedding and reranking models overview in the MongoDB documentation.

Getting started with Voyage AI

To access Voyage AI models, create a model API key in MongoDB Atlas:

  1. Sign up for a MongoDB Atlas account, or log in.
  2. In your Atlas project, select AI Model APIs in the navigation bar, click Create model API key, name the key, and click Create.
  3. Set the API key as an environment variable for convenience:
export VOYAGE_API_KEY=""

For more detail, see Voyage AI quick start in the MongoDB documentation.

You can obtain the embeddings by either using the official voyageai Python package or HTTP requests, as described in the following sections. Voyage AI also has an official TypeScript client. To use it with a model API key from Atlas, set its environment option to https://ai.mongodb.com/v1, as described in TypeScript client in the MongoDB documentation.

Voyage AI Python library

Install the voyageai package using the following command. To use a model API key from Atlas, you need version 0.3.7 or later.

pip install -U voyageai

Then, you can create a client object and start using it to embed your texts:

import voyageai

vo = voyageai.Client()
# This will automatically use the environment variable VOYAGE_API_KEY.
# Alternatively, you can use vo = voyageai.Client(api_key="")

texts = ["Sample text 1", "Sample text 2"]

result = vo.embed(texts, model="voyage-4", input_type="document")
print(result.embeddings[0])
print(result.embeddings[1])

result.embeddings is a list of two embedding vectors, each containing 1024 floating-point numbers. After running the preceding code, the two embeddings are printed on the screen:

[-0.013131560757756233, 0.019828535616397858, ...]   # embedding for "Sample text 1"
[-0.0069352793507277966, 0.020878976210951805, ...]  # embedding for "Sample text 2"

When creating the embeddings, you can specify a few other arguments to the embed() function.

For more information on the Python package, see Accessing Voyage AI models in the MongoDB documentation.

Voyage AI HTTP API

You can also get embeddings by sending HTTP requests to the Atlas Embedding and Reranking API. For example, you can send an HTTP request through the curl command in a terminal:

cURL
curl https://ai.mongodb.com/v1/embeddings \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $VOYAGE_API_KEY" \
  -d '{
    "input": ["Sample text 1", "Sample text 2"],
    "model": "voyage-4",
    "input_type": "document"
  }'

The response you would get is a JSON object containing the embeddings and the token usage:

{
  "object": "list",
  "data": [
    {
      "object": "embedding",
      "embedding": [-0.013131560757756233, 0.019828535616397858 /* ... */],
      "index": 0
    },
    {
      "object": "embedding",
      "embedding": [-0.0069352793507277966, 0.020878976210951805 /* ... */],
      "index": 1
    }
  ],
  "model": "voyage-4",
  "usage": {
    "total_tokens": 10
  }
}

Model API keys from Atlas work with ai.mongodb.com, except keys scoped to a geography, which use that geography's endpoint. If you have an API key from the Voyage AI platform instead, see Migrate your applications to use the Atlas Embedding and Reranking API in the MongoDB documentation.

For the full request and response reference, see Create text embeddings in the Atlas Embedding and Reranking API documentation.

AWS Marketplace

Voyage AI models are also available on AWS Marketplace through MongoDB's seller profile. For instructions, see Deploy Voyage AI models using AWS Marketplace in the MongoDB documentation.

Quickstart example

The following brief example shows how to use embeddings.

Suppose you have a small corpus of six documents to retrieve from

documents = [
    "The Mediterranean diet emphasizes fish, olive oil, and vegetables, believed to reduce chronic diseases.",
    "Photosynthesis in plants converts light energy into glucose and produces essential oxygen.",
    "20th-century innovations, from radios to smartphones, centered on electronic advancements.",
    "Rivers provide water, irrigation, and habitat for aquatic species, vital for ecosystems.",
    "Apple's conference call to discuss fourth fiscal quarter results and business updates is scheduled for Thursday, November 2, 2023 at 2:00 p.m. PT / 5:00 p.m. ET.",
    "Shakespeare's works, like 'Hamlet' and 'A Midsummer Night's Dream,' endure in literature.",
]

First, use Voyage AI to convert each document into an embedding vector.

import voyageai

vo = voyageai.Client()

# Embed the documents
doc_embds = vo.embed(documents, model="voyage-4", input_type="document").embeddings

The embeddings allow you to do semantic search / retrieval in the vector space. Given an example query,

query = "When is Apple's conference call scheduled?"

Next, convert it into an embedding and conduct a nearest neighbor search to find the most relevant document based on the distance in the embedding space.

import numpy as np

# Embed the query
query_embd = vo.embed([query], model="voyage-4", input_type="query").embeddings[0]

# Compute the similarity
# Voyage AI embeddings are normalized to length 1, so dot-product
# and cosine similarity are the same.
similarities = np.dot(doc_embds, query_embd)

retrieved_id = np.argmax(similarities)
print(documents[retrieved_id])

Note that input_type="document" and input_type="query" are used for embedding the document and query, respectively. For more about input_type, see When and how should I use the input_type parameter? in the FAQ.

The output is the fifth document, which is indeed the most relevant to the query:

Apple's conference call to discuss fourth fiscal quarter results and business updates is scheduled for Thursday, November 2, 2023 at 2:00 p.m. PT / 5:00 p.m. ET.

If you are looking for a detailed set of recipes on how to do RAG with embeddings, including vector databases, check out the RAG recipe.

FAQ

Pricing

For the most up-to-date pricing details, see Model pricing in the MongoDB documentation.

Was this page helpful?