Instructions to use jinaai/jina-embeddings-v3 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use jinaai/jina-embeddings-v3 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="jinaai/jina-embeddings-v3", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("jinaai/jina-embeddings-v3", trust_remote_code=True, device_map="auto") - sentence-transformers
How to use jinaai/jina-embeddings-v3 with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("jinaai/jina-embeddings-v3", trust_remote_code=True) sentences = [ "The weather is lovely today.", "It's so sunny outside!", "He drove to the stadium." ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [3, 3] - Notebooks
- Google Colab
- Kaggle
How to serve a jina-embeddings-v3 classification task using onnxruntime
Can you provide more detailed sample code? The example from README is a bit unclear, especially the output format is not explained.
I got a output, with shape ([batch_size, seq_len, 1024], [batch_size, 1024]). My questions are:
- How can I get the each label of each input text?
- How can I get the embedding representation of each input text?
Hi @luozhouyang
How can I get the embedding representation of each input text?
You should apply mean pooling to the token outputs ([batch_size, seq_len, 1024]). You can use this function:
https://huggingface.co/jinaai/xlm-roberta-flash-implementation/blob/12700ba4972d9e900313a85ae855f5a76fb9500e/modeling_xlm_roberta.py#L630
How can I get the each label of each input text?
If you want to get labels, you should train a classifier on top, or use some kind of a zero-shot classification technique. jina-embeddings-v3 is only responsible for the embedding and does not directly output a label.
@jupyterjazz Thanks!
Can I get the embedding representations of both input text and labels using jina-embeddings-v3 (with task=classification), and then compute the cosine similarities as the confidence score?