mteb/twentynewsgroups-clustering
Viewer • Updated • 10 • 17.2k • 1
How to use Y-Research-Group/CSR-NV_Embed_v2-Clustering-Biorxiv_TwentyNews with sentence-transformers:
from sentence_transformers import SparseEncoder
model = SparseEncoder("Y-Research-Group/CSR-NV_Embed_v2-Clustering-Biorxiv_TwentyNews", trust_remote_code=True)
queries = ["Which planet is known as the Red Planet?"]
documents = [
"Venus is often called Earth's twin because of its similar size and proximity.",
"Mars, known for its reddish appearance, is often referred to as the Red Planet.",
"Jupiter, the largest planet in our solar system, has a prominent red spot.",
]
query_embeddings = model.encode_query(queries)
document_embeddings = model.encode_document(documents)
similarities = model.similarity(query_embeddings, document_embeddings)
print(similarities)How to use Y-Research-Group/CSR-NV_Embed_v2-Clustering-Biorxiv_TwentyNews with Transformers:
# Use a pipeline as a high-level helper
from transformers import pipeline
pipe = pipeline("text-classification", model="Y-Research-Group/CSR-NV_Embed_v2-Clustering-Biorxiv_TwentyNews", trust_remote_code=True) # pip install -U transformers accelerate
# Load model directly
from transformers import AutoModel
model = AutoModel.from_pretrained("Y-Research-Group/CSR-NV_Embed_v2-Clustering-Biorxiv_TwentyNews", trust_remote_code=True, device_map="auto")For more details, including benchmark evaluation, hardware requirements, and inference performance, please refer to our Github.
📌 Tip: For NV-Embed-V2, using Transformers versions later than 4.47.0 may lead to performance degradation, as model_type=bidir_mistral in config.json is no longer supported.
We recommend using Transformers 4.47.0.
You can evaluate this model loaded by Sentence Transformers with the following code snippet:
import mteb
from sentence_transformers import SparseEncoder
model = SparseEncoder(
"CSR-NV_Embed_v2-Clustering-Biorxiv_TwentyNews",
trust_remote_code=True
)
model.prompts = {
"BiorxivClusteringP2P.v2": "Instruct: Identify the main category of Biorxiv papers based on the titles and abstracts\nQuery:",
"BiorxivClusteringS2S.v2": "Instruct: Identify the main category of Biorxiv papers based on the titles\nQuery:",
"TwentyNewsgroupsClustering": "Instruct: Identify the topic or theme of the given news articles\nQuery:"
}
task = mteb.get_tasks(tasks=["BiorxivClusteringP2P.v2", "BiorxivClusteringS2S.v2", "TwentyNewsgroupsClustering"])
evaluation = mteb.MTEB(tasks=task)
evaluation.run(
model,
eval_splits=["test"],
output_folder="./results/clustering",
show_progress_bar=True
encode_kwargs={"convert_to_sparse_tensor": False, "batch_size": 8},
) # MTEB don't support sparse tensors yet, so we need to convert to dense tensors
@inproceedings{wenbeyond,
title={Beyond Matryoshka: Revisiting Sparse Coding for Adaptive Representation},
author={Wen, Tiansheng and Wang, Yifei and Zeng, Zequn and Peng, Zhong and Su, Yudi and Liu, Xinyang and Chen, Bo and Liu, Hongwei and Jegelka, Stefanie and You, Chenyu},
booktitle={Forty-second International Conference on Machine Learning}
}
Base model
nvidia/NV-Embed-v2