minishlab/tokenlearn-cornstack-docs-coderankembed-v2
Viewer • Updated • 600k • 33 • 1
How to use Takara-DS1/miru-codev3-distill-fuse with Model2Vec:
from model2vec import StaticModel
model = StaticModel.from_pretrained("Takara-DS1/miru-codev3-distill-fuse")
embeddings = model.encode(["It's dangerous to go alone!", "It's a secret to everybody."])
print(embeddings.shape)Owned distill_fuse static code embedding bag (256-d, Model2Vec) — branch B of miru-codev3-dual.
Distilled from teachers 0.3 · tokenlearn + 0.7 · potion-code-16M-v2 on CornStack texts (code_static.train_distill_fuse). Potion is only a training teacher — not required at inference.
CoIR / MTEB NDCG@10 (×100), dense retrieval only. Self-reported.
from model2vec import StaticModel
model = StaticModel.from_pretrained("Takara-DS1/miru-codev3-distill-fuse")
emb = model.encode(["def add(a, b): return a + b"])
For the 512-d dual (tokenlearn ⊕ this bag), use Takara-DS1/miru-codev3-dual.
See the dual repo reproduce/REPRODUCE.md (step 2: train_distill_fuse --alpha 0.3).
MIT