LLM Reference

NV-Embed Models by NVIDIA AI

4 models2024–2025Up to 4k ctx

Last refreshed 2026-05-19. Next refresh: weekly.

Details

ResearcherNVIDIA AI
Models4
Released2024–2025
Max context4k

Links

Website

About

NV-Embed is NVIDIA's family of specialized embedding and reranking models, including NV-EmbedCode for code retrieval and NV-EmbedQA/RerankQA for question-answering tasks. NVIDIA NIM also hosts BAAI BGE models (such as BGE-M3) as first-class retrieval endpoints in its API catalog.

Decision facts

Best fit
embeddingrankingcoding
Capability starting point
NV-EmbedCode 7B v1 with 4k context
Lowest tracked input
Not tracked
Closest related family
NVIDIA Nemotron Nano 12B v2 VL

Current Variants

Use-when guidance is based on each model's tracked capabilities, context window, release date, and replacement status.

4 in view

Use when the workload needs embedding, 4k context, and 7B parameters.

2025-06embedding4k context7B parameters

Use when the workload needs embedding, 4k context, and 1B parameters.

2025-03embedding4k context1B parameters

Use when the workload needs ranking, 4k context, and 1B parameters.

2025-03ranking4k context1B parameters

Use when the workload needs embedding, 512 context, and 1B parameters.

2024-10embedding512 context1B parameters

Release Timeline

3 release groups
2025-06
1 current
NV-EmbedCode 7B v1
embedding4k context7B parameters
Current
2025-03
2 current
Llama 3.2 NV EmbedQA 1B v2
embedding4k context1B parameters
Current
Llama 3.2 NV RerankQA 1B v2
ranking4k context1B parameters
Current
2024-10
1 current
Llama 3.2 NV EmbedQA 1B v1
embedding512 context1B parameters
Current

Specifications(4 models)

NV-Embed model specifications comparison
ModelReleasedContextParameters
NV-EmbedCode 7B v12025-064k7B
Llama 3.2 NV EmbedQA 1B v22025-034k1B
Llama 3.2 NV RerankQA 1B v22025-034k1B
Llama 3.2 NV EmbedQA 1B v12024-105121B

Available From(1 provider)

Popular comparisons in this family