Using Llama 3.2 NV EmbedQA 1B v1 on NVIDIA NIM
Implementation guide · NV-Embed · NVIDIA AI
ServerlessOpen Weights
NVIDIA NIM exposes Llama 3.2 NV EmbedQA 1B v1 through model ID nvidia/llama-3.2-nv-embedqa-1b-v1. Use the setup steps, sourced pricing, capabilities, and official provider links below to validate this route before deployment.
Last refreshed 2026-05-06. Next refresh: weekly.
Quick Start
- 1
- 2Use the NVIDIA NIM SDK or REST API to call
nvidia/llama-3.2-nv-embedqa-1b-v1— see the documentation for request format. - 3
Code Examples
Pricing on NVIDIA NIM
Capabilities
No model capability flags are currently sourced.
About Llama 3.2 NV EmbedQA 1B v1
NVIDIA multilingual embedding model for question-answering retrieval, based on Llama 3.2. Outputs 2048-dimensional embeddings and supports 26 languages with 512-token context. Superseded by v2, which extends context to 4K tokens.
Model Specs
Released2024-10-08
Parameters1B
Context512
ArchitectureEncoder Only