How it works
An embedding model turns an input into a vector, a list of a few hundred to a few thousand numbers. Inputs with similar meaning land close together, and closeness is measured with cosine similarity or a dot product. This powers semantic search, 'more like this' recommendations, grouping similar support tickets, spotting duplicates and the retrieval step of RAG.
Hosted models from OpenAI, Google, Cohere and Voyage charge per token, while open models such as BGE, E5 or MiniLM run locally, even in a browser with ONNX Runtime. Documents and queries must be embedded with the same model, and switching models means re-embedding everything. The vectors are stored in a vector index, often inside a database you already run.
Embeddings pricing
Pay as you go
Hosted APIs cost roughly $0.02 to $0.20 per million tokens (OpenAI's small model at the low end). Open models run locally for free.
Embeddings pricing page (opens in a new tab)Approximate, checked September 2026.What the other tools cost
Related terms
More in AI and LLMs
Search and memory