An embedding represents an item as a list of numbers. Useful models arrange these representations so some kinds of similarity become measurable. Text with related meaning can end up nearby even when it uses different words. This is how “forgot my password” can connect to “account recovery” without matching a phrase exactly.
The map is learned
The geometry depends on the model, its training objective, and the data. Similarity is not a natural law hiding in the numbers. A representation designed for sentence retrieval may behave differently from one designed for another task.
The Sentence-BERT research addresses efficient sentence representations for similarity. It is a helpful foundation for understanding why comparing pooled representations can be more practical than running a costly pairwise model over every combination.
Cosine is an angle, not a lie detector
Cosine similarity compares direction. Other systems use dot products or distance measures. Whichever score you use, it measures a property of the representation. High similarity does not prove two claims are equally true, two documents are duplicates, or two people mean the same thing.
“This product is safe” and “This product is not safe” may share much of their language and topic. Negation can be crucial to the task while a representation still places the statements near each other. Retrieve candidates, then check the distinction that matters.
Search is a workflow
Embed documents, embed a query with a compatible model, retrieve likely matches, and inspect or rerank them. Keep source text and metadata alongside the vectors. The retrieved vector is not the document, and readers need the document.
Exact identifiers often benefit from lexical search. Combining lexical and vector retrieval can avoid semantic matches that ignore a serial number. The DPR paper is one reference for dense passage retrieval; it does not establish dense methods as universally optimal.
A score needs local calibration
Do not borrow a similarity threshold from an unrelated project. Collect representative positive and negative pairs. Inspect what happens near the threshold. Test languages, abbreviations, product names, and short fragments that appear in your real material.
If you change the embedding model, dimensions and geometry may change. Rebuild or migrate the index carefully and avoid mixing incompatible representations. Version the model and preprocessing so a future maintainer can reproduce the search behavior.
Vectors still carry data risk
An embedding should not be assumed to anonymize sensitive source text. Treat indexes as potentially sensitive and apply access controls. Retrieval filters should preserve the same document permissions as the source system.
Clustering can reveal recurring topics; duplicate detection can reduce review load; semantic search can improve discovery. Each is a different task with different errors. The best use is often to help a human find the right evidence faster.
Our word cloud uses editorial tags, not embeddings. It makes navigation playful without claiming semantic magic. For a real retrieval architecture, read RAG. A map is useful when you know what it maps.
KEEP EXPLORING
Spot an error? See our corrections channel and editorial policy.