Composite Neuro-Symbolic Search
Documentation status: architecture — see Maturity and evidence.
The current engine supports composite embedding queries that combine several numerical terms while retaining symbolic control over admissible results.
This is useful for queries whose intent cannot be represented by one embedding alone, for example:
find a Document
that is close to Alice
and close to the concept IncidentReport
Here Document may be a hard symbolic constraint while Alice and IncidentReport influence numerical ranking.
Query terms
A composite query contains one or more embedding terms:
(vector, weight)
All vectors used by a composed-vector strategy must have the same dimension.
Weights influence numerical composition/reranking. A zero-weight term contributes nothing.
Candidate generation strategies
The engine separates candidate generation from final scoring.
Composed-vector retrieval
A default strategy builds one weighted linear combination of the query vectors. The result can be L2-normalized before HNSW retrieval.
q = normalize(w1*v1 + w2*v2 + ... + wn*vn)
q -> one HNSW search
This is efficient when a linear composition is meaningful in the embedding space.
Multi-query fusion
The engine can also run one HNSW search per term and fuse the neighborhoods. This avoids requiring the terms to collapse into one geometric direction before retrieval.
Current fusion modes include:
- weighted distance aggregation;
- weighted reciprocal-rank fusion.
The union of per-term neighborhoods is preserved before final truncation, so a strong candidate from one term is not necessarily lost because it was absent from another term's local neighborhood.
Scoring strategies
Candidate generation and candidate scoring are independent. The current engine includes strategies corresponding to:
- HNSW retrieval distance;
- weighted average distance to the individual query terms;
- hybrid scoring that combines query-vector distance and per-term distance.
This separation allows experimentation with ranking without changing the symbolic identity model or the HNSW storage layer.
Hard symbolic filters
A candidate filter is applied after vector hits are mapped back to conceptual values. The current engine includes a Kind-compatibility filter.
This gives an important neuro-symbolic pattern:
soft numerical evidence:
type embedding + subject embedding + other signals
hard symbolic constraint:
candidate must be an instance of Document
Tests verify that a closer incompatible value is rejected and that a compatible conceptual candidate is retained. The numerical model influences which valid candidate is preferred, not what is valid.
Candidate over-fetching
When hard filters or reranking are used, retrieving only K numerical neighbors can be insufficient: some of those candidates may later be rejected.
The composite search therefore separates:
candidate count > final K
The larger candidate pool is generated first, then resolved, filtered, reranked and truncated to the requested result count.
Query-term exclusion
Composite search can exclude the vectors used as query terms from the result set. This is useful when searching for neighbors of known concepts rather than returning the concepts themselves.
Result semantics
A composite result carries both sides of the bridge:
conceptual candidate
vector identity
numerical score/distance
The score is ranking evidence. The conceptual value is the object on which symbolic reasoning and application behavior should operate.
What this does not mean
Composite vector search does not turn similarity into deduction. It does not redefine Kind, subtype relations or H-Logic truth. It is a candidate-generation and ranking mechanism that can be governed by symbolic constraints.