Graph and ontology search
Most search stacks make you choose. Either your documents live in a search engine and your ontology lives in a graph database, or you flatten the ontology into keywords and lose everything that made it an ontology.
Gnarl treats the graph as an index type. Edges are documents, reachability is a query, and a traversal composes with text, filters and vectors in the same request — across the same decentralized fan-out as everything else.
Edges are documents
A graph_edge field holds one edge per document:
{ "parent_of": { "source": 4021, "target": 8817 } }
Node identifiers are integers, which is what makes the adjacency compact: the engine keeps a native CSR (compressed sparse row) structure rather than resolving strings on every hop.
gnarl index create ontology --field 'parent_of:graph_edge'
Traversal
graph_traversal expands reachability from a source node:
{
"query": {
"graph_traversal": {
"field": "parent_of",
"source": 4021,
"hops": 3
}
}
}
Omit hops and it expands until it runs out of graph. max_nodes bounds the
frontier and defaults to 65536.
When the bound is exceeded the query is refused, not truncated. That is
deliberate and it is the same principle as coverage: a traversal
that quietly stopped at the limit would return a confident, wrong answer — a
subgraph presented as the subgraph. Raise max_nodes if you meant it, or narrow
hops if you did not.
The join that matters
Keeping the ontology and the corpus in one index would be a strange demand, so
you don't have to. graph_index traverses a different index's graph, and
node_field restricts the searched index to the nodes it reached:
{
"query": {
"bool": {
"must": [
{ "match": { "body": "ventilator weaning" } }
],
"filter": [
{
"graph_traversal": {
"graph_index": "snomed",
"field": "is_a",
"source": 4021,
"node_field": "concept_id"
}
}
]
}
}
}
Read that as: find documents matching this text whose concept_id is anywhere
beneath concept 4021 in the SNOMED hierarchy. The ontology is maintained in its
own index, by whoever owns it, and the content index only has to carry a numeric
concept id.
In a decentralized network that separation is the point. The ontology can live with the institution that curates it while the documents stay with the operators that produced them, and the query still resolves across both.
Taxonomy mode
Set taxonomy: true on a graph_edge field and the engine precomputes the
transitive closure at flush:
curl -X PUT localhost:8080/v1/indexes/snomed \
-H 'content-type: application/json' \
-d '{"schema":{"fields":{"is_a":{"type":"graph_edge","taxonomy":true}}}}'
gnarl index create --field takes NAME:TYPE only, so it cannot express field
options — a taxonomy field has to be created through the API.
You pay for it once, at write time, and ancestor questions stop being traversals. It also unlocks two queries that only make sense over a DAG with known structure.
Information content
information_content returns the structural information content of a node —
how specific a concept is, derived from its position in the hierarchy rather
than from any corpus. A leaf is informative; a root is not.
{ "query": { "information_content": { "field": "is_a", "source": 8817 } } }
Concept similarity
concept_similarity scores two concepts against each other:
{
"query": {
"concept_similarity": {
"field": "is_a",
"source": 8817,
"target": 9043,
"metric": "lin"
}
}
}
Two metrics, both standard in the ontology literature:
metric | What it measures |
|---|---|
resnik (default) | The information content of the most specific shared ancestor. Asks how much two concepts have in common, in absolute terms. |
lin | Resnik normalized by the information content of both concepts, so it lands in a comparable range regardless of how deep in the hierarchy you are. |
The practical difference: Resnik will tell you two deep, closely-related concepts share a lot; Lin will tell you whether that is a lot relative to how specific they are. Use Lin when you are comparing pairs from different parts of the tree.
Seeing what it did
A traversal has two phases and profile: true separates them:
"profile": {
"graph": {
"plan": "...",
"reachability_ms": 4,
"materialize_ms": 11,
"segment_edges": 214903,
"estimated_hits": 1180
}
}
reachability_ms is the walk; materialize_ms is fetching the documents it
reached. When the first number is small and the second is large, the graph is
not your problem — the result set is.
If you only want the size of a subgraph and not its contents, set size: 0 with
track_total_hits: true and skip materializing entirely.
Limits worth knowing before you design around them
- Node ids are integers. Mapping your own identifiers to them is your job, and the mapping has to be stable across every node that holds the graph.
- Edges are directed. An undirected relationship is two documents.
information_contentandconcept_similarityrequiretaxonomy: true, and refuse rather than silently falling back to an untaxonomized traversal.max_nodesrefuses at the bound. This is not a soft limit.
See also
- Routing and manifests — how a query reaches only the nodes that can answer it
- Query API reference