Skip to main content

Graph and ontology search

Most search stacks make you choose. Either your documents live in a search engine and your ontology lives in a graph database, or you flatten the ontology into keywords and lose everything that made it an ontology.

Gnarl treats the graph as an index type. Edges are documents, reachability is a query, and a traversal composes with text, filters and vectors in the same request — across the same decentralized fan-out as everything else.

Edges are documents​

A graph_edge field holds one edge per document:

{ "parent_of": { "source": 4021, "target": 8817 } }

Node identifiers are integers, which is what makes the adjacency compact: the engine keeps a native CSR (compressed sparse row) structure rather than resolving strings on every hop.

gnarl index create ontology --field 'parent_of:graph_edge'

Traversal​

graph_traversal expands reachability from a source node:

{
"query": {
"graph_traversal": {
"field": "parent_of",
"source": 4021,
"hops": 3
}
}
}

Omit hops and it expands until it runs out of graph. max_nodes bounds the frontier and defaults to 65536.

When the bound is exceeded the query is refused, not truncated. That is deliberate and it is the same principle as coverage: a traversal that quietly stopped at the limit would return a confident, wrong answer — a subgraph presented as the subgraph. Raise max_nodes if you meant it, or narrow hops if you did not.

The join that matters​

Keeping the ontology and the corpus in one index would be a strange demand, so you don't have to. graph_index traverses a different index's graph, and node_field restricts the searched index to the nodes it reached:

{
"query": {
"bool": {
"must": [
{ "match": { "body": "ventilator weaning" } }
],
"filter": [
{
"graph_traversal": {
"graph_index": "snomed",
"field": "is_a",
"source": 4021,
"node_field": "concept_id"
}
}
]
}
}
}

Read that as: find documents matching this text whose concept_id is anywhere beneath concept 4021 in the SNOMED hierarchy. The ontology is maintained in its own index, by whoever owns it, and the content index only has to carry a numeric concept id.

In a decentralized network that separation is the point. The ontology can live with the institution that curates it while the documents stay with the operators that produced them, and the query still resolves across both.

Taxonomy mode​

Set taxonomy: true on a graph_edge field and the engine precomputes the transitive closure at flush:

curl -X PUT localhost:8080/v1/indexes/snomed \
-H 'content-type: application/json' \
-d '{"schema":{"fields":{"is_a":{"type":"graph_edge","taxonomy":true}}}}'

gnarl index create --field takes NAME:TYPE only, so it cannot express field options — a taxonomy field has to be created through the API.

You pay for it once, at write time, and ancestor questions stop being traversals. It also unlocks two queries that only make sense over a DAG with known structure.

Information content​

information_content returns the structural information content of a node — how specific a concept is, derived from its position in the hierarchy rather than from any corpus. A leaf is informative; a root is not.

{ "query": { "information_content": { "field": "is_a", "source": 8817 } } }

Concept similarity​

concept_similarity scores two concepts against each other:

{
"query": {
"concept_similarity": {
"field": "is_a",
"source": 8817,
"target": 9043,
"metric": "lin"
}
}
}

Two metrics, both standard in the ontology literature:

metricWhat it measures
resnik (default)The information content of the most specific shared ancestor. Asks how much two concepts have in common, in absolute terms.
linResnik normalized by the information content of both concepts, so it lands in a comparable range regardless of how deep in the hierarchy you are.

The practical difference: Resnik will tell you two deep, closely-related concepts share a lot; Lin will tell you whether that is a lot relative to how specific they are. Use Lin when you are comparing pairs from different parts of the tree.

Seeing what it did​

A traversal has two phases and profile: true separates them:

"profile": {
"graph": {
"plan": "...",
"reachability_ms": 4,
"materialize_ms": 11,
"segment_edges": 214903,
"estimated_hits": 1180
}
}

reachability_ms is the walk; materialize_ms is fetching the documents it reached. When the first number is small and the second is large, the graph is not your problem — the result set is.

If you only want the size of a subgraph and not its contents, set size: 0 with track_total_hits: true and skip materializing entirely.

Limits worth knowing before you design around them​

  • Node ids are integers. Mapping your own identifiers to them is your job, and the mapping has to be stable across every node that holds the graph.
  • Edges are directed. An undirected relationship is two documents.
  • information_content and concept_similarity require taxonomy: true, and refuse rather than silently falling back to an untaxonomized traversal.
  • max_nodes refuses at the bound. This is not a soft limit.

See also​