Skip to content

How do I turn an OWL ontology and RDF data into a property graph?

You keep research data as RDF: an OWL ontology says there are researchers, publications and institutions and how they relate, and a Turtle file lists the actual researchers, papers and institutions. You want the same data in a property graph database, with vertices you can query by type and edges you can traverse.

GraFlo reads the ontology and proposes a manifest: each class becomes a vertex type, each datatype property becomes a vertex property, and each object property becomes an edge. It then reads the instances of each class from the data file and writes the graph.

flowchart LR
    Researcher((Researcher)) -- authorOf --> Publication((Publication))
    Researcher -- affiliatedWith --> Institution((Institution))
    Publication -- cites --> Publication

What you need

  • GraFlo installed (pip install graflo). The RDF libraries come with it.
  • A running ArangoDB. The repository ships a container for it; see docker/README.md.

The data

data/ontology.ttl declares three classes, seven datatype properties and three object properties. An excerpt:

ex:Researcher  a owl:Class ;
    rdfs:label "Researcher" .

ex:fullName     a owl:DatatypeProperty ;
    rdfs:domain ex:Researcher ;
    rdfs:range  xsd:string .

ex:authorOf  a owl:ObjectProperty ;
    rdfs:domain ex:Researcher ;
    rdfs:range  ex:Publication .

data/data.ttl holds the instances: 3 institutions, 4 researchers and 4 publications. Each researcher is affiliated with one institution and authors one publication; publications 2 to 4 each cite the one before.

ex:alice  a ex:Researcher ;
    ex:fullName       "Alice Smith" ;
    ex:orcid          "0000-0001-0001-0001" ;
    ex:affiliatedWith ex:mit ;
    ex:authorOf       ex:paper1 .

Steps

The steps are the parts of ingest.py.

1. Infer the schema from the ontology

conn_conf = ArangoConfig.from_docker_env()
engine = GraphEngine(target_db_flavor=conn_conf.connection_type)
schema, ingestion_model = engine.infer_schema_from_rdf(
    source=ONTOLOGY_FILE, schema_name="academic_kg"
)

infer_schema_from_rdf returns the schema and one resource per class. A resource is the recipe that turns one kind of record into vertices and edges.

2. Look at what was inferred

The script saves the schema to generated-manifest.yaml:

core_schema:
    edge_config:
        edges:
        -   relation: affiliatedWith
            source: Researcher
            target: Institution
        # ... authorOf, cites
    vertex_config:
        vertices:
        # ... Institution, Publication
        -   identity:
            -   _uri
            name: Researcher
            properties:
            -   name: _key
            -   name: _uri
            -   name: fullName
            -   name: orcid
  • Each vertex type gets two properties besides the datatype properties: _uri, the full IRI of the instance (http://example.org/alice), which is its identity, and _key, the last part of the IRI (alice).
  • Each object property becomes an edge from its rdfs:domain to its rdfs:range, named after the property. A property with several ranges gives one edge per range, and each object is linked under the classes it has: the connector lists the property in typed_objects, and the source adds a field <property>@<Class> holding the objects of that class.
  • Anonymous classes (owl:unionOf and the like) are skipped. Two classes with the same local name would share a vertex type, so inference refuses them.

The resource for Researcher makes a Researcher vertex, then reads the IRI in authorOf into a Publication vertex and adds the authorOf edge; the same for affiliatedWith. A researcher with several authorOf papers gets one edge per paper. Classes, properties and edges are listed in IRI order, so inferring again from the same ontology gives the same file.

3. Read the instances of each class from the data file

bindings = Bindings()
for resource in ingestion_model.resources:
    connector = SparqlConnector(rdf_class=NAMESPACE + resource.name, rdf_file=DATA_FILE)
    bindings.add_connector(connector)
    bindings.bind_resource(resource.name, connector)

A connector says where the records of a resource come from. Each one here reads the instances of one class from data.ttl and gives the resource one record per instance: its IRI and one field per property.

4. Ingest

engine.define_and_ingest(
    manifest=GraphManifest(
        graph_schema=schema, ingestion_model=ingestion_model, bindings=bindings
    ),
    target_db_config=conn_conf,
    ingestion_params=IngestionParams(clear_data=True),
    recreate_schema=True,
)

5. Run it

cd examples/10-infer-from-rdf
uv run python ingest.py

What you should see

The database holds:

Count Why
Researcher vertices 4 One per instance of ex:Researcher
Publication vertices 4 One per instance; a cited or authored paper is matched by its IRI, not added again
Institution vertices 3 Dave and Bob share ETH Zürich
authorOf edges 4 One per researcher
affiliatedWith edges 4 One per researcher
cites edges 3 Papers 2, 3 and 4 each cite one paper

Also possible

  • To read the data from a SPARQL endpoint instead of a file, give each SparqlConnector an endpoint_url in place of rdf_file.
  • When the ontology and the instances are in one file, engine.create_bindings_from_rdf(path) builds these connectors for you, all reading that file.
  • A blank node ([ ex:value 3 ; ex:unit "nm" ]) has no IRI, so it is keyed on its content: _uri is _: followed by a digest of its triples, the same on every read. Reading the file again updates the same vertices.
  • IRIs joined by owl:sameAs are read as one instance under the smallest IRI, with the others listed in a _same_as field; declare _same_as on the vertex to store it. SparqlConnector(..., same_as="keep") reads the statements as an ordinary sameAs field instead.

Files

The example lives in examples/10-infer-from-rdf.

generated-manifest.yaml
core_schema:
    edge_config:
        edges:
        -   relation: affiliatedWith
            source: Researcher
            target: Institution
        -   relation: authorOf
            source: Researcher
            target: Publication
        -   relation: cites
            source: Publication
            target: Publication
    vertex_config:
        vertices:
        -   identity:
            -   _uri
            name: Institution
            properties:
            -   name: _key
            -   name: _uri
            -   name: country
            -   name: instName
        -   identity:
            -   _uri
            name: Publication
            properties:
            -   name: _key
            -   name: _uri
            -   name: doi
            -   name: title
            -   name: year
        -   identity:
            -   _uri
            name: Researcher
            properties:
            -   name: _key
            -   name: _uri
            -   name: fullName
            -   name: orcid
metadata:
    name: academic_kg
ingest.py
"""How do I turn an OWL ontology and RDF data into a property graph?

Infers a schema and one resource per class from ``data/ontology.ttl``, saves the
schema to ``generated-manifest.yaml``, connects each resource to the instances of
its class in ``data/data.ttl`` and writes the graph to ArangoDB. Run it from this
directory:

    uv run python ingest.py
"""

from pathlib import Path

from suthing import FileHandle

from graflo import GraphManifest
from graflo.architecture.contract.bindings import Bindings, SparqlConnector
from graflo.connections import ArangoConfig
from graflo.hq import GraphEngine, IngestionParams

EXAMPLE_DIR = Path(__file__).resolve().parent
ONTOLOGY_FILE = EXAMPLE_DIR / "data" / "ontology.ttl"
DATA_FILE = EXAMPLE_DIR / "data" / "data.ttl"
NAMESPACE = "http://example.org/"

# 1. Infer the schema and the resources from the ontology.
conn_conf = ArangoConfig.from_docker_env()
engine = GraphEngine(target_db_flavor=conn_conf.connection_type)
schema, ingestion_model = engine.infer_schema_from_rdf(
    source=ONTOLOGY_FILE, schema_name="academic_kg"
)

# 2. Save what was inferred, to read or edit.
FileHandle.dump(
    schema.model_dump(exclude_defaults=True), EXAMPLE_DIR / "generated-manifest.yaml"
)

# 3. Read the instances of each class from the data file.
bindings = Bindings()
for resource in ingestion_model.resources:
    connector = SparqlConnector(rdf_class=NAMESPACE + resource.name, rdf_file=DATA_FILE)
    bindings.add_connector(connector)
    bindings.bind_resource(resource.name, connector)

# 4. Write the graph.
engine.define_and_ingest(
    manifest=GraphManifest(
        graph_schema=schema, ingestion_model=ingestion_model, bindings=bindings
    ),
    target_db_config=conn_conf,
    ingestion_params=IngestionParams(clear_data=True),
    recreate_schema=True,
)