How do I turn an OWL ontology and RDF data into a property graph?¶
You keep research data as RDF: an OWL ontology says there are researchers, publications and institutions and how they relate, and a Turtle file lists the actual researchers, papers and institutions. You want the same data in a property graph database, with vertices you can query by type and edges you can traverse.
GraFlo reads the ontology and proposes a manifest: each class becomes a vertex type, each datatype property becomes a vertex property, and each object property becomes an edge. It then reads the instances of each class from the data file and writes the graph.
flowchart LR
Researcher((Researcher)) -- authorOf --> Publication((Publication))
Researcher -- affiliatedWith --> Institution((Institution))
Publication -- cites --> Publication
What you need¶
- GraFlo installed (
pip install graflo). The RDF libraries come with it. - A running ArangoDB. The repository ships a container for it; see
docker/README.md.
The data¶
data/ontology.ttl declares three classes, seven datatype
properties and three object properties. An excerpt:
ex:Researcher a owl:Class ;
rdfs:label "Researcher" .
ex:fullName a owl:DatatypeProperty ;
rdfs:domain ex:Researcher ;
rdfs:range xsd:string .
ex:authorOf a owl:ObjectProperty ;
rdfs:domain ex:Researcher ;
rdfs:range ex:Publication .
data/data.ttl holds the instances: 3 institutions, 4
researchers and 4 publications. Each researcher is affiliated with one
institution and authors one publication; publications 2 to 4 each cite the one
before.
ex:alice a ex:Researcher ;
ex:fullName "Alice Smith" ;
ex:orcid "0000-0001-0001-0001" ;
ex:affiliatedWith ex:mit ;
ex:authorOf ex:paper1 .
Steps¶
The steps are the parts of ingest.py.
1. Infer the schema from the ontology¶
conn_conf = ArangoConfig.from_docker_env()
engine = GraphEngine(target_db_flavor=conn_conf.connection_type)
schema, ingestion_model = engine.infer_schema_from_rdf(
source=ONTOLOGY_FILE, schema_name="academic_kg"
)
infer_schema_from_rdf returns the schema and one resource per class. A
resource is the recipe that turns one kind of record into vertices and edges.
2. Look at what was inferred¶
The script saves the schema to generated-manifest.yaml:
core_schema:
edge_config:
edges:
- relation: affiliatedWith
source: Researcher
target: Institution
# ... authorOf, cites
vertex_config:
vertices:
# ... Institution, Publication
- identity:
- _uri
name: Researcher
properties:
- name: _key
- name: _uri
- name: fullName
- name: orcid
- Each vertex type gets two properties besides the datatype properties:
_uri, the full IRI of the instance (http://example.org/alice), which is its identity, and_key, the last part of the IRI (alice). - Each object property becomes an edge from its
rdfs:domainto itsrdfs:range, named after the property. A property with several ranges gives one edge per range, and each object is linked under the classes it has: the connector lists the property intyped_objects, and the source adds a field<property>@<Class>holding the objects of that class. - Anonymous classes (
owl:unionOfand the like) are skipped. Two classes with the same local name would share a vertex type, so inference refuses them.
The resource for Researcher makes a Researcher vertex, then reads the IRI in
authorOf into a Publication vertex and adds the authorOf edge; the same
for affiliatedWith. A researcher with several authorOf papers gets one edge
per paper. Classes, properties and edges are listed in IRI order, so inferring
again from the same ontology gives the same file.
3. Read the instances of each class from the data file¶
bindings = Bindings()
for resource in ingestion_model.resources:
connector = SparqlConnector(rdf_class=NAMESPACE + resource.name, rdf_file=DATA_FILE)
bindings.add_connector(connector)
bindings.bind_resource(resource.name, connector)
A connector says where the records of a resource come from. Each one here reads
the instances of one class from data.ttl and gives the resource one record per
instance: its IRI and one field per property.
4. Ingest¶
engine.define_and_ingest(
manifest=GraphManifest(
graph_schema=schema, ingestion_model=ingestion_model, bindings=bindings
),
target_db_config=conn_conf,
ingestion_params=IngestionParams(clear_data=True),
recreate_schema=True,
)
5. Run it¶
What you should see¶
The database holds:
| Count | Why | |
|---|---|---|
Researcher vertices |
4 | One per instance of ex:Researcher |
Publication vertices |
4 | One per instance; a cited or authored paper is matched by its IRI, not added again |
Institution vertices |
3 | Dave and Bob share ETH Zürich |
authorOf edges |
4 | One per researcher |
affiliatedWith edges |
4 | One per researcher |
cites edges |
3 | Papers 2, 3 and 4 each cite one paper |
Also possible¶
- To read the data from a SPARQL endpoint instead of a file, give each
SparqlConnectoranendpoint_urlin place ofrdf_file. - When the ontology and the instances are in one file,
engine.create_bindings_from_rdf(path)builds these connectors for you, all reading that file. - A blank node (
[ ex:value 3 ; ex:unit "nm" ]) has no IRI, so it is keyed on its content:_uriis_:followed by a digest of its triples, the same on every read. Reading the file again updates the same vertices. - IRIs joined by
owl:sameAsare read as one instance under the smallest IRI, with the others listed in a_same_asfield; declare_same_ason the vertex to store it.SparqlConnector(..., same_as="keep")reads the statements as an ordinarysameAsfield instead.
What to read next¶
- A graph from a PostgreSQL database: the same idea for relational tables.
- Credentials outside the manifest: how a manifest names a source without its password.
- Glossary: schema inference.
Files¶
The example lives in examples/10-infer-from-rdf.
generated-manifest.yaml
core_schema:
edge_config:
edges:
- relation: affiliatedWith
source: Researcher
target: Institution
- relation: authorOf
source: Researcher
target: Publication
- relation: cites
source: Publication
target: Publication
vertex_config:
vertices:
- identity:
- _uri
name: Institution
properties:
- name: _key
- name: _uri
- name: country
- name: instName
- identity:
- _uri
name: Publication
properties:
- name: _key
- name: _uri
- name: doi
- name: title
- name: year
- identity:
- _uri
name: Researcher
properties:
- name: _key
- name: _uri
- name: fullName
- name: orcid
metadata:
name: academic_kg
ingest.py
"""How do I turn an OWL ontology and RDF data into a property graph?
Infers a schema and one resource per class from ``data/ontology.ttl``, saves the
schema to ``generated-manifest.yaml``, connects each resource to the instances of
its class in ``data/data.ttl`` and writes the graph to ArangoDB. Run it from this
directory:
uv run python ingest.py
"""
from pathlib import Path
from suthing import FileHandle
from graflo import GraphManifest
from graflo.architecture.contract.bindings import Bindings, SparqlConnector
from graflo.connections import ArangoConfig
from graflo.hq import GraphEngine, IngestionParams
EXAMPLE_DIR = Path(__file__).resolve().parent
ONTOLOGY_FILE = EXAMPLE_DIR / "data" / "ontology.ttl"
DATA_FILE = EXAMPLE_DIR / "data" / "data.ttl"
NAMESPACE = "http://example.org/"
# 1. Infer the schema and the resources from the ontology.
conn_conf = ArangoConfig.from_docker_env()
engine = GraphEngine(target_db_flavor=conn_conf.connection_type)
schema, ingestion_model = engine.infer_schema_from_rdf(
source=ONTOLOGY_FILE, schema_name="academic_kg"
)
# 2. Save what was inferred, to read or edit.
FileHandle.dump(
schema.model_dump(exclude_defaults=True), EXAMPLE_DIR / "generated-manifest.yaml"
)
# 3. Read the instances of each class from the data file.
bindings = Bindings()
for resource in ingestion_model.resources:
connector = SparqlConnector(rdf_class=NAMESPACE + resource.name, rdf_file=DATA_FILE)
bindings.add_connector(connector)
bindings.bind_resource(resource.name, connector)
# 4. Write the graph.
engine.define_and_ingest(
manifest=GraphManifest(
graph_schema=schema, ingestion_model=ingestion_model, bindings=bindings
),
target_db_config=conn_conf,
ingestion_params=IngestionParams(clear_data=True),
recreate_schema=True,
)