Skip to content

GraFlo ontology

A manifest is a YAML file, but catalogs, lineage tools and SPARQL endpoints work with RDF. GraFlo ships an OWL ontology, with the prefix gf:, that describes the manifest itself: its schema, resources, transforms and bindings. This page shows what the vocabulary covers and how to convert a manifest to RDF and back, from Python or the shell, without losing anything the manifest says.

The ontology describes manifests, not your domain. Building a manifest from your own OWL or RDFS ontology (ex:Machine, ex:WorkOrder, ...) is a different task, done by RdfInferenceManager; see the RDF inference example (10).

flowchart TB
    subgraph domain ["Your domain"]
        UOWL["OWL/RDFS ontology<br/>ex:Machine, ex:WorkOrder"]
        UOWL --> RIM["RdfInferenceManager"]
        RIM --> LPG["GraFlo Schema + IngestionModel"]
    end

    subgraph meta ["The manifest as RDF"]
        YAML["GraphManifest YAML"]
        YAML --> SER["ManifestRdfSerializer"]
        SER --> GFOWL["gf: GraphManifest RDF"]
        GFOWL --> DES["ManifestRdfDeserializer"]
        DES --> YAML
    end

    LPG -. "same Pydantic types" .- YAML

Ontology identifiers

Role IRI
Ontology document https://ontology.growgraph.dev/graflo
Version IRI https://ontology.growgraph.dev/graflo/1.7.0
Version info 1.7.0
Vocabulary prefix gf: https://ontology.growgraph.dev/graflo/

The Turtle source lives in the package at graflo/rdf/ontology/graflo.ttl. Constants are also exposed in Python as graflo.rdf.namespace (GF_ONTOLOGY_IRI, GF_VERSION, GF_VERSION_IRI, GF_BASE).

Interactive visualization

The explorer below is a class graph from graflo.ttl. Classes are grouped into the blocks a manifest is made of; the bands are derived from what gf:GraphManifest points at, not hand-assigned. Within a band, columns run left to right from general to specific: a class sits one column right of whatever contains it (gf:Schema → gf:CoreSchema → gf:VertexConfig → gf:Vertex → gf:Field) or generalises it (gf:Actor → gf:VertexProducingActor → gf:VertexActor), and specialization always wins, so a subclass is never level with its superclass. Classes at the same distance stay in the same column — gf:Vertex and gf:Edge are peers. The layout is deterministic: the same ontology always draws the same picture.

Only the taxonomy is drawn by default; select a class to reveal its properties, or switch the filter to Display all. gf:GrafloArtifact is the superclass of nearly every class, so its links are hidden by default — tick Show GrafloArtifact to bring them back. Drag, scroll to zoom, click to focus. Regenerate with uv run python docs/_build/scripts/build_ontology_viz.py after ontology edits.

If the embedded viewer is blank in an IDE browser preview, use Open full screen in a normal browser tab.

Open full screen

What the vocabulary covers

Schema block

  • gf:GraphManifest, gf:Schema, gf:CoreSchema, gf:GraphMetadata, gf:DatabaseProfile
  • gf:VertexConfig, gf:EdgeConfig, gf:Vertex, gf:Edge, gf:Field, gf:Identity
  • gf:FieldType individuals (gf:INT, gf:STRING, …)

Ingestion block

  • gf:IngestionModel, gf:Resource, gf:ProtoTransform, gf:Transform
  • gf:DressConfig, gf:KeySelectionConfig, gf:EdgeInferSpec
  • Pipeline actor steps (blank nodes): gf:VertexActor, gf:EdgeActor, gf:TransformActor, gf:DescendActor, gf:VertexRouterActor (Python aliases *ActorStep in graflo.rdf.namespace)

Bindings block

  • gf:Bindings, and gf:BoundConnector with one subclass per connector model: gf:FileConnector, gf:TableConnector, gf:SparqlConnector, gf:APIConnector, gf:KafkaConnector
  • gf:ResourceConnectorBinding, gf:ConnectorConnectionBinding, gf:StagingProxyBinding

Semantic grounding (optional)

An element may be anchored to an external vocabulary through a semantics: block on the schema metadata, a vertex, an edge, or a field:

vertices:
-   name: person
    identity: [email]
    semantics:
        iri: https://schema.org/Person
        exact_match: [http://xmlns.com/foaf/0.1/Person]
        synonyms: [individual, human]
    properties:
    -   name: speed
        type: FLOAT
        semantics:
            unit: m/s

This maps to gf:semanticIri, skos:exactMatch, skos:altLabel, and — fields only — gf:unit. The block is purely descriptive: identity, storage naming and ingestion behave identically whether or not it is present. unit is rejected outside a field, where it would be meaningless.

Grounding survives a fold. When two definitions of one type or property are combined — by merge_vertices, or by a merge_manifests equivalence — exact_match and synonyms union, since they are sets of claims. A single-valued iri cannot: two sides denoting different concepts denote neither exactly, so a disagreement clears it rather than electing one. unit is the exception that refuses outright — unlike an iri, two units mean the combined property would hold numerically incomparable values, which is a defect in the data rather than in its description.

Declared, symmetric and native inverses

gf:EdgeConfig points at its declared inverse pairs through gf:hasInverse; each gf:EdgeInverse carries gf:relation and gf:inverseRelation (the pair is unordered), ordered by gf:artifactIndex. Relations declared as their own inverse are gf:symmetricRelation literals on the gf:EdgeConfig. A gf:DatabaseProfile names each relation whose inverse the database maintains with gf:nativeInverseRelation. The inverse type's name is not repeated on the profile — it is the declared pair's other relation. The third realization needs no term: a materialized inverse is an ordinary gf:Edge, and the emit_inverse flag that feeds it rides in the step's gf:stepPayload like every other step option.

These are reified, string-valued terms describing a manifest, not OWL axioms over your relations. Going the other way, schema inference from an ontology does read owl:inverseOf and owl:SymmetricProperty, and turns them into edge_config.inverses and edge_config.symmetric.

List field types

gf:FieldType has one individual for each of the nine FieldType members, gf:UUID and gf:LIST included, and a list's element type is given by gf:itemType (domain gf:Field, range gf:FieldType).

properties:
    - name: tags
      type: LIST
      item_type: STRING

Naming convention (optional)

Schema metadata may declare how the schema spells the identifiers it invents:

metadata:
    name: shop
    naming:
        vertex_case: pascal          # Customer, OrderLine
        relation_case: camel         # placedBy, reportsTo
        property_case: preserve      # exactly as the source names them
        singular_vertex_names: true

This maps to gf:hasNamingConvention and a gf:NamingConvention node carrying gf:vertexCase, gf:relationCase, gf:propertyCase (each a gf:NameCase individual) and gf:singularVertexNames.

Two things the block is careful about. property_case defaults to preserve because a property name binds to a key in the source document — restyling one without the matching transform.rename leaves a property that never receives a value, and nothing reports it. And like semantics:, the block is descriptive: it records the convention the names follow so a later author extending the schema need not infer it, but nothing consults it at runtime and declaring it does not rewrite anything.

Enumerations (named individuals): gf:DBType (ArangoDB, Neo4j, …), transform target/strategy, key-selection mode, edge duplicate policy, bound source kind.

PROV-O hooks: gf:GraphManifest ⊑ prov:Entity, gf:ProtoTransform ⊑ prov:Activity (subclasses such as gf:Transform inherit this; for lineage tooling).

Manifest instance URIs

When you serialize a manifest, you pass a base_uri that identifies that manifest document (not the ontology). The serializer mints stable paths under it, for example:

Path under base_uri RDF type
(base_uri) gf:GraphManifest
schema gf:Schema
schema/core/vertex-config gf:VertexConfig
schema/core/edge-config gf:EdgeConfig
schema/core/vertex/Person gf:Vertex
schema/core/edge/Person_knows_Person gf:Edge
ingestion gf:IngestionModel
ingestion/resource/my_resource gf:Resource
ingestion/transform/my_transform gf:ProtoTransform
bindings gf:Bindings
bindings/connector/<hash> the gf:BoundConnector subclass of the connector, such as gf:FileConnector

Pipeline steps are blank nodes typed with the appropriate gf:*Actor class (and gf:Actor); the full step dict is stored in gf:stepPayload as JSON so round-trip preserves shorthand YAML shapes (vertex: person, nested descend, transform.call, …).

List order (resources, transforms, connectors, vertices, fields, pipeline steps) is preserved via gf:artifactIndex.

Python API

from graflo import GraphManifest
from graflo.rdf import ManifestRdfDeserializer, ManifestRdfSerializer

manifest = GraphManifest.from_yaml("manifest.yaml")
base = "https://growgraph.dev/manifests/academic/v1"

serializer = ManifestRdfSerializer(include_ontology=True)
ttl = serializer.to_turtle(manifest, base)

restored = ManifestRdfDeserializer().from_turtle(
    ttl,
    manifest_uri=base.rstrip("/"),
)
  • include_ontology=True (default) embeds graflo.ttl triples in the output graph — useful for self-contained Turtle files.
  • to_json_ld, to_graph — same graph, other serializations.

CLI

# Manifest to RDF
graflo manifest-to-rdf manifest.yaml \
  --base-uri https://growgraph.dev/manifests/academic/v1 \
  --format turtle \
  --output academic.ttl

# RDF to manifest YAML
graflo rdf-to-manifest academic.ttl \
  --manifest-uri https://growgraph.dev/manifests/academic/v1 \
  --output manifest.restored.yaml

--format is turtle (default), json-ld, nt or xml; without --output the result is printed. --no-include-ontology leaves the ontology's own triples out. rdf-to-manifest reads the same formats and n3, chosen with --input-format.

Round-trip fidelity

Area Behavior
Scalars, enums, transforms Full via literals and gf individuals
pipeline actor steps Full via gf:stepPayload JSON
params, connector extras JSON literals on payload properties
YAML aliases (schema / graph, pipeline / apply) Canonical names only in restored YAML
Runtime PrivateAttr state Not serialized; call finish_init() after load
Vertex filters Serialized in gf:vertexPayload JSON; not decomposed into filter AST

The guaranteed invariant matches the rest of GraFlo config: semantic canonical round-trip (parse → RDF → parse equals minimal canonical dict), not byte-identical YAML.

JSON-LD

graflo/rdf/ontology/graflo-context.jsonld maps common JSON keys to gf: IRIs for tools that consume JSON-LD directly. The serializer’s to_json_ld() output can be combined with this context in downstream pipelines.