Skip to content

How do I try GraFlo without a database, and export a graph to files?

You want to see the graph GraFlo builds from your data before you set up a database. Or you have a graph in a database and want a copy on disk, to keep it or to load it into another database later.

GraFlo can write a graph to a directory instead of a database. The directory, called a file backend, holds the schema and the records in compressed files. GraFlo reads it back as if it were a database, so you can load it into any database later. This example uses the data and the manifest of the CSV example (01); only the target changes.

flowchart LR
    csv[CSV files] -- ingest.py --> files[(artifacts/csv-backend)]
    neo4j[(Neo4j)] -- export.py --> dump[(artifacts/neo4j-backend)]
    files -- migrate.py --> arango[(ArangoDB)]
    dump -- migrate.py --> arango

What you need

  • GraFlo installed (pip install graflo).
  • No database for step 1. Step 2 reads from Neo4j and step 3 writes to ArangoDB; the repository ships containers for both, see docker/README.md.

The data

The two files of example 01, unchanged: data/people.csv lists three people with their age, and data/departments.csv says which department each of them works in. manifest.yaml is example 01's manifest: person is identified by id, department by name, and each person has an edge to their department.

Steps

1. Write the graph to a directory

ingest.py is example 01's script with another target: a GraFloBackendConfig that names a directory, instead of an ArangoConfig.

backend = GraFloBackendConfig(output_dir=Path("artifacts/csv-backend"))

engine = GraphEngine(target_db_flavor=backend.connection_type)
engine.define_and_ingest(
    manifest=manifest,
    target_db_config=backend,
    ingestion_params=IngestionParams(clear_data=True),
    recreate_schema=True,
)
cd examples/14-file-backend-export
uv run python ingest.py

2. Copy a graph from a database to files

export.py copies whatever graph the Neo4j container holds. To have something to copy, load example 01 into Neo4j first; its README shows the two lines to change.

source = Neo4jConfig.from_docker_env()
backend = GraFloBackendConfig(output_dir=Path("artifacts/neo4j-backend"))

engine = GraphEngine(target_db_flavor=backend.connection_type)
engine.migrate_graph(source, backend)

migrate_graph reads the schema and every record from the source and writes them to the target. Here the target is a directory, so the copy lands in artifacts/neo4j-backend.

uv run python export.py
uv run python inspect_backend.py artifacts/neo4j-backend

3. Load the files into a database

The directory is also a source. migrate.py runs the same call the other way round, from a directory to ArangoDB, and replaces the graph that is there:

source = GraFloBackendConfig(output_dir=backend_dir)
target = ArangoConfig.from_docker_env()

engine = GraphEngine(target_db_flavor=target.connection_type)
engine.migrate_graph(source, target)
uv run python migrate.py                          # artifacts/csv-backend
uv run python migrate.py artifacts/neo4j-backend

What you should see

After step 1 the directory holds the schema, an index, and one or more gzip-compressed JSON Lines files per type, one record per line:

artifacts/csv-backend/
├── INDEX.json          record count and file names per type
├── schema.yaml         the schema, no data
├── vertices/
│   ├── department.000.jsonl.gz
│   ├── person.000.jsonl.gz
│   └── person.001.jsonl.gz
└── edges/
    └── person____department.000.jsonl.gz

inspect_backend.py counts the records:

uv run python inspect_backend.py
artifacts/csv-backend (schema hr)
vertices:
  person      6 records, 3 distinct identities (id)
  department  3 records, 3 distinct identities (name)
edges:
  person -> department    3 records

The file backend appends every record it receives; it does not merge records that have the same identity. Both files mention each person, so person holds six records. The record {"age":"27","id":"1","name":"John Hancock"} from people.csv and the record {"id":"1","name":"John Hancock"} from departments.csv have the same id, 1. A database stores them as one vertex; the file backend keeps both records, so each id appears twice. When migrate.py loads the directory into ArangoDB, the database merges them: it holds three person vertices, three department vertices and three edges, as after example 01.

Also possible

  • Name the database you will load the files into when you write them, with target_flavor_hint on GraFloBackendConfig; see Writing for a known target.
  • GraphEngine.export_graph returns the schema and the records as Python objects and writes nothing.

Files

The example lives in examples/14-file-backend-export.

manifest.yaml
schema:
    metadata:
        name: hr
    graph:
        vertex_config:
            vertices:
            -   name: person
                properties:
                -   id
                -   name
                -   age
                identity:
                -   id
            -   name: department
                properties:
                -   name
                identity:
                -   name
        edge_config:
            edges:
            -   source: person
                target: department
    db_profile: {}
ingestion_model:
    resources:
    -   name: people
        pipeline:
        -   vertex: person
    -   name: departments
        pipeline:
        -   vertex: person
            from:
                id: person_id
                name: person
        -   vertex: department
            from:
                name: department
bindings:
    connectors:
    -   regex: "^people.*\\.csv$"
        sub_path: data
        resource_name: people
    -   regex: "^dep.*\\.csv$"
        sub_path: data
        resource_name: departments
export.py
"""Copy the graph in a Neo4j database to files.

Reads the schema and every vertex and edge from the Neo4j container started
from ``docker/neo4j`` and writes them to the directory
``artifacts/neo4j-backend``. Run it from this directory:

    uv run python export.py
"""

from pathlib import Path

from graflo.connections import GraFloBackendConfig, Neo4jConfig
from graflo.hq import GraphEngine

source = Neo4jConfig.from_docker_env()
backend = GraFloBackendConfig(output_dir=Path("artifacts/neo4j-backend"))

engine = GraphEngine(target_db_flavor=backend.connection_type)
engine.migrate_graph(source, backend)
print(f"Wrote {backend.output_dir}")
ingest.py
"""How do I try GraFlo without a database, and export a graph to files?

Reads ``manifest.yaml`` (the manifest of example 01) and writes the graph built
from the two CSV files in ``data/`` to the directory ``artifacts/csv-backend``
instead of a database. Run it from this directory:

    uv run python ingest.py
"""

from pathlib import Path

from suthing import FileHandle

from graflo import GraphManifest
from graflo.connections import GraFloBackendConfig
from graflo.hq import GraphEngine
from graflo.hq.caster import IngestionParams

manifest = GraphManifest.from_config(FileHandle.load("manifest.yaml"))
manifest.finish_init()

# The target is a directory: the only change from example 01's ingest.py.
backend = GraFloBackendConfig(output_dir=Path("artifacts/csv-backend"))

engine = GraphEngine(target_db_flavor=backend.connection_type)
engine.define_and_ingest(
    manifest=manifest,
    target_db_config=backend,
    ingestion_params=IngestionParams(clear_data=True),
    recreate_schema=True,
)
print(f"Wrote {backend.output_dir}")
inspect_backend.py
"""Show what a file backend directory holds: records and distinct vertices per type.

The file backend appends every record it receives; it does not merge records
that have the same identity, as a database does. This script therefore counts
both the records and the distinct identities of each vertex type. Run it from
this directory, after ``ingest.py`` or ``export.py``:

    uv run python inspect_backend.py
    uv run python inspect_backend.py artifacts/neo4j-backend
"""

from pathlib import Path

import click

from graflo.architecture.backend import GraFloBackendReader


@click.command()
@click.argument(
    "backend_dir",
    type=click.Path(exists=True, file_okay=False, path_type=Path),
    default="artifacts/csv-backend",
)
def main(backend_dir: Path) -> None:
    """Print record and identity counts for every vertex and edge type."""
    reader = GraFloBackendReader(backend_dir)
    schema = reader.read_schema()
    vertex_config = schema.core_schema.vertex_config

    click.echo(f"{backend_dir} (schema {schema.metadata.name})")
    click.echo("vertices:")
    for vertex in vertex_config.vertices:
        fields = vertex_config.identity_fields(vertex.name)
        records = [
            doc for batch in reader.iter_vertex_batches(vertex.name) for doc in batch
        ]
        distinct = {tuple(doc.get(field) for field in fields) for doc in records}
        click.echo(
            f"  {vertex.name:<12}{len(records)} records, "
            f"{len(distinct)} distinct identities ({', '.join(fields)})"
        )
    click.echo("edges:")
    for edge in schema.core_schema.edge_config.edges:
        records = [
            row for batch in reader.iter_edge_batches(edge.edge_id) for row in batch
        ]
        label = f"{edge.source} -> {edge.target}"
        click.echo(f"  {label:<24}{len(records)} records")


if __name__ == "__main__":
    main()
migrate.py
"""Load a graph from files into ArangoDB.

Reads a file backend directory (``artifacts/csv-backend`` unless you name
another) and writes its schema and records to the ArangoDB container started
from ``docker/arango``. Run it from this directory:

    uv run python migrate.py
    uv run python migrate.py artifacts/neo4j-backend
"""

from pathlib import Path

import click

from graflo.connections import ArangoConfig, GraFloBackendConfig
from graflo.hq import GraphEngine


@click.command()
@click.argument(
    "backend_dir",
    type=click.Path(exists=True, file_okay=False, path_type=Path),
    default="artifacts/csv-backend",
)
def main(backend_dir: Path) -> None:
    """Replace the graph in ArangoDB with the one stored in BACKEND_DIR."""
    source = GraFloBackendConfig(output_dir=backend_dir)
    target = ArangoConfig.from_docker_env()

    engine = GraphEngine(target_db_flavor=target.connection_type)
    engine.migrate_graph(source, target)
    click.echo(f"Loaded {backend_dir} into ArangoDB")


if __name__ == "__main__":
    main()