How do I try GraFlo without a database, and export a graph to files?¶
You want to see the graph GraFlo builds from your data before you set up a database. Or you have a graph in a database and want a copy on disk, to keep it or to load it into another database later.
GraFlo can write a graph to a directory instead of a database. The directory, called a file backend, holds the schema and the records in compressed files. GraFlo reads it back as if it were a database, so you can load it into any database later. This example uses the data and the manifest of the CSV example (01); only the target changes.
flowchart LR
csv[CSV files] -- ingest.py --> files[(artifacts/csv-backend)]
neo4j[(Neo4j)] -- export.py --> dump[(artifacts/neo4j-backend)]
files -- migrate.py --> arango[(ArangoDB)]
dump -- migrate.py --> arango
What you need¶
- GraFlo installed (
pip install graflo). - No database for step 1. Step 2 reads from Neo4j and step 3 writes to
ArangoDB; the repository ships containers for both, see
docker/README.md.
The data¶
The two files of example 01, unchanged: data/people.csv
lists three people with their age, and
data/departments.csv says which department each of
them works in. manifest.yaml is example 01's manifest:
person is identified by id, department by name, and each person has an
edge to their department.
Steps¶
1. Write the graph to a directory¶
ingest.py is example 01's script with another target: a
GraFloBackendConfig that names a directory, instead of an ArangoConfig.
backend = GraFloBackendConfig(output_dir=Path("artifacts/csv-backend"))
engine = GraphEngine(target_db_flavor=backend.connection_type)
engine.define_and_ingest(
manifest=manifest,
target_db_config=backend,
ingestion_params=IngestionParams(clear_data=True),
recreate_schema=True,
)
2. Copy a graph from a database to files¶
export.py copies whatever graph the Neo4j container holds. To
have something to copy, load example 01 into Neo4j first; its README shows the
two lines to change.
source = Neo4jConfig.from_docker_env()
backend = GraFloBackendConfig(output_dir=Path("artifacts/neo4j-backend"))
engine = GraphEngine(target_db_flavor=backend.connection_type)
engine.migrate_graph(source, backend)
migrate_graph reads the schema and every record from the source and writes
them to the target. Here the target is a directory, so the copy lands in
artifacts/neo4j-backend.
3. Load the files into a database¶
The directory is also a source. migrate.py runs the same call
the other way round, from a directory to ArangoDB, and replaces the graph that
is there:
source = GraFloBackendConfig(output_dir=backend_dir)
target = ArangoConfig.from_docker_env()
engine = GraphEngine(target_db_flavor=target.connection_type)
engine.migrate_graph(source, target)
What you should see¶
After step 1 the directory holds the schema, an index, and one or more gzip-compressed JSON Lines files per type, one record per line:
artifacts/csv-backend/
├── INDEX.json record count and file names per type
├── schema.yaml the schema, no data
├── vertices/
│ ├── department.000.jsonl.gz
│ ├── person.000.jsonl.gz
│ └── person.001.jsonl.gz
└── edges/
└── person____department.000.jsonl.gz
inspect_backend.py counts the records:
artifacts/csv-backend (schema hr)
vertices:
person 6 records, 3 distinct identities (id)
department 3 records, 3 distinct identities (name)
edges:
person -> department 3 records
The file backend appends every record it receives; it does not merge records
that have the same identity. Both files mention each person, so person holds
six records. The record {"age":"27","id":"1","name":"John Hancock"} from
people.csv and the record {"id":"1","name":"John Hancock"} from
departments.csv have the same id, 1. A database stores them as one vertex;
the file backend keeps both records, so each id appears twice. When
migrate.py loads the directory into ArangoDB, the database merges them: it
holds three person vertices, three department vertices and three edges, as
after example 01.
Also possible¶
- Name the database you will load the files into when you write them, with
target_flavor_hintonGraFloBackendConfig; see Writing for a known target. GraphEngine.export_graphreturns the schema and the records as Python objects and writes nothing.
What to read next¶
- My data has no obvious key: find what identifies a record, and write the result to a file backend.
- Graph export and replay: the same steps for your own graph.
- Graph export and migration: the directory layout, which databases can be read as a graph, and the limits.
Files¶
The example lives in examples/14-file-backend-export.
manifest.yaml
schema:
metadata:
name: hr
graph:
vertex_config:
vertices:
- name: person
properties:
- id
- name
- age
identity:
- id
- name: department
properties:
- name
identity:
- name
edge_config:
edges:
- source: person
target: department
db_profile: {}
ingestion_model:
resources:
- name: people
pipeline:
- vertex: person
- name: departments
pipeline:
- vertex: person
from:
id: person_id
name: person
- vertex: department
from:
name: department
bindings:
connectors:
- regex: "^people.*\\.csv$"
sub_path: data
resource_name: people
- regex: "^dep.*\\.csv$"
sub_path: data
resource_name: departments
export.py
"""Copy the graph in a Neo4j database to files.
Reads the schema and every vertex and edge from the Neo4j container started
from ``docker/neo4j`` and writes them to the directory
``artifacts/neo4j-backend``. Run it from this directory:
uv run python export.py
"""
from pathlib import Path
from graflo.connections import GraFloBackendConfig, Neo4jConfig
from graflo.hq import GraphEngine
source = Neo4jConfig.from_docker_env()
backend = GraFloBackendConfig(output_dir=Path("artifacts/neo4j-backend"))
engine = GraphEngine(target_db_flavor=backend.connection_type)
engine.migrate_graph(source, backend)
print(f"Wrote {backend.output_dir}")
ingest.py
"""How do I try GraFlo without a database, and export a graph to files?
Reads ``manifest.yaml`` (the manifest of example 01) and writes the graph built
from the two CSV files in ``data/`` to the directory ``artifacts/csv-backend``
instead of a database. Run it from this directory:
uv run python ingest.py
"""
from pathlib import Path
from suthing import FileHandle
from graflo import GraphManifest
from graflo.connections import GraFloBackendConfig
from graflo.hq import GraphEngine
from graflo.hq.caster import IngestionParams
manifest = GraphManifest.from_config(FileHandle.load("manifest.yaml"))
manifest.finish_init()
# The target is a directory: the only change from example 01's ingest.py.
backend = GraFloBackendConfig(output_dir=Path("artifacts/csv-backend"))
engine = GraphEngine(target_db_flavor=backend.connection_type)
engine.define_and_ingest(
manifest=manifest,
target_db_config=backend,
ingestion_params=IngestionParams(clear_data=True),
recreate_schema=True,
)
print(f"Wrote {backend.output_dir}")
inspect_backend.py
"""Show what a file backend directory holds: records and distinct vertices per type.
The file backend appends every record it receives; it does not merge records
that have the same identity, as a database does. This script therefore counts
both the records and the distinct identities of each vertex type. Run it from
this directory, after ``ingest.py`` or ``export.py``:
uv run python inspect_backend.py
uv run python inspect_backend.py artifacts/neo4j-backend
"""
from pathlib import Path
import click
from graflo.architecture.backend import GraFloBackendReader
@click.command()
@click.argument(
"backend_dir",
type=click.Path(exists=True, file_okay=False, path_type=Path),
default="artifacts/csv-backend",
)
def main(backend_dir: Path) -> None:
"""Print record and identity counts for every vertex and edge type."""
reader = GraFloBackendReader(backend_dir)
schema = reader.read_schema()
vertex_config = schema.core_schema.vertex_config
click.echo(f"{backend_dir} (schema {schema.metadata.name})")
click.echo("vertices:")
for vertex in vertex_config.vertices:
fields = vertex_config.identity_fields(vertex.name)
records = [
doc for batch in reader.iter_vertex_batches(vertex.name) for doc in batch
]
distinct = {tuple(doc.get(field) for field in fields) for doc in records}
click.echo(
f" {vertex.name:<12}{len(records)} records, "
f"{len(distinct)} distinct identities ({', '.join(fields)})"
)
click.echo("edges:")
for edge in schema.core_schema.edge_config.edges:
records = [
row for batch in reader.iter_edge_batches(edge.edge_id) for row in batch
]
label = f"{edge.source} -> {edge.target}"
click.echo(f" {label:<24}{len(records)} records")
if __name__ == "__main__":
main()
migrate.py
"""Load a graph from files into ArangoDB.
Reads a file backend directory (``artifacts/csv-backend`` unless you name
another) and writes its schema and records to the ArangoDB container started
from ``docker/arango``. Run it from this directory:
uv run python migrate.py
uv run python migrate.py artifacts/neo4j-backend
"""
from pathlib import Path
import click
from graflo.connections import ArangoConfig, GraFloBackendConfig
from graflo.hq import GraphEngine
@click.command()
@click.argument(
"backend_dir",
type=click.Path(exists=True, file_okay=False, path_type=Path),
default="artifacts/csv-backend",
)
def main(backend_dir: Path) -> None:
"""Replace the graph in ArangoDB with the one stored in BACKEND_DIR."""
source = GraFloBackendConfig(output_dir=backend_dir)
target = ArangoConfig.from_docker_env()
engine = GraphEngine(target_db_flavor=target.connection_type)
engine.migrate_graph(source, target)
click.echo(f"Loaded {backend_dir} into ArangoDB")
if __name__ == "__main__":
main()