Skip to content

Database connections

GraFlo needs to know which database to write the graph to: its address, the credentials, and the name of the graph inside it. None of this is in the manifest; you pass it in Python as a config object. This guide shows how to build that object from environment variables, in code or from a YAML file, so that you can point any script or example at your own database.

What you need

  • GraFlo installed (pip install graflo).
  • The address and credentials of a running database, or a directory for the file backend.

Steps

1. Pick the class for your database

Every class is imported from graflo.connections.

Database Class Variable prefix Typical URI Field that names the graph
ArangoDB ArangoConfig ARANGO_ http://localhost:8529 database
Neo4j Neo4jConfig NEO4J_ bolt://localhost:7687 database
TigerGraph TigergraphConfig TIGERGRAPH_ http://localhost:14240 schema_name
FalkorDB FalkordbConfig FALKORDB_ redis://localhost:6379 database
Memgraph MemgraphConfig MEMGRAPH_ bolt://localhost:7687 database
NebulaGraph NebulaConfig NEBULA_ nebula://localhost:9669 schema_name (the space)
PostgreSQL PostgresConfig POSTGRES_ postgresql://localhost:5432 database and schema_name
File backend GraFloBackendConfig GRAFLO_BACKEND_ none output_dir

If you leave the field that names the graph empty, GraFlo uses the manifest's schema.metadata.name, adjusted to the characters the database accepts; see Graph namespace and schema. PostgreSQL is different: database must name a database that exists, and GraFlo creates the tables in the schema schema_name, public when it is not set.

A URI without a port gets the database's default. For Neo4j the default follows the scheme: bolt_port when it is set and otherwise 7687 for bolt and neo4j URIs, 7474 for http and https.

2. Fill in the settings

Four ways lead to the same object. Environment variables suit most programs; the others follow.

From environment variables

export ARANGO_URI=http://localhost:8529
export ARANGO_USERNAME=root
export ARANGO_PASSWORD=change-me
export ARANGO_DATABASE=plant
from graflo.connections import ArangoConfig

conn_conf = ArangoConfig.from_env()

from_env() reads each field from the variable named by the prefix and the field name in upper case: ARANGO_URI, ARANGO_USERNAME, ARANGO_PASSWORD, ARANGO_DATABASE. Fields that only one database has are read the same way, for example NEO4J_BOLT_PORT, TIGERGRAPH_SECRET (token authentication) and NEBULA_VERSION (3 or 5). The graph, space or schema name is read from TIGERGRAPH_SCHEMA_NAME, NEBULA_SCHEMA_NAME and POSTGRES_SCHEMA_NAME. In code and in a YAML file the field also accepts the key schema; the environment does not.

When one program talks to two databases of the same kind, give each set of variables its own qualifier:

export SENSORS_ARANGO_URI=http://sensors-db.example:8529
export SENSORS_ARANGO_DATABASE=sensor_feed
sensors_conf = ArangoConfig.from_env(prefix="SENSORS")

prefix="SENSORS" reads SENSORS_ARANGO_URI. The two other qualifiers place the word elsewhere: profile="DEV" reads ARANGO_DEV_URI, and suffix="DEV" reads ARANGO_URI_DEV. Use one qualifier per call.

In code

from graflo.connections import TigergraphConfig

conn_conf = TigergraphConfig(
    uri="http://localhost:14240",
    username="tigergraph",
    schema_name="plant",
)

A keyword argument wins over the environment, and a field you do not pass is still read from it. Here the password comes from TIGERGRAPH_PASSWORD, while the name of the graph is fixed in code.

For PostgreSQL, PostgresConfig.from_dsn("postgresql://user:password@host:5432/plant") fills the address, credentials and database from one connection string.

In a YAML file, for the command line

graflo ingest --db-config-path db.yaml and graflo migrate-schema apply --db-config-path db.yaml read the database from a file. db_type picks the class: arango, neo4j, tigergraph, falkordb, memgraph, nebula, postgres or graflo_backend. The other keys are the fields of that class.

db_type: neo4j
uri: bolt://localhost:7687
username: neo4j
database: plant

The file is read with DBConfig.from_dict, and fields it leaves out are read from the environment. Keep the password out of the file and set NEO4J_PASSWORD instead.

From the containers of the repository

A clone of the repository has a Docker Compose setup per database under docker/; docker/README.md explains how to start them. The examples connect to them with:

conn_conf = ArangoConfig.from_docker_env()

from_docker_env() reads the settings file of the container in docker/arango/ (docker/neo4j/ for Neo4jConfig, and so on). It works only in a clone, because those files belong to the repository, not to the installed package; pass docker_dir= to read another directory. GraFloBackendConfig has no container: give it output_dir instead.

3. Pass the config to the engine

Hand the config to GraphEngine together with a loaded manifest, as in the quick start:

from graflo.hq import GraphEngine

engine = GraphEngine(target_db_flavor=conn_conf.connection_type)
engine.define_and_ingest(manifest=manifest, target_db_config=conn_conf)

connection_type is the kind of database the config describes. Passing it to GraphEngine makes the engine prepare the schema for that database.

What you should see

Print the config before you connect to check what was read. For the environment variables of step 2:

print(conn_conf.connection_type, conn_conf.uri, conn_conf.database)
arango http://localhost:8529 plant

A field that shows None was found neither in the call nor in the environment.

Databases you read from

A PostgreSQL database, a SPARQL endpoint, a REST API or a Kafka topic that records come from is not configured this way. The manifest names it by a label (conn_proxy), and at run time you register the settings for that label with a connection provider, so the manifest never holds a password. The connection proxy example (11) shows it for PostgreSQL, and the API environment example (12) for REST APIs.