Getting started
Getting started takes you from nothing to a running Datastrato Enterprise with a catalog attached and metadata flowing from it, in three stages: install, connect a source, then decide what to set up next.
Quick Start
1. Get a license. Trial licenses come from the Datastrato trial page and production licenses from your account team. Either way you receive a release package holding the license file and a pull key for the release registry. See Licensing.
2. Install. Follow Install with Helm through signing in to the UI. It covers the
cluster requirements, the values file, and creating your first metalake, and it takes about thirty
minutes. The commands below reuse the shell variables set there (ADMIN_PASSWORD and METALAKE)
and the port forward to the server on localhost:8090.
3. Connect a catalog, as described in the next section.
Connect a Catalog
A catalog attaches one metadata source. PostgreSQL is the quickest to verify, and the shape is the same for every source. Set the connection details, then create the catalog.
PG_HOST={postgres_host}
PG_DB={postgres_database}
PG_USER={postgres_user}
PG_PASSWORD='{postgres_password}'
curl -s -u "admin:$ADMIN_PASSWORD" -X POST -H "Content-Type: application/json" -d "{
\"name\": \"pg_demo\",
\"type\": \"relational\",
\"provider\": \"jdbc-postgresql\",
\"comment\": \"\",
\"properties\": {
\"jdbc-url\": \"jdbc:postgresql://$PG_HOST:5432/$PG_DB\",
\"jdbc-database\": \"$PG_DB\",
\"jdbc-user\": \"$PG_USER\",
\"jdbc-password\": \"$PG_PASSWORD\",
\"jdbc-driver\": \"org.postgresql.Driver\"
}
}" \
"http://localhost:8090/api/metalakes/$METALAKE/catalogs" | jq '{code, name: .catalog.name, message}'
The response shows "code": 0. jdbc-database is required for PostgreSQL and SQL Server even though
the URL already names the database; without it the catalog is created but every operation on it
fails.
List the schemas to confirm metadata is flowing.
curl -s -u "admin:$ADMIN_PASSWORD" \
"http://localhost:8090/api/metalakes/$METALAKE/catalogs/pg_demo/schemas" | jq .
What comes back is read from the source, not from Gravitino's own store. The catalog now appears in the UI as well, and catalogs can be created there instead of through the API.
Next Steps
Connect the sources you actually care about. Every catalog type is documented under Connect sources, and Connecting a catalog covers what they have in common.
Set up identity. The install signs in with a local admin account. A production deployment
authenticates users against your identity provider. See
Authentication, then SCIM provisioning
to have users and groups arrive from your directory.
Point an engine at it. Connectors on Kubernetes covers getting the Trino, Spark, or Flink connector into an engine running in your cluster. Engines that speak Iceberg REST need the Iceberg REST service published through the ingress, which Install with Helm lists under Configure.
Move off the built-in database. The install keeps the entity store in a PostgreSQL the chart deploys, which suits a trial. For production, use a database you operate and back up. See Relational backend storage.