Skip to main content

Credential Vending

Background

Gravitino credential vending is used to generate temporary or static credentials for accessing data. With credential vending, Gravitino provides a unified way to control access to diverse data sources across different platforms.

Supported Catalogs

Catalog typeVends
FilesetS3, OSS, GCS, ADLS
HiveS3, OSS, GCS, ADLS
IcebergS3, OSS, GCS, ADLS
GlueS3
JDBCJDBC user and password
PaimonS3, OSS, JDBC user and password

S3 is Amazon S3, OSS is Alibaba Cloud OSS, GCS is Google Cloud Storage, and ADLS is Azure Data Lake Storage. The Gravitino Spark, Flink, and Trino connectors consume vended credentials automatically for these catalogs.

Quick Start

Vend scoped S3 credentials to Spark through the IRC. Create the catalog through the Gravitino REST catalog API:

curl -X POST http://localhost:8090/api/metalakes/{metalake}/catalogs \
-H "Content-Type: application/json" \
-d '{
"name": "iceberg_catalog",
"type": "RELATIONAL",
"provider": "lakehouse-iceberg",
"properties": {
"catalog-backend": "jdbc",
"uri": "jdbc:postgresql://{postgres_host}:5432/{database}",
"jdbc-driver": "org.postgresql.Driver",
"jdbc-user": "{jdbc_user}",
"jdbc-password": "{jdbc_password}",
"jdbc-initialize": "true",
"warehouse": "s3://{bucket_name}/{warehouse_path}",
"io-impl": "org.apache.iceberg.aws.s3.S3FileIO",
"credential-providers": "s3-token",
"s3-access-key-id": "{access_key_id}",
"s3-secret-access-key": "{secret_access_key}",
"s3-region": "{region_name}",
"s3-role-arn": "{role_arn}"
}
}'

Point Spark at the IRC and request vended credentials with the delegation header:

./bin/spark-sql -v \
--packages org.apache.iceberg:iceberg-spark-runtime-3.5_2.12:1.10.0,org.apache.iceberg:iceberg-aws-bundle:1.10.0 \
--conf spark.sql.extensions=org.apache.iceberg.spark.extensions.IcebergSparkSessionExtensions \
--conf spark.sql.catalog.rest=org.apache.iceberg.spark.SparkCatalog \
--conf spark.sql.catalog.rest.type=rest \
--conf spark.sql.catalog.rest.uri=http://127.0.0.1:9001/iceberg/ \
--conf spark.sql.catalog.rest.prefix=iceberg_catalog \
--conf spark.sql.catalog.rest.header.X-Iceberg-Access-Delegation=vended-credentials

The role in s3-role-arn needs a trust policy and S3 permissions before this works. See s3-token.

For Trino instead of Spark:

connector.name=iceberg
iceberg.catalog.type=rest
iceberg.rest-catalog.uri=http://127.0.0.1:9001/iceberg/
iceberg.rest-catalog.prefix=iceberg_catalog
iceberg.rest-catalog.vended-credentials-enabled=true
fs.native-s3.enabled=true
s3.region={region_name}

For the full Trino setup, see Connect Trino to the IRC.

Setting Properties

Credential vending properties go in the catalog's properties map when you create it through the Gravitino REST catalog API, alongside the warehouse location and the other catalog settings.

The Gravitino Iceberg REST Catalog (IRC) is the exception. It can also read a catalog straight from gravitino.conf, using the same property names prefixed with gravitino.iceberg-rest.:

s3-role-arn                             # as a catalog property
gravitino.iceberg-rest.s3-role-arn # in gravitino.conf

Catalogs defined in gravitino.conf are not registered in a metalake, so Gravitino access control does not apply to them. Privileges are granted on catalogs in a metalake, and there is nothing to grant them on.

General Configurations

PropertyDescriptionDefault valueRequired
credential-providersThe credential provider types, separated by comma. If omitted, Gravitino infers some providers from the other properties present.(none)Yes, unless inferred
credential-cache-expire-ratioRatio of the credential's expiration time when Gravitino removes the credential from the cache.0.15No
credential-cache-max-sizeMax size for the credential cache.10000No

Values for credential-providers

ValueStorageVends
s3-tokenS3Temporary STS credentials, scoped to the table path
aws-irsaS3Credentials from an IAM role for service accounts, for EKS
s3-secret-keyS3The configured static access key and secret
oss-tokenOSSTemporary STS credentials, scoped to the table path
oss-secret-keyOSSThe configured static access key and secret
adls-tokenADLSA user delegation SAS token
azure-account-keyADLSThe configured static storage account key
gcs-tokenGCSA downscoped access token
jdbc-user-passwordJDBCThe configured JDBC username and password

Each value has its own properties, listed in the sections below. To vend for more than one storage type on a catalog, separate values with a comma. Custom providers can be added by implementing CredentialProvider, described under Custom Credentials.

When credential-providers Is Omitted

If a catalog does not set credential-providers, Gravitino infers providers from the credential properties present:

Properties presentProvider enabled
s3-access-key-id and s3-secret-access-keys3-secret-key
oss-access-key-id and oss-secret-access-keyoss-secret-key
azure-storage-account-name and azure-storage-account-keyazure-account-key
gcs-service-account-filegcs-token

JDBC catalogs additionally infer jdbc-user-password from jdbc-user and jdbc-password.

Four providers have no inference rule and must always be set explicitly: s3-token, oss-token, adls-token, and aws-irsa. In particular, setting s3-role-arn without credential-providers does not enable s3-token. The catalog falls back to s3-secret-key and vends the static access key instead, which is long-lived and not scoped to the table path. Set credential-providers explicitly whenever you want token-based vending.

S3

s3-token

Gravitino calls STS AssumeRole and returns temporary credentials scoped to the table path.

PropertyDescriptionDefault valueRequired
s3-role-arnARN of the role Gravitino assumes, in the form arn:aws:iam::{account_id}:role/{role_name}.(none)Yes
s3-regionRegion of the S3 service, like us-west-2.(none)No
s3-token-expire-in-secsSession lifetime of the vended credentials. Cannot exceed the role's maximum session duration.3600No
s3-external-idExternal ID passed on AssumeRole, for cross-account trust policies that require one.(none)No
s3-token-service-endpointAlternative STS endpoint, for S3-compatible storage such as MinIO.(none)No

Also set s3-access-key-id and s3-secret-access-key. Gravitino uses them to call AssumeRole, not to reach data, and they are never sent to the engine.

Trust Policy on the Role

The role in s3-role-arn must allow the s3-access-key-id principal to assume it. Without this, AssumeRole is rejected and no credential is vended.

{
"Version": "2012-10-17",
"Statement": [{
"Effect": "Allow",
"Principal": { "AWS": "arn:aws:iam::{account_id}:user/{gravitino_user}" },
"Action": "sts:AssumeRole"
}]
}

Permission Policy on the Role

The vended credentials inherit this policy, narrowed to the table path. Without S3 access to the warehouse prefix, the credentials are vended but cannot read or write.

{
"Version": "2012-10-17",
"Statement": [
{
"Effect": "Allow",
"Action": ["s3:GetObject", "s3:PutObject", "s3:DeleteObject"],
"Resource": "arn:aws:s3:::{bucket_name}/{warehouse_path}/*"
},
{
"Effect": "Allow",
"Action": ["s3:ListBucket", "s3:GetBucketLocation"],
"Resource": "arn:aws:s3:::{bucket_name}"
}
]
}

aws-irsa

For Gravitino running on EKS. Instead of an access key, Gravitino uses its pod's IAM role to call AssumeRole, so no static keys exist anywhere in the setup.

PropertyDescriptionDefault valueRequired
s3-role-arnARN of the role to assume, in the form arn:aws:iam::{account_id}:role/{role_name}.(none)For path-scoped credentials
s3-regionAWS region for STS operations.(none)No
s3-token-expire-in-secsSession lifetime of the vended credentials. Cannot exceed the role's maximum session duration.3600No
s3-token-service-endpointAlternative STS endpoint, for S3-compatible storage.(none)No

Set s3-role-arn to get credentials scoped to the table path, with an IAM policy generated per table covering its data, metadata, and write locations. Without it, the vended credentials carry the full permissions of the pod's role.

The role in s3-role-arn needs the same two policies as s3-token above, except the trust policy names the pod's IAM role rather than an IAM user.

IRSA itself must already be configured on the pod's Kubernetes service account. When it is, EKS injects a signed service account token into the pod and sets AWS_WEB_IDENTITY_TOKEN_FILE to its path, which is what the AWS SDK uses to obtain credentials. If vending fails, check that this variable is present in the pod.

s3-secret-key

Returns the catalog's configured access key and secret to the client, unchanged, in the loadTable response.

The key is long-lived, carries whatever permissions its IAM user has, and is not scoped to the table path. Any client that can load a table receives it, and it stays valid after the query finishes. Prefer s3-token, which returns temporary credentials scoped to the table path. Use s3-secret-key to confirm the vending path works before configuring a role.

PropertyDescriptionDefault valueRequired
s3-access-key-idThe static access key ID used to access S3 data.(none)Yes
s3-secret-access-keyThe static secret access key used to access S3 data.(none)Yes

OSS

oss-token

Gravitino calls Alibaba Cloud STS AssumeRole and returns temporary credentials scoped to the table path.

Also set oss-access-key-id and oss-secret-access-key. Gravitino uses them to call AssumeRole, not to reach data, and they are never sent to the engine.

PropertyDescriptionDefault valueRequired
oss-access-key-idThe static access key ID used to access OSS data.(none)Yes
oss-secret-access-keyThe static secret access key used to access OSS data.(none)Yes
oss-role-arnThe ARN of the role to access the OSS data.(none)Yes
oss-regionThe region of the OSS service, like oss-cn-hangzhou, only used when credential-providers is oss-token.(none)No
oss-external-idThe OSS external id to generate the token.(none)No
oss-token-expire-in-secsThe OSS security token expire time in secs.3600No

Trust Policy on the RAM Role

The role in oss-role-arn must allow the oss-access-key-id principal to assume it.

{
"Version": "1",
"Statement": [{
"Effect": "Allow",
"Action": "sts:AssumeRole",
"Principal": { "RAM": ["acs:ram::{account_id}:user/{gravitino_user}"] }
}]
}

Permission Policy on the RAM Role

The vended credentials inherit this policy, narrowed to the table path.

{
"Version": "1",
"Statement": [
{
"Effect": "Allow",
"Action": ["oss:GetObject", "oss:PutObject", "oss:DeleteObject"],
"Resource": "acs:oss:*:*:{bucket_name}/{warehouse_path}/*"
},
{
"Effect": "Allow",
"Action": ["oss:ListObjects", "oss:GetBucketInfo"],
"Resource": "acs:oss:*:*:{bucket_name}"
}
]
}

oss-secret-key

Returns the catalog's configured access key and secret to the client, unchanged.

The key is long-lived, carries whatever permissions its RAM user has, and is not scoped to the table path. Any client that can load a table receives it, and it stays valid after the query finishes. Prefer oss-token. Use oss-secret-key to confirm the vending path works before configuring a role.

PropertyDescriptionDefault valueRequired
oss-access-key-idThe static access key ID used to access OSS data.(none)Yes
oss-secret-access-keyThe static secret access key used to access OSS data.(none)Yes

ADLS

adls-token

Gravitino requests an Azure user delegation SAS and returns it to the client, scoped to the table path.

Azure grants access through role assignments rather than policy documents. The Microsoft Entra ID service principal identified by azure-tenant-id, azure-client-id, and azure-client-secret needs two roles:

  • Storage Blob Delegator, assigned at the storage account, which allows it to request the user delegation key that signs the SAS.
  • Storage Blob Data Contributor, assigned on the container or the warehouse path, which determines what the vended SAS can do. Use Storage Blob Data Reader for read-only access.

Without the delegator role the SAS cannot be issued at all. Without a data role the SAS is issued but grants nothing.

PropertyDescriptionDefault valueRequired
azure-storage-account-nameThe static storage account name used to access ADLS data.(none)Yes
azure-tenant-idAzure Active Directory (AAD) tenant ID.(none)Yes
azure-client-idAzure Active Directory (AAD) client ID used for authentication.(none)Yes
azure-client-secretAzure Active Directory (AAD) client secret used for authentication.(none)Yes
adls-token-expire-in-secsThe ADLS SAS token expire time in secs.3600No

azure-account-key

Returns the catalog's configured storage account key to the client, unchanged.

A storage account key grants full access to every container in the storage account, not just the warehouse path, and it does not expire. Any client that can load a table receives it. Prefer adls-token, which is scoped and time-limited. Use azure-account-key to confirm the vending path works before configuring a service principal.

PropertyDescriptionDefault valueRequired
azure-storage-account-nameThe static storage account name used to access ADLS data.(none)Yes
azure-storage-account-keyThe static storage account key used to access ADLS data.(none)Yes

GCS

gcs-token

Gravitino downscopes its own credentials using GCS credential access boundaries and returns a token scoped to the table path.

There is no role to assume. The identity is the service account in gcs-service-account-file, or the application default credentials when that is unset. Grant that service account Storage Object User (roles/storage.objectUser) on the warehouse bucket, or Storage Object Viewer for read-only access. Downscoping narrows from those permissions, so the vended token can never exceed what the service account itself holds.

PropertyDescriptionDefault valueRequired
gcs-service-account-fileThe location of the GCS credential file.GCS Application default credential.No

For the IRC, ensure that the credential file is accessible by that server. For example, the server may be running on a GCE machine, or you may set the environment variable export GOOGLE_APPLICATION_CREDENTIALS=/xx/application_default_credentials.json even when gcs-service-account-file is already configured.

Requesting Vended Credentials

How a client asks depends on which interface it uses.

Over the IRC

Credentials are vended only when the client asks for them. Spark, Flink, and other IRC clients ask with a header:

X-Iceberg-Access-Delegation: vended-credentials

In Spark, set it as a catalog config key:

spark.sql.catalog.{name}.header.X-Iceberg-Access-Delegation=vended-credentials

Trino asks with a catalog property instead, and sends the header for you:

iceberg.rest-catalog.vended-credentials-enabled=true

Over the Gravitino REST Catalog API

Hive, Glue, JDBC, Paimon, and Fileset catalogs are reached through the Gravitino REST catalog API, which has no delegation header. Credential properties are hidden from the catalog GET response, so clients fetch them from the Gravitino credential endpoint instead. It works for any metadata object:

GET /api/metalakes/{metalake}/objects/{type}/{full_name}/credentials

For a catalog, {type} is catalog and {full_name} is the catalog name:

GET /api/metalakes/{metalake}/objects/catalog/{catalog}/credentials

The Gravitino Spark and Flink connectors call this for you and inject the returned credentials, so no client configuration is needed.

Read or Write Scope

Over the IRC, a credential is vended for writing when the caller is entitled to modify the table, and for reading otherwise. Narrowing the caller's roles with the X-Gravitino-Active-Roles header narrows this as well, so a caller whose active roles no longer carry MODIFY_TABLE is vended a read-only credential. See Narrowing Access with Active Roles.

Custom Credentials

Gravitino supports custom credentials. You can implement the org.apache.gravitino.credential.CredentialProvider interface to support custom credentials, and place the corresponding jar in the classpath of the IRC or the Fileset catalog.

Deployment

The credential provider implementations ship in separate jars. Whichever component vends the credentials needs the right jar on its classpath, or the provider cannot be created and no credentials are vended.

Vending componentJarClasspath
IRCgravitino-iceberg-{cloud}-bundleSee Deployment; it differs by deployment mode
Iceberg cataloggravitino-iceberg-{cloud}-bundlecatalogs/lakehouse-iceberg/libs/
Fileset cataloggravitino-{cloud}-bundlecatalogs/fileset/libs/
Hive cataloggravitino-{cloud}catalogs/hive/libs/
Glue cataloggravitino-awscatalogs/glue/libs/
Paimon cataloggravitino-aws for S3, gravitino-aliyun for OSScatalogs/lakehouse-paimon/libs/

Substitute {cloud} with aws, gcp, aliyun, or azure. Note the two jar families: the -bundle variants also carry Hadoop and cloud SDK packages, which the Fileset catalog and the IRC need. The Hive, Glue, and Paimon catalogs only vend credentials, so they take the plain gravitino-{cloud} jar.

The Gravitino Iceberg cloud bundle jars already include the Iceberg cloud bundle jars, so there is no need to download and include those separately.

Vending JDBC user and password requires no additional jar.

Bundle jars on Maven Central:

Upgrading From a Release Earlier Than 1.3.0

Sensitive catalog properties such as s3-access-key-id, s3-secret-access-key, jdbc-user, and jdbc-password are excluded from GET /api/metalakes/{metalake}/catalogs/{catalog}. Clients written against earlier releases that read those properties directly lose access to them.

For a zero-downtime migration, set the following in gravitino.conf:

gravitino.catalog.credential.backfillToProperties = true

The Gravitino server then re-includes the hidden properties in its catalog GET responses. Turn it off once all clients use the Gravitino credential endpoint, since it exposes credentials in plaintext.