Skip to main content

Audit log

The audit log is the server's record of who did what to which object, whether it worked, and where the request came from. Changes, reads, refused requests, and requests that failed all produce a line in it, with the few exceptions described in Where Records Come From.

The Helm chart turns the audit log on and writes it to the logs volume of each server pod. Each line is one record, in a tab-separated text format by default or in JSON when a pipeline needs structured input.

Quick start​

  1. Confirm the audit log is on. The chart sets gravitino.audit.enabled to true, so an install from the chart is already writing one. A server configured by hand writes nothing until that property is true.

  2. Watch it. The log lives in a directory named for the pod, under the logs volume.

    Follow the audit log
    kubectl exec -n {namespace} deploy/{release}-gravitino -c gravitino -- \
    sh -c 'tail -f /opt/gravitino/logs/$HOSTNAME/gravitino_audit.log'
  3. Find the changes that matter. Operation types are fixed names, so a plain search finds every grant, revoke, and ownership change in the current file.

    Find grants, revokes, and ownership changes
    kubectl exec -n {namespace} deploy/{release}-gravitino -c gravitino -- \
    sh -c 'grep -E "GRANT_|REVOKE_|OVERRIDE_PRIVILEGES|SET_OWNER" /opt/gravitino/logs/$HOSTNAME/gravitino_audit.log'
  4. Plan where it goes next. The chart keeps rotated files for 30 days or 10 GB, whichever comes first. Anything that must be kept longer, or searched across pods, belongs in a log pipeline, as described in Sending the Log Elsewhere.

What a record is​

Where records come from​

Records come from two places. When a request performs an operation on an object, such as loading a table or granting a role, the operation writes the record, and its operation type names what was done. When a request performs no such operation, the HTTP layer writes the record instead, with operation type UNKNOWN and the method, path, and status code in its custom info. Requests to /configs, to the metrics endpoints, and to parts of the API that are not operations on an object are recorded this way.

The HTTP layer writes its record only when no operation already has, so a request usually appears once. When an operation succeeds and the response then ends in a 4xx or 5xx status anyway, the log keeps both the operation's SUCCESS record and an HTTP-level FAILURE record, because both happened. A request that sets the owner of several objects at once writes a SET_OWNER record for each object.

Some requests are not recorded at all:

  • Health probes: /api/health, /health, and /health.html on the server port, /iceberg/health on the IRC port, /scim/health on the SCIM endpoint, and the paths under each.
  • Successful SCIM requests that do not change or read a user or group. SCIM requests that fail are recorded.
  • The static files of the server's built-in web UI, at / and under /ui.

Success and failure​

An operation that throws writes a record with status FAILURE and the same operation type a success would have had. A request with no operation that ends in a 4xx or 5xx status writes an HTTP-level FAILURE record, which is how requests rejected before reaching an operation appear, such as a request whose credentials fail authentication.

When the server refuses a request because the caller lacks a privilege, the record has operation type AUTHORIZATION_DENIAL instead of the operation that was refused. Its identifier is the object the privilege check evaluated, and its custom info names the API method and the privilege expression the caller did not satisfy. A request the Gravitino IRC endpoint refuses for lack of a privilege is recorded differently: as an HTTP-level FAILURE record with http.status=403, source GRAVITINO_ICEBERG_REST_SERVER, and the refused path in http.uri, but no identifier.

What a record leaves out​

A record names the operation and the object, not the content of the request. In particular:

  • An ALTER_* record does not say what changed, and a CREATE_* record does not carry the new object's definition or properties.
  • GRANT_PRIVILEGES, REVOKE_PRIVILEGES, and OVERRIDE_PRIVILEGES identify the role, not the privileges or the object they apply to.
  • ASSOCIATE_TAGS_FOR_METADATA_OBJECT and ASSOCIATE_POLICIES_FOR_METADATA_OBJECT identify the object, not the tags or policies added or removed.
  • SET_OWNER identifies the object, not the new owner.
  • RUN_JOB identifies the new job, not the template it ran.
  • GET_FILESET_LOCATION records that a path was resolved, not which file operation GVFS went on to perform.

Role grants to users and groups are the exception: they carry the role names. SCIM changes to users and groups carry a summary of the change, as described in Custom Info.

The Gravitino MCP server keeps an audit log of its own tool calls in its own pod. The API requests it makes to the server on a user's behalf appear in this log like any other client's.

Record format​

Text format​

The default format writes eight tab-separated fields per line. A record with no custom info has an empty last field.

Example text records
[2026-10-07 19:31:27.492]	admin	UNKNOWN	null	SUCCESS	GRAVITINO_SERVER	10.42.0.46	{metalake=acme, http.method=GET, http.uri=/api/system/iceberg-rest, http.status=200}
[2026-10-07 19:31:27.499] admin LIST_CATALOG acme SUCCESS GRAVITINO_SERVER 10.42.0.46 {count=15}
[2026-10-07 19:31:27.546] admin LOAD_CATALOG acme.iceberg_s3 SUCCESS GRAVITINO_SERVER 10.42.0.46
[2026-10-07 19:31:29.190] unknown UNKNOWN null SUCCESS GRAVITINO_SERVER 10.42.0.29 {http.method=GET, http.uri=/configs, http.status=200}
PositionFieldContent
1TimestampWhen the operation finished, as [yyyy-MM-dd HH:mm:ss.SSS] in the server's time zone, with no offset.
2UserThe authenticated user who made the request. See User.
3Operation typeWhat was done, from the Operation Catalog.
4IdentifierThe object it was done to, or null. See Identifier.
5StatusSUCCESS or FAILURE.
6SourceThe endpoint that served the request. See Source.
7Client addressWhere the request came from. See Client Address.
8Custom infoExtra facts as {key=value, ...}, or empty. See Custom Info.

The timestamp carries no zone, so its meaning depends on the server's time zone. Pods from the chart run in UTC, and the timestamps above, taken from a live server, are UTC.

JSON format​

JsonAuditFormatter writes one JSON object per line. It carries the same facts as the text format under named keys, adds the zone offset to the timestamp, and escapes values, so a parser never has to guess where a field ends. Use it for anything a machine will read. The second line of the text example above, in JSON:

Example JSON record
{"timestamp":"2026-10-07T19:31:27.499Z","user":"admin","operation":"LIST_CATALOG","operationType":"LIST_CATALOG","identifier":"acme","status":"SUCCESS","operationStatus":"SUCCESS","eventSource":"GRAVITINO_SERVER","remoteAddress":"10.42.0.46","customInfo":{},"resultCount":15}
KeyContent
timestampWhen the operation finished, ISO 8601 with milliseconds and the zone offset.
userSame as the text format's user field.
operationTypeSame as the text format's operation type field. Key on this one.
operationAn older, coarser name for the operation kept for existing consumers. Several types share one value here, such as LIST_USERS_PAGED and COUNT_USERS, which both appear as LIST_USERS.
identifierSame as the text format's identifier field, or null.
operationStatusSame as the text format's status field.
statusThe same outcome under its older name.
eventSourceSame as the text format's source field.
remoteAddressSame as the text format's client address field.
customInfoAn object of string keys and values, empty when there are none.
resultCountThe number of items a list operation returned. Present only on list operations, and not on SCIM list records.

Switch formats in the chart values and upgrade the release:

Switch to the JSON format
gravitino:
audit:
formatter:
className: org.apache.gravitino.audit.JsonAuditFormatter

Fields​

User​

The user is the authenticated principal the request ran as. A request served under the simple authenticator without a user name records anonymous. An HTTP-level record for a request with no authenticated principal, such as an unauthenticated call to /configs or a request that failed authentication, records unknown.

A service that calls the server under its own user, such as the Trino connector or an ingestion job, is recorded under that user. A person signed in under the same user cannot be told apart from it in the log, so give each service a user of its own.

Identifier​

The identifier is the full dotted name of the object, starting with the metalake. Users, groups, roles, tags, policies, job templates, and jobs live in a reserved system catalog of the metalake, and their identifiers say so.

ObjectIdentifier
Metalake{metalake}
Catalog{metalake}.{catalog}
Schema{metalake}.{catalog}.{schema}
Table, view, fileset, topic, model, function{metalake}.{catalog}.{schema}.{name}
Table served through IRC{metalake}.{catalog}.{namespace_levels}.{table}
User{metalake}.system.user.{user}
Group{metalake}.system.group.{group}
Role{metalake}.system.role.{role}
Tag{metalake}.system.tag.{tag}
Policy{metalake}.system.policy.{policy}
Job template{metalake}.system.job_template.{template}
Job{metalake}.system.job.{job_id}
User or group changed through SCIM_instance.system.user.{user} or _instance.system.group.{group}

A list operation identifies the container it listed, so LIST_TABLE carries the schema and LIST_CATALOG carries the metalake. HTTP-level records have no identifier and show null.

Source​

The source names the endpoint that served the request.

SourceRequests
GRAVITINO_SERVERThe Gravitino REST API, the SCIM endpoint, /configs, and the metrics endpoints on the server port.
GRAVITINO_ICEBERG_REST_SERVERThe Gravitino IRC endpoint, and the metrics endpoints on its port.
GRAVITINO_LANCE_REST_SERVERThe Lance REST endpoint, which writes HTTP-level records only. The Gravitino API calls it makes to serve a request appear as GRAVITINO_SERVER records.

Client address​

The client address is the first entry of the request's X-Forwarded-For header when there is one, and the address of the connection otherwise. Inside the cluster that is the caller's pod IP. Behind an ingress that sets X-Forwarded-For, it is the address the ingress received the request from. When the load balancer in front of the ingress does not preserve the caller's address, every request from outside the cluster records the same internal address, such as 10.42.0.1, and the field no longer tells callers apart.

The server takes X-Forwarded-For as given. A client that can reach the server without passing through a proxy that rewrites the header can put any address there, so treat the client address as trustworthy only when every route to the server goes through such a proxy.

Two record types differ slightly. IRC operation records without X-Forwarded-For use the connection's host name, which is its address unless the server resolves host names. SCIM operation records with no address leave the field empty rather than writing unknown.

Custom info​

Custom info carries the facts that belong to one kind of record. Keys appear only on the records they apply to.

KeysPresent On
Query parameters of the requestEvery record from a request that had a query string, except SCIM requests. Up to 50 parameter names are kept, and each value is cut at 256 characters with ...(truncated) appended.
countList operations, as the number of items returned. The JSON format carries it as resultCount instead, except on SCIM list records, where it stays in custom info.
roleNamesGRANT_USER_ROLES, REVOKE_USER_ROLES, GRANT_GROUP_ROLES, and REVOKE_GROUP_ROLES, as a comma-separated list.
source, resourceType, id, externalIdRequests from an identity provider through SCIM. source is always scim, resourceType is User or Group, and id and externalId are the SCIM identifiers.
changesSCIM ALTER_USER and ALTER_GROUP: put for a full replace, or the patched attributes separated by ;, such as active=false.
membersAdded, membersRemovedSCIM ALTER_GROUP, as comma-separated SCIM member IDs.
reasonFailed SCIM requests, as the error message cut at 512 characters, or the error type when there is no message.
Request headersEvery operation record with source GRAVITINO_ICEBERG_REST_SERVER, one key per HTTP header the client sent. The Authorization header is redacted.
icebergEncryption.*Table operations on encrypted Iceberg tables, describing the encryption decision: the policy, key provider, key ID, enforcement, and reason.
auth.method, auth.expressionAUTHORIZATION_DENIAL: the API method refused and the privilege expression it was evaluated against.
http.method, http.uri, http.statusHTTP-level records with operation type UNKNOWN.

Redaction​

Before a record is written, the value of any custom info key that looks sensitive is replaced with ***. A key is sensitive when it is one of authorization, cookie, x-amz-security-token, s3.access-key-id, or jdbc-password, or when, ignoring case and with everything but letters and digits removed, it contains password, secret, token, credential, apikey, accesskey, privatekey, auth, or signature. The fixed keys http.method, http.uri, http.status, auth.method, and auth.expression are never redacted.

Redaction looks at key names only. A secret passed under an innocuous name, such as a credential in a query parameter called q, is written as sent.

Operation catalog​

Every operation type a record can carry is listed here, grouped by the object it acts on. Each type appears with status SUCCESS or FAILURE; the table says when the record is written. Types marked IRC only come from the Gravitino IRC endpoint; every other type comes from the Gravitino REST API, and the types an IRC request shares with it carry source GRAVITINO_ICEBERG_REST_SERVER when they come from IRC.

Metalakes​

The identifier is {metalake}. LIST_METALAKE has no identifier.

Operation TypeRecorded When
CREATE_METALAKEA metalake is created.
ALTER_METALAKEA metalake's name, comment, or properties change.
DROP_METALAKEA metalake is dropped.
LOAD_METALAKEA metalake is loaded.
LIST_METALAKEMetalakes are listed. Carries the number returned.
ENABLE_METALAKEA metalake is enabled.
DISABLE_METALAKEA metalake is disabled.

Catalogs​

The identifier is {metalake}.{catalog}, or {metalake} for LIST_CATALOG.

Operation TypeRecorded When
CREATE_CATALOGA catalog is created.
ALTER_CATALOGA catalog's name, comment, or properties change.
DROP_CATALOGA catalog is dropped.
LOAD_CATALOGA catalog is loaded.
LIST_CATALOGCatalogs in a metalake are listed. Carries the number returned.
ENABLE_CATALOGA catalog is enabled.
DISABLE_CATALOGA catalog is disabled.

Schemas​

The identifier is {metalake}.{catalog}.{schema}, or the catalog for LIST_SCHEMA. IRC requests on namespaces record the same operation types with source GRAVITINO_ICEBERG_REST_SERVER.

Operation TypeRecorded When
CREATE_SCHEMAA schema or IRC namespace is created.
ALTER_SCHEMAA schema's properties or an IRC namespace's properties change.
DROP_SCHEMAA schema or IRC namespace is dropped.
LOAD_SCHEMAA schema or IRC namespace is loaded.
LIST_SCHEMASchemas or IRC namespaces are listed. Carries the number returned.
SCHEMA_EXISTSAn IRC client checks whether a namespace exists. IRC only.

Tables​

The identifier is {metalake}.{catalog}.{schema}.{table}, or the schema for LIST_TABLE. For IRC requests the namespace levels take the place of {schema}.

Operation TypeRecorded When
CREATE_TABLEA table is created.
ALTER_TABLEA table changes. Through IRC, a table commit.
RENAME_TABLEAn IRC client renames a table. IRC only; a rename through the Gravitino API is ALTER_TABLE.
DROP_TABLEA table is dropped.
PURGE_TABLEA table is purged, removing its data along with its metadata.
LOAD_TABLEA table is loaded.
LIST_TABLETables in a schema are listed. Carries the number returned.
REGISTER_TABLEAn IRC client registers an existing table from its metadata file. IRC only.
TABLE_EXISTSAn IRC client checks whether a table exists. IRC only.
LOAD_TABLE_CREDENTIALAn IRC client fetches vended storage credentials for a table. IRC only.
PLAN_TABLE_SCANAn IRC client asks the server to plan a table scan. IRC only.

Partitions​

The identifier is the table. The partition name is not part of the record.

Operation TypeRecorded When
ADD_PARTITIONA partition is added.
DROP_PARTITIONA partition is dropped.
PURGE_PARTITIONA partition is purged.
LOAD_PARTITIONA partition is loaded.
PARTITION_EXISTSA caller checks whether a partition exists.
LIST_PARTITIONPartitions of a table are listed. Carries the number returned.
LIST_PARTITION_NAMESPartition names of a table are listed. Carries the number returned.

Views​

The identifier is {metalake}.{catalog}.{schema}.{view}, or the schema for LIST_VIEW.

Operation TypeRecorded When
CREATE_VIEWA view is created.
ALTER_VIEWA view changes through the Gravitino API.
REPLACE_VIEWAn IRC client replaces a view's definition. IRC only.
RENAME_VIEWAn IRC client renames a view. IRC only.
DROP_VIEWA view is dropped.
LOAD_VIEWA view is loaded.
LIST_VIEWViews in a schema are listed. Carries the number returned.
VIEW_EXISTSAn IRC client checks whether a view exists. IRC only.

Functions​

The identifier is {metalake}.{catalog}.{schema}.{function}, or the schema for the two list operations.

Operation TypeRecorded When
REGISTER_FUNCTIONA function is registered.
ALTER_FUNCTIONA function changes.
DROP_FUNCTIONA function is dropped.
GET_FUNCTIONA function is loaded.
LIST_FUNCTIONFunction names in a schema are listed. Carries the number returned.
LIST_FUNCTION_INFOSFunctions in a schema are listed with their details. Carries the number returned. A failed call is recorded as LIST_FUNCTION with status FAILURE.

Filesets​

The identifier is {metalake}.{catalog}.{schema}.{fileset}, or the schema for LIST_FILESET.

Operation TypeRecorded When
CREATE_FILESETA fileset is created.
ALTER_FILESETA fileset changes.
DROP_FILESETA fileset is dropped.
LOAD_FILESETA fileset is loaded.
LIST_FILESETFilesets in a schema are listed. Carries the number returned.
LIST_FILESET_FILESFiles under a fileset are listed. Carries the number returned.
GET_FILESET_LOCATIONA client resolves a path inside a fileset to its storage location, as GVFS does before it reads or writes a file.

Topics​

The identifier is {metalake}.{catalog}.{schema}.{topic}, or the schema for LIST_TOPIC.

Operation TypeRecorded When
CREATE_TOPICA topic is created.
ALTER_TOPICA topic changes.
DROP_TOPICA topic is dropped.
LOAD_TOPICA topic is loaded.
LIST_TOPICTopics in a schema are listed. Carries the number returned.

Models​

The identifier is {metalake}.{catalog}.{schema}.{model}, including for version operations, which do not record the version number or alias. LIST_MODEL carries the schema.

Operation TypeRecorded When
REGISTER_MODELA model is registered.
REGISTER_AND_LINK_MODEL_VERSIONA model is registered and its first version linked in one call.
ALTER_MODELA model changes.
DELETE_MODELA model is deleted.
GET_MODELA model is loaded.
LIST_MODELModels in a schema are listed. Carries the number returned.
LINK_MODEL_VERSIONA version is linked to a model.
ALTER_MODEL_VERSIONA model version changes.
DELETE_MODEL_VERSIONA model version is deleted.
GET_MODEL_VERSIONA model version is loaded by number or alias.
GET_MODEL_VERSION_URIThe storage URI of a model version is resolved.
LIST_MODEL_VERSIONSVersion numbers of a model are listed. Carries the number returned.
LIST_MODEL_VERSION_INFOSVersions of a model are listed with their details. Carries the number returned.

Statistics​

The identifier is the table or other object that owns the statistics.

Operation TypeRecorded When
UPDATE_STATISTICSStatistics on an object are written.
LIST_STATISTICSStatistics on an object are listed. Carries the number returned.
DROP_STATISTICSStatistics on an object are dropped.
UPDATE_PARTITION_STATISTICSPartition statistics on a table are written.
LIST_PARTITION_STATISTICSPartition statistics on a table are listed. Carries the number returned.
DROP_PARTITION_STATISTICSPartition statistics on a table are dropped.

Tags​

The identifier is {metalake}.system.tag.{tag}. ASSOCIATE_TAGS_FOR_METADATA_OBJECT, GET_TAG_FOR_METADATA_OBJECT, and the two LIST_TAGS_*_FOR_METADATA_OBJECT types carry the object's identifier instead, and LIST_TAG and LIST_TAGS_INFO carry {metalake}.

Operation TypeRecorded When
CREATE_TAGA tag is created.
ALTER_TAGA tag changes.
DELETE_TAGA tag is deleted.
GET_TAGA tag is loaded.
LIST_TAGTag names in a metalake are listed. Carries the number returned.
LIST_TAGS_INFOTags in a metalake are listed with their details. Carries the number returned.
ASSOCIATE_TAGS_FOR_METADATA_OBJECTTags are added to or removed from an object.
GET_TAG_FOR_METADATA_OBJECTOne tag on an object is loaded.
LIST_TAGS_FOR_METADATA_OBJECTTag names on an object are listed. Carries the number returned.
LIST_TAGS_INFO_FOR_METADATA_OBJECTTags on an object are listed with their details. Carries the number returned.
LIST_METADATA_OBJECTS_FOR_TAGObjects carrying a tag are listed. Carries the number returned.

Policies​

The identifier is {metalake}.system.policy.{policy}. ASSOCIATE_POLICIES_FOR_METADATA_OBJECT, GET_POLICY_FOR_METADATA_OBJECT, and LIST_POLICY_INFOS_FOR_METADATA_OBJECT carry the object's identifier instead, and LIST_POLICY and LIST_POLICY_INFO carry {metalake}.

Operation TypeRecorded When
CREATE_POLICYA policy is created.
ALTER_POLICYA policy changes.
DELETE_POLICYA policy is deleted.
ENABLE_POLICYA policy is enabled.
DISABLE_POLICYA policy is disabled.
GET_POLICYA policy is loaded.
LIST_POLICYPolicy names in a metalake are listed. Carries the number returned.
LIST_POLICY_INFOPolicies in a metalake are listed with their details. Carries the number returned.
ASSOCIATE_POLICIES_FOR_METADATA_OBJECTPolicies are attached to or detached from an object.
GET_POLICY_FOR_METADATA_OBJECTOne policy on an object is loaded.
LIST_POLICY_INFOS_FOR_METADATA_OBJECTPolicies on an object are listed with their details. Carries the number returned.
LIST_METADATA_OBJECTS_FOR_POLICYObjects a policy is attached to are listed. Carries the number returned.

Users​

The identifier is {metalake}.system.user.{user}, or {metalake} for the list and count operations. Requests from an identity provider through SCIM record the same operation types with source=scim in the custom info, and their identifiers use the fixed name _instance in place of the metalake, as in _instance.system.user.{user}.

Operation TypeRecorded When
ADD_USERA user is added, directly or by SCIM provisioning.
REMOVE_USERA user is removed, directly or by SCIM deprovisioning.
ALTER_USERSCIM replaces or patches a user. Deactivation is recorded here as changes=active=false.
GET_USERA user is loaded by name.
GET_USER_BY_IDSCIM loads a user by its SCIM ID.
LIST_USERSUsers in a metalake are listed. Carries the number returned.
LIST_USERS_PAGEDOne page of users is listed, by the UI or by SCIM. Carries the number returned, which SCIM records keep in custom info.
LIST_USER_NAMESUser names in a metalake are listed. Carries the number returned.
COUNT_USERSUsers in a metalake are counted.
GRANT_USER_ROLESRoles are granted to a user. Carries the role names.
REVOKE_USER_ROLESRoles are revoked from a user. Carries the role names.

Groups​

The identifier is {metalake}.system.group.{group}, or {metalake} for the list and count operations. SCIM requests carry source=scim and the _instance identifier prefix, as for users.

Operation TypeRecorded When
ADD_GROUPA group is added, directly or by SCIM provisioning.
REMOVE_GROUPA group is removed, directly or by SCIM deprovisioning.
ALTER_GROUPSCIM replaces or patches a group. Carries the member IDs added and removed.
GET_GROUPA group is loaded by name.
GET_GROUP_BY_IDSCIM loads a group by its SCIM ID.
LIST_GROUPSGroups in a metalake are listed. Carries the number returned.
LIST_GROUPS_PAGEDOne page of groups is listed, by the UI or by SCIM. Carries the number returned, which SCIM records keep in custom info.
LIST_GROUP_NAMESGroup names in a metalake are listed. Carries the number returned.
COUNT_GROUPSGroups in a metalake are counted.
GRANT_GROUP_ROLESRoles are granted to a group. Carries the role names.
REVOKE_GROUP_ROLESRoles are revoked from a group. Carries the role names.

Roles and privileges​

The identifier is {metalake}.system.role.{role}, or {metalake} for LIST_ROLE_NAMES. The privilege operations identify the role that changed, not the object the privilege applies to.

Operation TypeRecorded When
CREATE_ROLEA role is created.
DELETE_ROLEA role is deleted.
GET_ROLEA role is loaded.
LIST_ROLE_NAMESRole names in a metalake are listed. Carries the number returned.
GRANT_PRIVILEGESPrivileges on an object are granted to a role.
REVOKE_PRIVILEGESPrivileges on an object are revoked from a role.
OVERRIDE_PRIVILEGESA role's privileges are replaced as a set.

Ownership​

The identifier is the object whose owner is read or set.

Operation TypeRecorded When
GET_OWNERAn object's owner is read.
SET_OWNERAn object's owner is set. A request that sets the owner of several objects writes one record per object.

Jobs​

Template operations carry {metalake}.system.job_template.{template}, job operations carry {metalake}.system.job.{job_id}, and the two list operations carry {metalake}.

Operation TypeRecorded When
REGISTER_JOB_TEMPLATEA job template is registered.
ALTER_JOB_TEMPLATEA job template changes.
DELETE_JOB_TEMPLATEA job template is deleted.
GET_JOB_TEMPLATEA job template is loaded.
LIST_JOB_TEMPLATESJob templates in a metalake are listed. Carries the number returned.
RUN_JOBA job is started from a template. The identifier is the new job's ID.
GET_JOBA job's status is read.
CANCEL_JOBA job is cancelled.
LIST_JOBSJobs in a metalake are listed. Carries the number returned.

Requests without an operation​

These two types come from the server itself rather than from an operation on an object.

Operation TypeRecorded When
AUTHORIZATION_DENIALThe server refuses a request because the caller lacks the privileges for it. Always FAILURE. The identifier is the object the check evaluated, and the custom info carries auth.method and auth.expression.
UNKNOWNA request that no operation recorded, written by the HTTP layer so that the request still appears once. No identifier. The custom info carries http.method, http.uri, and http.status.

Where the log lives​

Each server pod writes its own file, gravitino_audit.log, in a directory named for the pod on the logs volume:

Audit log paths
/opt/gravitino/logs/{pod_name}/gravitino_audit.log
/opt/gravitino/logs/{sub_path}/{pod_name}/gravitino_audit.log # when logPersistence.subPath is set

The file rotates every day and whenever it reaches 256 MB. Rotated files are compressed and named gravitino_audit_{yyyyMMdd}.{n}.log.gz beside it. A rotated file is deleted once it is older than logMaxAge or once the rotated audit files together exceed auditLogMaxTotalSize, whichever happens first, oldest first.

The logs volume is a persistent volume claim, so the log survives a pod being replaced. Setting logPersistence.enabled to false replaces it with a scratch volume that is lost with the pod. With more than one server replica, every pod writes to the same volume, which then has to be ReadWriteMany; each pod's records are in its own directory, and a complete picture means reading all of them.

ValueEffectDefault
gravitino.audit.enabledTurns the audit log on or off.true
gravitino.audit.formatter.classNameSelects the record format.org.apache.gravitino.audit.v2.SimpleFormatterV2
gravitino.log4j2Properties.auditLogMaxTotalSizeTotal size of rotated audit archives kept on the volume.10GB
gravitino.log4j2Properties.logMaxAgeAge after which rotated archives are deleted. Applies to the server, lineage, and search logs as well.30d
gravitino.logPersistence.sizeSize of the logs volume, shared by all four logs and by job staging files.14Gi
gravitino.logPersistence.subPathDirectory on the volume to write under, for releases that share one volume.(empty)

Sending the log elsewhere​

The audit log on the volume is bounded by the retention above and is split across pods. Keeping records for a compliance period, searching them, or alerting on them means shipping them to a log platform or SIEM. There are three ways to do it.

Write to standard output​

The audit log is written through a Log4j2 logger named gravitino.audit. Adding a console appender to that logger sends every record to the container's standard output as well as to the file, where any cluster log collector picks it up. Use the JSON format, so the collector can tell audit records apart from the server's own log lines by their shape.

Send audit records to standard output
gravitino:
audit:
formatter:
className: org.apache.gravitino.audit.JsonAuditFormatter
additionalLog4j2Properties:
appender.audit_console.type: Console
appender.audit_console.name: auditConsole
appender.audit_console.layout.type: PatternLayout
appender.audit_console.layout.pattern: "%msg%n"
logger.audit.name: gravitino.audit
logger.audit.appenderRef.audit_console.ref: auditConsole

Keep the logger.audit.name line even though the server's own configuration already sets it. The chart applies additionalLog4j2Properties to the metrics service's Log4j2 configuration as well, which has no audit logger of its own; without the name there, Log4j2 rejects the configuration and the metrics service fails to start.

Send the Audit Log to Datadog and Send the Audit Log to Splunk walk through this route end to end with each platform's own collector.

Read the file from a sidecar​

A log shipper running beside the server, added through gravitino.extraContainers with the gravitino-logs volume mounted at /opt/gravitino/logs, can tail gravitino_audit.log and forward each line. The file stays the single source and nothing about the server changes, at the cost of running and configuring the shipper.

Write a custom writer​

The writer is an interface, org.apache.gravitino.audit.AuditLogWriter. An implementation on the server classpath, named in gravitino.audit.writer.className, receives every formatted record and can deliver it anywhere. The writer runs on the request thread, as described in Delivery, so a slow writer slows every request. Redaction happens inside the two built-in formatters' output, not in the record, so a writer that reads customInfo() directly receives sensitive values unredacted and has to apply AuditLogRedactor itself.

Volume​

Every request writes a record, reads included, so the log grows with traffic and not just with change. Three sources write records continuously on an otherwise idle system:

  • The Trino connector. The connector refreshes its view of the metalake every gravitino.metadata.refresh-interval-seconds, ten seconds by default. Each refresh writes one LIST_CATALOG record and one LOAD_CATALOG record per catalog the connector exposes, all under the connector's user. When the connector discovers the IRC endpoint itself rather than being given its address, each refresh also writes an HTTP-level record for the discovery call. With 15 catalogs that is 17 records every ten seconds, about 147,000 a day, for each Trino cluster. A longer refresh interval on the connector reduces it in proportion.
  • Metrics scraping. Each scrape of /metrics or /prometheus/metrics writes an HTTP-level record.
  • The UI. The UI calls /configs repeatedly, and each call writes an HTTP-level record under user unknown.

A text record without custom info is about 100 bytes, so 150,000 records a day is about 15 MB before compression, well inside the chart's retention limits. Records from the IRC endpoint are larger, because they carry the request headers. The volume matters more downstream, where a SIEM bills by ingest and an investigator has to read past the background records; filtering on operation type and user before shipping is the usual remedy.

Delivery​

The server writes each record on the thread that served the request, before the response is complete. The audit log does not go through the event listener queues, so the queue settings on Events do not apply to it, and no record is lost to a full queue or to a shutdown. The cost is that writing the record is part of every request's latency, and a slow disk or a slow custom writer slows the server.

When the writer fails on a record, the server log reports Failed to write audit log, the record is lost, and the request carries on. A gap in the audit log is visible only in the server log, so watch it for that message.

Configuration​

The chart manages these properties.

Configuration ItemDescriptionDefault Value
gravitino.audit.enabledWhether to write an audit log. The Helm chart sets it to true.false
gravitino.audit.formatter.classNameFormatter that turns an event into a record.org.apache.gravitino.audit.v2.SimpleFormatterV2
gravitino.audit.writer.classNameWriter that puts the record somewhere.org.apache.gravitino.audit.FileAuditWriter

The writer's file location, rotation, and retention are not server properties. FileAuditWriter hands each record to the gravitino.audit Log4j2 logger, and the audit_file appender in log4j2.properties decides where it goes. The chart's audit.writer.file values are not read, and the server logs a warning at startup for each of them that it finds.