Send the audit log to Datadog
Datadog collects the audit log the same way it collects any container log: the server writes each
audit record to its container's standard output as one JSON line, and the Datadog Agent running in
the cluster reads that output and forwards it. Datadog parses each record into searchable
attributes, so a search such as @operationType:GRANT_PRIVILEGES @user:alice works without any
parsing rules.
Nothing is installed in the Gravitino pod, and the audit log file on the logs volume keeps being written as before. The Audit Log page describes every field a record carries.
Quick start
-
Write audit records to standard output. Save this override and apply it to the release.
audit-stdout.yamlgravitino:
audit:
formatter:
className: org.apache.gravitino.audit.JsonAuditFormatter
additionalLog4j2Properties:
appender.audit_console.type: Console
appender.audit_console.name: auditConsole
appender.audit_console.layout.type: PatternLayout
appender.audit_console.layout.pattern: "%msg%n"
logger.audit.name: gravitino.audit
logger.audit.appenderRef.audit_console.ref: auditConsoleApply the override to the releasehelm upgrade {release} {chart} -n {namespace} --reuse-values -f audit-stdout.yaml --wait --timeout 10m -
Install the Datadog Agent. Skip this step if the cluster already runs it with log collection on. Otherwise follow Install the Datadog Agent.
-
Find a record. Search Datadog's logs for
@operationType:*. Every result is an audit record. -
Map failures to a severity. Without this step Datadog shows every failed request as an emergency. See Set the Severity.
Write audit records to standard output
The override in the Quick Start does two things. It switches the audit log to the JSON format, which
Datadog parses on arrival, and it attaches a console appender to the gravitino.audit Log4j2 logger
alongside the file appender the chart already configures. Each record then goes both to
gravitino_audit.log and to standard output.
Keep the logger.audit.name line. The chart applies additionalLog4j2Properties to the metrics
service as well as the server, and the metrics service has no audit logger of its own; without the
name, Log4j2 rejects its configuration and the metrics service does not start.
Confirm the change here before involving Datadog.
kubectl logs -n {namespace} deploy/{release}-gravitino -c gravitino -f | grep --line-buffered '^{"timestamp"'
Install the Datadog agent
The Agent runs on every node as a DaemonSet and reads container output from the node. Install it from Datadog's Helm chart with log collection turned on. The API key goes into a Kubernetes secret so that it never appears in a values file; enter it at the prompt.
read -rsp "Datadog API key: " DD_API_KEY; echo
kubectl create namespace datadog
kubectl create secret generic datadog-secret -n datadog --from-literal api-key="$DD_API_KEY"
unset DD_API_KEY
helm repo add datadog https://helm.datadoghq.com && helm repo update datadog
datadog:
apiKeyExistingSecret: datadog-secret
site: {datadog_site}
clusterName: {cluster_name}
logs:
enabled: true
containerCollectAll: true
containerExcludeLogs: "kube_namespace:.*"
containerIncludeLogs: "kube_namespace:{namespace}"
helm install datadog-agent datadog/datadog -n datadog -f datadog-values.yaml --wait --timeout 10m
{datadog_site} is the Datadog site the account is on, such as datadoghq.com or us5.datadoghq.com.
The two container*Logs lines limit collection to the namespace Gravitino runs in; drop them to
collect from the whole cluster.
On a k3s cluster, also set the containerd socket and accept the kubelet's self-signed certificate, or the Agent sees no containers:
datadog:
criSocketPath: /run/k3s/containerd/containerd.sock
kubelet:
tlsVerify: false
Confirm that the Agent is reading the server's output.
kubectl exec -n datadog ds/datadog-agent -c agent -- agent status | grep -A2 '{release}-gravitino.*/gravitino$'
Finding audit records
Every audit record has an operationType attribute and nothing else from the server does, so
@operationType:* selects audit records and leaves out the server's own log lines and the other
services in the namespace. Datadog uses the record's timestamp as the event time, and keys such as
http.method in custom info become nested attributes such as @customInfo.http.method.
| Question | Search |
|---|---|
| Who changed privileges or ownership? | @operationType:(GRANT_PRIVILEGES OR REVOKE_PRIVILEGES OR OVERRIDE_PRIVILEGES OR SET_OWNER) |
| Who changed role assignments? | @operationType:(GRANT_USER_ROLES OR REVOKE_USER_ROLES OR GRANT_GROUP_ROLES OR REVOKE_GROUP_ROLES) |
| What was refused for lack of a privilege? | @operationType:AUTHORIZATION_DENIAL |
| What failed? | @operationStatus:FAILURE |
| Which requests failed authentication? | @operationStatus:FAILURE @customInfo.http.status:401 |
| What did one user do? | @operationType:* @user:{user} |
| What happened to one object? | @operationType:* @identifier:{metalake}.{catalog}* |
| What came from provisioning? | @customInfo.source:scim |
Set the severity
Datadog reads a field named status as a record's severity, and every audit record has one. It shows
SUCCESS records as OK, and because it reads any value starting with F as fatal, it shows
FAILURE records as Emergency, its highest severity. A single mistyped password then appears as an
emergency, and an alert on high-severity logs fires on it.
A log pipeline that sets the severity from operationStatus fixes this. Create a pipeline whose
filter is @operationType:*, and give it two processors in this order:
-
A category processor. Set the target attribute to
audit.severity, with two categories:@operationStatus:FAILUREnamederror, and@operationStatus:SUCCESSnamedinfo. -
A status remapper. Set it to read
audit.severity.
New records then arrive with severity Error or Info. Records already indexed keep the severity they had.
Reducing ingest
Every request writes a record, reads included, and Datadog bills by the volume it ingests and
indexes. On an idle system most records come from background callers described in
Volume: the Trino connector reloading every catalog every ten seconds, the
UI reading /configs, and metrics scrapes.
Datadog can drop records before indexing them with an exclusion filter on the index. A filter that matches the Trino connector's refresh, for example, removes most of an idle system's volume:
@operationType:(LIST_CATALOG OR LOAD_CATALOG) @user:{trino_connector_user}
Excluded records are still ingested, and they can still be archived, but they are not searchable. Exclude only what the audit trail does not need: the reads a known service makes on a schedule, not the reads people make.
Other logs from the namespace
With collection scoped to the namespace, Datadog also receives the server's own log lines and the
output of every other pod there. Some of those arrive marked Error without being errors: PostgreSQL,
for example, writes its routine checkpoint messages to standard error, and Datadog marks everything
on standard error as an error. Filter on @operationType:* when the question is about the audit
trail, and on service:{release}-gravitino when it is about the server.