Skip to main content

Send the audit log to Datadog

Datadog collects the audit log the same way it collects any container log: the server writes each audit record to its container's standard output as one JSON line, and the Datadog Agent running in the cluster reads that output and forwards it. Datadog parses each record into searchable attributes, so a search such as @operationType:GRANT_PRIVILEGES @user:alice works without any parsing rules.

Nothing is installed in the Gravitino pod, and the audit log file on the logs volume keeps being written as before. The Audit Log page describes every field a record carries.

Quick start​

  1. Write audit records to standard output. Save this override and apply it to the release.

    audit-stdout.yaml
    gravitino:
    audit:
    formatter:
    className: org.apache.gravitino.audit.JsonAuditFormatter
    additionalLog4j2Properties:
    appender.audit_console.type: Console
    appender.audit_console.name: auditConsole
    appender.audit_console.layout.type: PatternLayout
    appender.audit_console.layout.pattern: "%msg%n"
    logger.audit.name: gravitino.audit
    logger.audit.appenderRef.audit_console.ref: auditConsole
    Apply the override to the release
    helm upgrade {release} {chart} -n {namespace} --reuse-values -f audit-stdout.yaml --wait --timeout 10m
  2. Install the Datadog Agent. Skip this step if the cluster already runs it with log collection on. Otherwise follow Install the Datadog Agent.

  3. Find a record. Search Datadog's logs for @operationType:*. Every result is an audit record.

  4. Map failures to a severity. Without this step Datadog shows every failed request as an emergency. See Set the Severity.

Write audit records to standard output​

The override in the Quick Start does two things. It switches the audit log to the JSON format, which Datadog parses on arrival, and it attaches a console appender to the gravitino.audit Log4j2 logger alongside the file appender the chart already configures. Each record then goes both to gravitino_audit.log and to standard output.

Keep the logger.audit.name line. The chart applies additionalLog4j2Properties to the metrics service as well as the server, and the metrics service has no audit logger of its own; without the name, Log4j2 rejects its configuration and the metrics service does not start.

Confirm the change here before involving Datadog.

Watch the server output for audit records
kubectl logs -n {namespace} deploy/{release}-gravitino -c gravitino -f | grep --line-buffered '^{"timestamp"'

Install the Datadog agent​

The Agent runs on every node as a DaemonSet and reads container output from the node. Install it from Datadog's Helm chart with log collection turned on. The API key goes into a Kubernetes secret so that it never appears in a values file; enter it at the prompt.

Store the API key and add the Datadog repository
read -rsp "Datadog API key: " DD_API_KEY; echo
kubectl create namespace datadog
kubectl create secret generic datadog-secret -n datadog --from-literal api-key="$DD_API_KEY"
unset DD_API_KEY
helm repo add datadog https://helm.datadoghq.com && helm repo update datadog
datadog-values.yaml
datadog:
apiKeyExistingSecret: datadog-secret
site: {datadog_site}
clusterName: {cluster_name}
logs:
enabled: true
containerCollectAll: true
containerExcludeLogs: "kube_namespace:.*"
containerIncludeLogs: "kube_namespace:{namespace}"
Install the Datadog Agent
helm install datadog-agent datadog/datadog -n datadog -f datadog-values.yaml --wait --timeout 10m

{datadog_site} is the Datadog site the account is on, such as datadoghq.com or us5.datadoghq.com. The two container*Logs lines limit collection to the namespace Gravitino runs in; drop them to collect from the whole cluster.

On a k3s cluster, also set the containerd socket and accept the kubelet's self-signed certificate, or the Agent sees no containers:

datadog-values.yaml additions for k3s
datadog:
criSocketPath: /run/k3s/containerd/containerd.sock
kubelet:
tlsVerify: false

Confirm that the Agent is reading the server's output.

Check the Agent status for the gravitino container
kubectl exec -n datadog ds/datadog-agent -c agent -- agent status | grep -A2 '{release}-gravitino.*/gravitino$'

Finding audit records​

Every audit record has an operationType attribute and nothing else from the server does, so @operationType:* selects audit records and leaves out the server's own log lines and the other services in the namespace. Datadog uses the record's timestamp as the event time, and keys such as http.method in custom info become nested attributes such as @customInfo.http.method.

QuestionSearch
Who changed privileges or ownership?@operationType:(GRANT_PRIVILEGES OR REVOKE_PRIVILEGES OR OVERRIDE_PRIVILEGES OR SET_OWNER)
Who changed role assignments?@operationType:(GRANT_USER_ROLES OR REVOKE_USER_ROLES OR GRANT_GROUP_ROLES OR REVOKE_GROUP_ROLES)
What was refused for lack of a privilege?@operationType:AUTHORIZATION_DENIAL
What failed?@operationStatus:FAILURE
Which requests failed authentication?@operationStatus:FAILURE @customInfo.http.status:401
What did one user do?@operationType:* @user:{user}
What happened to one object?@operationType:* @identifier:{metalake}.{catalog}*
What came from provisioning?@customInfo.source:scim

Set the severity​

Datadog reads a field named status as a record's severity, and every audit record has one. It shows SUCCESS records as OK, and because it reads any value starting with F as fatal, it shows FAILURE records as Emergency, its highest severity. A single mistyped password then appears as an emergency, and an alert on high-severity logs fires on it.

A log pipeline that sets the severity from operationStatus fixes this. Create a pipeline whose filter is @operationType:*, and give it two processors in this order:

  1. A category processor. Set the target attribute to audit.severity, with two categories: @operationStatus:FAILURE named error, and @operationStatus:SUCCESS named info.

  2. A status remapper. Set it to read audit.severity.

New records then arrive with severity Error or Info. Records already indexed keep the severity they had.

Reducing ingest​

Every request writes a record, reads included, and Datadog bills by the volume it ingests and indexes. On an idle system most records come from background callers described in Volume: the Trino connector reloading every catalog every ten seconds, the UI reading /configs, and metrics scrapes.

Datadog can drop records before indexing them with an exclusion filter on the index. A filter that matches the Trino connector's refresh, for example, removes most of an idle system's volume:

Exclusion filter for the Trino connector refresh
@operationType:(LIST_CATALOG OR LOAD_CATALOG) @user:{trino_connector_user}

Excluded records are still ingested, and they can still be archived, but they are not searchable. Exclude only what the audit trail does not need: the reads a known service makes on a schedule, not the reads people make.

Other logs from the namespace​

With collection scoped to the namespace, Datadog also receives the server's own log lines and the output of every other pod there. Some of those arrive marked Error without being errors: PostgreSQL, for example, writes its routine checkpoint messages to standard error, and Datadog marks everything on standard error as an error. Filter on @operationType:* when the question is about the audit trail, and on service:{release}-gravitino when it is about the server.