Skip to main content

Send the audit log to Splunk

Splunk collects the audit log through the Splunk OpenTelemetry Collector for Kubernetes, Splunk's own way of collecting container logs from a cluster. The server writes each audit record to its container's standard output as one JSON line, the collector reads that output on every node and sends it to a Splunk HTTP Event Collector (HEC) endpoint, and Splunk extracts the record's fields at search time. A search such as operationType=GRANT_PRIVILEGES user=alice works with no add-on and no field extraction rules.

Nothing is installed in the Gravitino pod, and the audit log file on the logs volume keeps being written as before. The Audit Log page describes every field a record carries.

Quick start​

  1. Write audit records to standard output. Save this override and apply it to the release.

    audit-stdout.yaml
    gravitino:
    audit:
    formatter:
    className: org.apache.gravitino.audit.JsonAuditFormatter
    additionalLog4j2Properties:
    appender.audit_console.type: Console
    appender.audit_console.name: auditConsole
    appender.audit_console.layout.type: PatternLayout
    appender.audit_console.layout.pattern: "%msg%n"
    logger.audit.name: gravitino.audit
    logger.audit.appenderRef.audit_console.ref: auditConsole
    Apply the override to the release
    helm upgrade {release} {chart} -n {namespace} --reuse-values -f audit-stdout.yaml --wait --timeout 10m
  2. Create a HEC token in Splunk. The token needs permission to write to the index the audit records go to.

  3. Install the collector. Skip this step if the cluster already runs it with log collection on and sending to that Splunk. Otherwise follow Install the Collector.

  4. Find a record. Search index={index} sourcetype="kube:container:gravitino" operationType=*. Every result is an audit record.

Write audit records to standard output​

The override in the Quick Start does two things. It switches the audit log to the JSON format, which Splunk extracts fields from automatically, and it attaches a console appender to the gravitino.audit Log4j2 logger alongside the file appender the chart already configures. Each record then goes both to gravitino_audit.log and to standard output.

Keep the logger.audit.name line. The chart applies additionalLog4j2Properties to the metrics service as well as the server, and the metrics service has no audit logger of its own; without the name, Log4j2 rejects its configuration and the metrics service does not start.

To confirm the change before involving Splunk, watch the server's output for JSON records:

Watch for JSON audit records
kubectl logs -n {namespace} deploy/{release}-gravitino -c gravitino -f | grep --line-buffered '^{"timestamp"'

Install the collector​

The collector runs on every node as a DaemonSet and reads container output from the node. Install it from Splunk's Helm chart with platform logs turned on. The HEC token goes into a Kubernetes secret so that it never appears in a values file; enter it at the prompt.

Store the HEC token and add the chart repository
read -rsp "Splunk HEC token: " HEC_TOKEN; echo
kubectl create namespace splunk-otel
kubectl create secret generic splunk-otel-hec -n splunk-otel --from-literal splunk_platform_hec_token="$HEC_TOKEN"
unset HEC_TOKEN
helm repo add splunk-otel-collector-chart https://signalfx.github.io/splunk-otel-collector-chart
helm repo update splunk-otel-collector-chart
splunk-otel-values.yaml
clusterName: {cluster_name}
secret:
create: false
name: splunk-otel-hec
validateSecret: false
splunkPlatform:
endpoint: https://{splunk_host}:8088/services/collector
index: {index}
logsEnabled: true
metricsEnabled: false
logsCollection:
containers:
containerRuntime: containerd
excludeAgentLogs: true
Install the collector
helm install splunk-otel splunk-otel-collector-chart/splunk-otel-collector -n splunk-otel \
-f splunk-otel-values.yaml --wait --timeout 10m

{splunk_host} is the Splunk instance or load balancer that serves HEC, and {index} is the index the HEC token writes to. Set containerRuntime to the runtime the cluster's nodes use: containerd on most clusters, cri-o on OpenShift. When the HEC endpoint presents a certificate the collector does not trust, such as a self-signed one on a test instance, add insecureSkipVerify: true under splunkPlatform; leave it out in production.

The collector sends the logs of every container in the cluster. Every audit record carries the sourcetype kube:container:gravitino, after the container's name, which is how searches below separate the audit trail from everything else.

Finding audit records​

Every audit record has an operationType field and nothing else from the server does, so adding operationType=* to a search on the gravitino sourcetype leaves out the server's own log lines. Keys inside custom info become fields named after their path, such as customInfo.http.status, and a field name containing dots has to be quoted in a search.

Each search below starts with index={index} sourcetype="kube:container:gravitino".

QuestionSearch
Who changed privileges or ownership?operationType IN (GRANT_PRIVILEGES, REVOKE_PRIVILEGES, OVERRIDE_PRIVILEGES, SET_OWNER)
Who changed role assignments?operationType IN (GRANT_USER_ROLES, REVOKE_USER_ROLES, GRANT_GROUP_ROLES, REVOKE_GROUP_ROLES)
What was refused for lack of a privilege?operationType=AUTHORIZATION_DENIAL
What failed?operationStatus=FAILURE
Which requests failed authentication?operationStatus=FAILURE "customInfo.http.status"=401
What did one user do?operationType=* user={user}
What happened to one object?operationType=* identifier="{metalake}.{catalog}*"
What came from provisioning?"customInfo.source"=scim

Splunk's _time for each event is the time the container wrote the line. The record's own timestamp field is the time the operation finished; the two differ by no more than the time it took to write the record.

Reducing ingest​

Every request writes a record, reads included, and Splunk licenses by the volume it ingests. On an idle system most records come from background callers described in Volume: the Trino connector reloading every catalog every ten seconds, the UI reading /configs, and metrics scrapes.

The cheapest place to drop them is the collector, before they leave the cluster, with a filter on user and operationType that matches only the scheduled reads of a known service, such as the Trino connector's user loading catalogs. Drop only what the audit trail does not need: the reads a known service makes on a schedule, not the reads people make.