Send the audit log to Splunk
Splunk collects the audit log through the Splunk OpenTelemetry Collector for Kubernetes, Splunk's
own way of collecting container logs from a cluster. The server writes each audit record to its
container's standard output as one JSON line, the collector reads that output on every node and sends
it to a Splunk HTTP Event Collector (HEC) endpoint, and Splunk extracts the record's fields at search
time. A search such as operationType=GRANT_PRIVILEGES user=alice works with no add-on and no field
extraction rules.
Nothing is installed in the Gravitino pod, and the audit log file on the logs volume keeps being written as before. The Audit Log page describes every field a record carries.
Quick start
-
Write audit records to standard output. Save this override and apply it to the release.
audit-stdout.yamlgravitino:
audit:
formatter:
className: org.apache.gravitino.audit.JsonAuditFormatter
additionalLog4j2Properties:
appender.audit_console.type: Console
appender.audit_console.name: auditConsole
appender.audit_console.layout.type: PatternLayout
appender.audit_console.layout.pattern: "%msg%n"
logger.audit.name: gravitino.audit
logger.audit.appenderRef.audit_console.ref: auditConsoleApply the override to the releasehelm upgrade {release} {chart} -n {namespace} --reuse-values -f audit-stdout.yaml --wait --timeout 10m -
Create a HEC token in Splunk. The token needs permission to write to the index the audit records go to.
-
Install the collector. Skip this step if the cluster already runs it with log collection on and sending to that Splunk. Otherwise follow Install the Collector.
-
Find a record. Search
index={index} sourcetype="kube:container:gravitino" operationType=*. Every result is an audit record.
Write audit records to standard output
The override in the Quick Start does two things. It switches the audit log to the JSON format, which
Splunk extracts fields from automatically, and it attaches a console appender to the
gravitino.audit Log4j2 logger alongside the file appender the chart already configures. Each
record then goes both to gravitino_audit.log and to standard output.
Keep the logger.audit.name line. The chart applies additionalLog4j2Properties to the metrics
service as well as the server, and the metrics service has no audit logger of its own; without the
name, Log4j2 rejects its configuration and the metrics service does not start.
To confirm the change before involving Splunk, watch the server's output for JSON records:
kubectl logs -n {namespace} deploy/{release}-gravitino -c gravitino -f | grep --line-buffered '^{"timestamp"'
Install the collector
The collector runs on every node as a DaemonSet and reads container output from the node. Install it from Splunk's Helm chart with platform logs turned on. The HEC token goes into a Kubernetes secret so that it never appears in a values file; enter it at the prompt.
read -rsp "Splunk HEC token: " HEC_TOKEN; echo
kubectl create namespace splunk-otel
kubectl create secret generic splunk-otel-hec -n splunk-otel --from-literal splunk_platform_hec_token="$HEC_TOKEN"
unset HEC_TOKEN
helm repo add splunk-otel-collector-chart https://signalfx.github.io/splunk-otel-collector-chart
helm repo update splunk-otel-collector-chart
clusterName: {cluster_name}
secret:
create: false
name: splunk-otel-hec
validateSecret: false
splunkPlatform:
endpoint: https://{splunk_host}:8088/services/collector
index: {index}
logsEnabled: true
metricsEnabled: false
logsCollection:
containers:
containerRuntime: containerd
excludeAgentLogs: true
helm install splunk-otel splunk-otel-collector-chart/splunk-otel-collector -n splunk-otel \
-f splunk-otel-values.yaml --wait --timeout 10m
{splunk_host} is the Splunk instance or load balancer that serves HEC, and {index} is the index
the HEC token writes to. Set containerRuntime to the runtime the cluster's nodes use: containerd
on most clusters, cri-o on OpenShift. When the HEC endpoint presents a certificate the collector
does not trust, such as a self-signed one on a test instance, add insecureSkipVerify: true under
splunkPlatform; leave it out in production.
The collector sends the logs of every container in the cluster. Every audit record carries the
sourcetype kube:container:gravitino, after the container's name, which is how searches below
separate the audit trail from everything else.
Finding audit records
Every audit record has an operationType field and nothing else from the server does, so adding
operationType=* to a search on the gravitino sourcetype leaves out the server's own log lines.
Keys inside custom info become fields named after their path, such as customInfo.http.status, and
a field name containing dots has to be quoted in a search.
Each search below starts with index={index} sourcetype="kube:container:gravitino".
| Question | Search |
|---|---|
| Who changed privileges or ownership? | operationType IN (GRANT_PRIVILEGES, REVOKE_PRIVILEGES, OVERRIDE_PRIVILEGES, SET_OWNER) |
| Who changed role assignments? | operationType IN (GRANT_USER_ROLES, REVOKE_USER_ROLES, GRANT_GROUP_ROLES, REVOKE_GROUP_ROLES) |
| What was refused for lack of a privilege? | operationType=AUTHORIZATION_DENIAL |
| What failed? | operationStatus=FAILURE |
| Which requests failed authentication? | operationStatus=FAILURE "customInfo.http.status"=401 |
| What did one user do? | operationType=* user={user} |
| What happened to one object? | operationType=* identifier="{metalake}.{catalog}*" |
| What came from provisioning? | "customInfo.source"=scim |
Splunk's _time for each event is the time the container wrote the line. The record's own
timestamp field is the time the operation finished; the two differ by no more than the time it took
to write the record.
Reducing ingest
Every request writes a record, reads included, and Splunk licenses by the volume it ingests. On an
idle system most records come from background callers described in Volume:
the Trino connector reloading every catalog every ten seconds, the UI reading /configs, and metrics
scrapes.
The cheapest place to drop them is the collector, before they leave the cluster, with a filter on
user and operationType that matches only the scheduled reads of a known service, such as the
Trino connector's user loading catalogs. Drop only what the audit trail does not need: the reads a
known service makes on a schedule, not the reads people make.