Skip to content
Pipelines and Pizza 🍕
Go back

Loki in Production: Labels, Per-Stream Retention, and the LogQL Alerts We Run

15 min read

Table of Contents

Open Table of Contents

Where We Left Off

Last article we deployed Loki in SimpleScalable mode — three write pods, three read pods, two backend pods, all writing to Nutanix Objects via the S3 API. That’s the deployment. This article is the operating manual.

The choices that matter for the people using Loki — what to put in a label, what to leave out, how long to keep what, how to write a query that finishes — happen here, in the values that the deployed pods read. Get them right and Loki is invisible. Get them wrong and you’re either losing data, paging the on-call for noise, or paying for storage you don’t need.


The Label Set We Actually Run

Loki indexes labels, not log contents. Every unique combination of label values creates a stream. Streams are what the ingesters hold in memory and what the index points at. Too many streams and your write path runs out of memory; too few and your queries can’t find anything without a brute-force scan.

Our cap is max_global_streams_per_user: 50000. Right now we’re at about 16,000 — comfortably under, with headroom for the Windows fleet onboarding wave that’s projected to push us to roughly 25k.

The label set we run, grouped by source:

Pod logs (from loki.source.kubernetes)

LabelCardinalitySource
namespacelow (~30)__meta_kubernetes_namespace
podmedium (changes on restart)__meta_kubernetes_pod_name
containerlow__meta_kubernetes_pod_container_name
nodelow (number of K8s nodes)__meta_kubernetes_pod_node_name
applow (number of distinct apps)__meta_kubernetes_pod_label_app_kubernetes_io_name
sourceconstantstatic label, “kubernetes”
jobconstant“loki.source.kubernetes.pod_logs”
clusterlowexternal label, “conveyor-platform”
dclowexternal label, “east” or “west”

pod is the highest-cardinality label here because pod names change every time a Deployment rolls out (my-app-5f8b9c-x7d2kmy-app-5f8b9c-q4l9p). The chunks for an old pod become inactive once the pod is gone, and the compactor reclaims them on the standard schedule. Day-to-day, it works.

Kube-audit (from the API server log file)

LabelCardinalitySource
jobconstant“kube-audit”
sourceconstant“audit”
verblow (get, list, create, update, patch, delete, watch)parsed from JSON
usermediumparsed from JSON user.username
resourcemediumparsed from JSON objectRef.resource
audit_nslowparsed from JSON objectRef.namespace
status_codevery low (HTTP codes)parsed from JSON responseStatus.code
audit_levelvery low (Request, RequestResponse, etc.)parsed from JSON level

The CRD/lease/self-subject-review filter (covered in the Alloy production post) keeps the audit volume sane. Everything that survives the filter is genuinely useful for security alerting.

Network syslog (from the alloy-network Deployment)

LabelCardinalitySource
sourceconstant“network_syslog”
device_typelow“nutanix” / “rubrik” / “dnac” / “network” / “ise”
hostnamemediumparsed from RFC 3164/5424 header
severitylowparsed from <PRI> for Cisco, regex for Nutanix

device_type is the lever we use most. Switches are network, Nutanix CVMs are nutanix, Cisco ISE is ise, Rubrik backup appliances are rubrik. Each gets its own retention rule.

Windows EventLog

LabelCardinalitySource
jobconstant“windows_eventlog”
levelvery low“Information” / “Verbose” / “Warning” / “Error” / “Critical”
channellowApplication / Security / System / etc.
hostmediumone per Windows server

level is the workhorse. It’s the only field that determines how long the Windows event lives — Error and Critical get a full year, everything else gets 90 or 180 days. Storage cost ends up dominated by the volume of Information-level events, which is why those have the shortest retention.


Per-Stream Retention: 14 Rules That Earn Their Keep

This is the section I had the most fun with in this article, because the retention table is genuinely useful and it’s not generic. Every rule has a reason.

Our retention_period is a 365-day global default. Then per-stream rules selectively trim down (or extend) for specific log sources. Loki’s compactor evaluates the per-stream rules in priority order and applies the most-specific match.

limits_config:
  retention_period: 365d
  retention_stream:
    # --- kube-audit ---
    - selector: '{job="kube-audit"}'
      priority: 1
      period: 90d

    # --- Pod logs: noisy infra ---
    - selector: '{job="loki.source.kubernetes.pod_logs",namespace="kube-system",container="calico-node"}'
      priority: 2
      period: 30d
    - selector: '{job="loki.source.kubernetes.pod_logs",namespace="observability",container="loki"}'
      priority: 2
      period: 90d
    - selector: '{job="loki.source.kubernetes.pod_logs",namespace="observability",container="nginx"}'
      priority: 2
      period: 90d
    - selector: '{job="loki.source.kubernetes.pod_logs",namespace="observability",container="grafana-sc-dashboard"}'
      priority: 2
      period: 90d
    - selector: '{job="loki.source.kubernetes.pod_logs",namespace="runners"}'
      priority: 2
      period: 90d

    # --- Pod logs catch-all ---
    - selector: '{job="loki.source.kubernetes.pod_logs"}'
      priority: 1
      period: 180d

    # --- windows_eventlog ---
    - selector: '{job="windows_eventlog",level="Information"}'
      priority: 2
      period: 90d
    - selector: '{job="windows_eventlog",level="Verbose"}'
      priority: 2
      period: 90d
    - selector: '{job="windows_eventlog",level="Warning"}'
      priority: 2
      period: 180d
    - selector: '{job="windows_eventlog",level=~"Error|Critical"}'
      priority: 2
      period: 365d

    # --- network syslog by device type ---
    - selector: '{source="network_syslog",device_type="nutanix"}'
      priority: 2
      period: 90d
    - selector: '{source="network_syslog",device_type="rubrik"}'
      priority: 2
      period: 90d
    - selector: '{source="network_syslog",device_type="dnac"}'
      priority: 2
      period: 180d
    - selector: '{source="network_syslog",device_type="network"}'   # switches, firewalls
      priority: 2
      period: 365d
    - selector: '{source="network_syslog",device_type="ise"}'        # auth logs
      priority: 2
      period: 365d

A few decisions in here are worth talking about.

calico-node at 30 days. Calico is the CNI. It’s chatty. On a busy node, the daemon logs status messages every few seconds. We don’t need a year of CNI status to debug anything; the last 30 days covers any rolling-upgrade or BGP-peering investigation we’ve ever needed. The volume difference is significant — keeping calico-node at the catch-all 180 days would consume real storage for zero query value.

Self-logs at 90 days. Loki, the Mimir nginx gateway, the Grafana dashboard sidecar — these all log a lot, but their logs only matter when something is broken with the observability stack itself. Three months covers the worst incident window.

ARC runners at 90 days. Our GitHub Actions runner pods are ephemeral by design. The runner logs from three months ago aren’t going to help us debug a CI job. Trim them.

Pod logs catch-all at 180 days, kube-audit at 90 days. This is the one I’d revisit if I were starting over. Kube-audit at 90 days is shorter than the pod logs they correspond to, which means in some forensic scenarios you can see what the pod did but not what the API server did about it. We picked 90 days for audit because the volume is high and 90 days satisfies our internal policy. If you have a different policy or different volume, this is a knob worth turning.

Switch/firewall and ISE auth logs at 365 days. These are the security and compliance logs. Regulators have opinions. The volume from syslog is modest enough that 365 days isn’t expensive.

Nutanix CVM syslog at 90 days. Operational signal, not security. Three months is plenty for any post-incident investigation.

The general principle: set per-stream retention by what the log is for, not where it comes from. Security logs get long retention. Operational logs get medium retention. Noise gets short retention. The compactor will do the rest.


LogQL Alerts We Actually Rely On

Some of the most valuable LogQL queries we run are alerting rules, not dashboards. The Loki ruler evaluates these continuously and fires alerts into Alertmanager when they trip.

Here’s the audit-log alert group, lifted directly from observability/loki/rules/audit-log-alerts.yaml:

groups:
  - name: audit-log-alerts
    rules:
      - alert: UnauthorizedAPIAccess
        expr: |
          sum by (audit_ns) (
            count_over_time(
              {job="kube-audit"} | json | responseStatus_code=~"401|403" [5m]
            )
          ) > 10
        for: 5m
        labels:
          severity: warning
        annotations:
          summary: "High rate of unauthorized API access attempts"

      - alert: SensitiveResourceAccess
        expr: |
          count_over_time(
            {job="kube-audit"}
              | json
              | objectRef_resource="secrets"
              | verb=~"create|update|delete|patch"
            [5m]
          ) > 0
        labels:
          severity: warning
        annotations:
          summary: "Sensitive resource modification detected"

      - alert: ClusterAdminBindingCreated
        expr: |
          count_over_time(
            {job="kube-audit"}
              |~ "cluster-admin"
              | json
              | objectRef_resource="clusterrolebindings"
              | verb="create"
            [15m]
          ) > 0
        labels:
          severity: critical
        annotations:
          summary: "cluster-admin ClusterRoleBinding created"

ClusterAdminBindingCreated is the one I’d point at if someone asked “what does this stack actually do for us beyond the dashboards?” If anybody — engineer, attacker, runaway controller — creates a new cluster-admin ClusterRoleBinding, the on-call gets paged within 15 minutes. We have the audit log because the API server emits it, we have Loki because we ship it, we have this alert because someone wrote three lines of LogQL. The audit pipeline pays for itself the first time this rule fires for a real reason.

The syslog alert group covers the network side:

groups:
  - name: syslog-alerts
    rules:
      - alert: NetworkDeviceCriticalSyslog
        # severity is a stream label (parsed from the Cisco %FACILITY-SEVERITY-MNEMONIC
        # pattern and mapped to names in the Alloy pipeline) — raw syslog lines aren't
        # JSON, so a `| json` stage here would error on every line
        expr: |
          count_over_time(
            {job="network_syslog", severity=~"emergency|alert|critical"} [5m]
          ) > 0
        labels:
          severity: critical
        annotations:
          summary: "Critical syslog from {{ $labels.hostname }}"

      - alert: FirewallDenySpike
        expr: |
          sum by (hostname) (
            count_over_time(
              {job="network_syslog", device_type="firewall"} |~ "(?i)(deny|drop|block|reject)" [5m]
            )
          ) > 500
        for: 5m
        labels:
          severity: warning

      - alert: InterfaceFlap
        expr: |
          sum by (hostname) (
            count_over_time(
              {job="network_syslog"}
                |~ "(?i)(line protocol.*down|link.*down|interface.*changed state to down)"
              [15m]
            )
          ) > 3
        labels:
          severity: warning

NetworkDeviceCriticalSyslog fires on Cisco syslog severity 0–2, which the Alloy pipeline maps to the emergency/alert/critical label values. Three hundred switches all over the fleet can’t realistically be watched by humans. This rule watches them.

FirewallDenySpike catches the pattern where a misconfiguration or a probe scan suddenly causes thousands of deny events. The (?i) makes the regex case-insensitive, which matters because firewall vendors don’t all agree on capitalization.

InterfaceFlap is the small-but-helpful one. Three interface state changes in 15 minutes usually means a cable is going bad or a transceiver is failing. Catching it before the user complaints arrive is a small win every time.

A couple of LogQL patterns worth noticing across all of these:

  • Always start with a label selector{job="kube-audit"}, {job="network_syslog"}. Never {} |~ "error". The label selector is what determines which chunks Loki has to scan.
  • Line filter before parser. |= "secrets" | json | objectRef_resource="secrets" — the line filter is Loki’s fastest stage and discards most lines before the parser ever runs; the parsed-field filter then adds precision (it won’t match “secrets” appearing elsewhere in the line). A bare | json over a whole stream makes Loki parse every line.
  • count_over_time over a window is the standard idiom for “how many events in the last N minutes.” Pair with sum by (label) to break down by stream.

The Ingestion-Rate Gotcha That Bit Us Early

We tripped over this one in the first month, and the confusing part wasn’t the fix — it was that the config option means the opposite of what we assumed.

ingestion_rate_strategy defaults to global, and global sounds like “the cap is checked against aggregate volume.” In accounting terms it is — but the enforcement is distributed: each distributor applies a local limiter of ingestion_rate_mb / N, so that N of them together add up to the global cap. Three write pods, 50 MB/s configured, each one enforcing about 16.6 MB/s.

Now add kube-proxy. Its connection load balancing is L4 — it picks a backend pod when a TCP connection opens and pins every subsequent packet to that pod for the life of the connection. Alloy’s loki.write holds long-lived connections. If a heavy sender’s connection lands on one distributor, that pod absorbs the whole stream — and starts returning 429s at 16.6 MB/s while its two peers sit idle. Plenty of headroom in aggregate; a hard ceiling on the one pod doing the work.

Our first instinct was to flip the strategy to local. Read the fine print before you do that: local hands every distributor the full configured cap, so three distributors will happily accept 150 MB/s combined — the number in your config quietly stops meaning anything. We kept global, because we want ingestion_rate_mb to mean “total,” and fixed the actual problem instead: we raised the cap, sizing it so that cap/N comfortably covers the worst single-sender burst. Onboarding backfill waves set that number, not the steady-state average — steady state here is single-digit MB/s, but a freshly onboarded host fleet pushing its backlog through one sticky connection is what the limiter actually sees.

loki:
  loki:
    limits_config:
      ingestion_rate_strategy: global   # keep the default — the cap means aggregate
      ingestion_rate_mb: 50             # sized so rate/N covers a backfill burst
      ingestion_burst_size_mb: 100
      per_stream_rate_limit: 10MB       # still caps any single runaway source
      per_stream_rate_limit_burst: 30MB

The per-stream limits are the other half of the answer. With the global cap raised, a single misbehaving source can’t eat the new headroom — 10 MB/s per stream is generous for anything legitimate and a brick wall for a log loop.


Structured Metadata for High-Cardinality Fields

Schema v13 added something called structured metadata, which is the right home for high-cardinality fields that you want to filter on but never group by.

Examples:

  • Trace IDs. Every request has a unique trace ID. If you make trace_id a label, you create one stream per request — your stream count explodes within an hour. But you do want to be able to find the logs for a specific trace.
  • Request IDs. Same problem.
  • User session IDs. Same problem.

The pre-v13 advice was to grep for these inside the log content with |~. That worked but it was slow because Loki had to scan every chunk for the regex.

With structured metadata, you can attach trace_id, request_id, session_id as metadata key-value pairs on each log entry without making them labels. Queries can filter on metadata directly:

{namespace="api", app="auth"} | trace_id = "abc123def456"

Loki uses the labels to find the right chunks, then scans the structured metadata stored alongside each entry in those chunks to find the matches — without creating a stream per trace, and without parsing the log line itself.

In Alloy, you set structured metadata in a loki.process stage:

loki.process "extract_metadata" {
  stage.json {
    expressions = {
      trace_id   = "trace_id",
      request_id = "request_id",
    }
  }
  stage.structured_metadata {
    values = {
      trace_id   = "",
      request_id = "",
    }
  }
  forward_to = [loki.write.local.receiver]
}

This is the right answer for any high-cardinality field that you want to query but don’t want to count as a label. If you find yourself wanting to add a label with thousands or millions of distinct values, structured metadata is what you want instead.


On Multi-Tenancy: We Don’t Use It

A note for anyone reading this expecting a deep multi-tenancy section like every other Loki tutorial.

We run single-tenant — auth_enabled: false in the deployment values from the previous article. Everyone using this Loki is on the same platform team. We don’t need to isolate data between groups; we don’t need per-tenant ingestion limits; we don’t need per-tenant retention. One tenant. Done.

If you do need multi-tenancy:

  • Flip auth_enabled: true.
  • Configure X-Scope-OrgID on every Alloy loki.write block: tenant_id = "team-name".
  • Configure per-tenant overrides in a runtime config file (mounted ConfigMap) for ingestion limits and retention.
  • Configure Grafana data sources to send the right X-Scope-OrgID header per data source (one Loki data source per tenant, or one with a template variable).

The architecture from the previous post doesn’t change; only the auth and per-request headers do. You can flip this on later if your needs change.

The Loki docs cover multi-tenancy in detail. I haven’t run it in production so I won’t pretend to have opinions about which corners it has.


Wrapping Up

Loki’s value lives in the operational choices, not the deployment. Our specific calls:

  • Labels stay low-cardinality. The set above gives us every filter we actually use, and our stream count sits comfortably under the 50k cap.
  • Per-stream retention by purpose, not by source. Security and compliance logs at 365 days, operational logs at 90–180 days, infrastructure noise at 30 days.
  • LogQL alerts on the audit and syslog streams earn their keep — cluster-admin binding creation, severity 0–2 network syslog, firewall deny spikes, interface flaps.
  • ingestion_rate_strategy: global — keep the default, but remember enforcement is rate/N per distributor and kube-proxy stickiness can pin a sender to one of them. Size the cap for the worst single-pod burst and let per-stream limits catch the runaways.
  • Structured metadata for high-cardinality fields — trace IDs, request IDs, anything you want to filter on but never group by. Use v13 schema and the stage.structured_metadata Alloy stage.
  • Single-tenant. Multi-tenancy is for cases where you’re running Loki for multiple unrelated consumers. We aren’t.

Next post (article 7) we move to Mimir. Same conceptual stack — agent ships data, scale-out components, Nutanix Objects backend — but the operational levers for metrics are very different from the ones for logs. Cardinality, scrape intervals, the ingester ring, and what happens when somebody adds a request_id label to a Prometheus counter.

Happy automating!