Skip to main content

Helm Deployment and Configuration

Deprecated SPIRE runtime

SPIFFE/SPIRE runtime support is deprecated. Use mTLS with ServiceRadar's deployment-managed CA. Explicit SPIRE configuration remains compatible during this deprecation phase. See Migrating off SPIRE. Existing spiffe:// certificate URI identities remain supported.

This guide shows how to deploy ServiceRadar via the bundled Helm chart. For sweep behavior, tuning, and concepts, see Network Sweeps and SYN Scanner Tuning and Conntrack Mitigation.

Chart version

<chart-version> below is a placeholder. Look up the current release before deploying:

helm show chart oci://registry.carverauto.dev/serviceradar/charts/serviceradar | grep '^version'

This page deliberately does not name a specific version: a hardcoded example goes stale silently, and readers reasonably copy it as fact.

Install/upgrade

  • Namespace: create once: kubectl create ns serviceradar (or change namespace in chart values).
  • Deploy from the official OCI chart (recommended):
    • helm upgrade --install serviceradar oci://registry.carverauto.dev/serviceradar/charts/serviceradar --version <chart-version> -n serviceradar --create-namespace -f my-values.yaml
  • Test unreleased chart changes from a source checkout (chart development only; operators use the OCI chart above):
    • helm upgrade --install serviceradar ./helm/serviceradar -n serviceradar -f my-values.yaml
  • Quick overrides without a file: add --set flags (examples below).
  • MCP (/mcp) is off by default. Enable with --set webNg.mcpEnabled="true" (the chart must template SERVICERADAR_MCP_* from that key; extraEnv cannot shadow it). Demo overlay values-demo.yaml already turns it on. Client setup (Codex, Claude Code, Grok) is in MCP Integration.

OCI chart quick start

  • Inspect chart metadata and defaults:
    • helm show chart oci://registry.carverauto.dev/serviceradar/charts/serviceradar --version <chart-version>
    • helm show values oci://registry.carverauto.dev/serviceradar/charts/serviceradar --version <chart-version> > default-values.yaml (reference only; put just the keys you change in my-values.yaml)
  • Image tags follow the chart by default:
    • If you leave global.imageTag empty (the default), every first-party ServiceRadar image except serviceradar-cnpg (pinned by digest) uses the chart's appVersion, pulled as registry.carverauto.dev/serviceradar/serviceradar-<component>:v<chart-version>. The chart and the application it deploys are released together, so --version <chart-version> alone selects matching images and needs no further configuration.
  • Treat global.imageTag as an override, not a release selector:
    • If you set it, it must be v<chart-version> for the same --version you install. Any other tag drifts from the core.migrations.expectedVersion and templates that chart version ships. To change releases, change --version.
  • Pin images explicitly (immutable rollouts):
    • Pin per-service digests of those v<chart-version> images with image.digests.*.

HA profile overlay

  • values.yaml stays conservative by default. Most stateful or queue-backed services start at 1 replica unless you opt into a larger topology.
  • values-ha.yaml ships inside the published chart as a purpose-named HA overlay. Extract it from the chart version you install with helm pull oci://registry.carverauto.dev/serviceradar/charts/serviceradar --version <chart-version> --untar --untardir serviceradar-<chart-version>, then apply it with -f serviceradar-<chart-version>/serviceradar/values-ha.yaml, before your own -f my-values.yaml, as the starting point for a multi-replica deployment.
  • The HA overlay runs these at 3 replicas:
    • core
    • webNg
    • agentGateway
    • datasvc
    • logCollector
    • logCollector.tcpCollector
    • trapd
  • flowCollector stays at replicaCount: 1 (IPFIX/NetFlow template state is process-local) with Recreate and a 1 GiB RWO data PVC so rehome/ownership/readiness markers survive pod replacement. Stream HA is config.stream_replicas (JetStream), not pod count.
  • bmpCollector is not scaled by the HA overlay unless another values file sets it.
  • The profile disables PVC-backed local state for the multi-replica services above where shared NATS/JetStream state is the real source of truth (flow-collector is the deliberate exception).

Optional public endpoint inventory

  • k8sInventory.enabled (default false) deploys a cluster-plane collector that maps LoadBalancer / Gateway API public addresses to Services, routes, and backend pods. Helm creates the ServiceAccount, read-only ClusterRole, and Deployment together—do not create the SA by hand for normal installs.
  • Full IR workflow, Argo CD vs manual install notes, and RBAC details: Kubernetes Public Endpoint Inventory.
  • Demo values (values-demo.yaml) enable it with clusterId: demo once the serviceradar-k8s-inventory image is available for that release tag.

JetStream sizing values

  • nats.jetstream.profile selects a sizing profile: small (default, max_file_store 36G for the default 36Gi PVC), medium (100G for a 100Gi PVC) or large (500G for a 500Gi PVC), plus tenant-2g (2G) for hosted tenant plans. With nats.replicas of 1 or 2 the shipped profiles use their single-server size table (the Docker Compose sizes). A profile sets max_file_store and the default size of every stream, KV bucket and object store ServiceRadar creates. Every size key below is unset by default so the profile supplies it; an explicit value wins, and so does nats.jetstream.maxFileStore. The chart renders max_file_store as an exact byte count (NATS reads G as 10^9 and Gi as 2^30).
  • The chart checks the budget at render time: helm template and helm upgrade fail, listing every reservation, when the most loaded NATS server would reserve more than 85% of max_file_store. nats.jetstream.allowOvercommit: true skips that check. flows and ARANCINI_CAUSAL are always counted at the collector size, even when the collector is disabled.
  • The chart also fails when max_file_store exceeds 94% of nats.persistence.size, and allowOvercommit does not skip it. A profile never resizes the NATS PVC; moving a live install to a larger profile follows the volume-expansion runbook at docs/nats-jetstream-profile-runbook.md in the repository.
  • EventWriter streams: core.eventWriter.streams.<stream>.maxBytes for metrics, k8s_inventory, analytics_predictions, mtr_results, scan_results, trivy_reports and notifications (the NOTIFICATIONS stream), and the EventWriter fallbacks for flows and ARANCINI_CAUSAL (with .replicas). Object stores: webNg.pluginStorage.jetstreamMaxBucketBytes, webNg.fieldSurveyArtifactStore.jetstreamMaxBucketBytes and core.threatIntelRawPayloadStore.jetstreamMaxBucketBytes.
  • The shared events stream is created and reconciled by multiple services. The important knobs are:
    • logCollector.streamReplicas
    • logCollector.streamMaxBytes (profile default: 2 GiB in small)
    • trapd.streamReplicas
  • Dedicated flows stream (owned by flow-collector; isolated from logs/OTEL on events):
    • flowCollector.config.stream_name (default flows)
    • flowCollector.config.stream_replicas
    • flowCollector.config.stream_max_bytes (profile default: 8 GiB in small)
    • flowCollector.config.stream_max_age_secs (default 6h)
    • EventWriter consumers use concrete flows.raw.<name> leaves only; do not put ownership wildcards such as flows.raw.> / flows.> / *.> in collector subjects (collector validation rejects them). Host-slice subjects are not EventWriter flow consumers; attribution joining is out of scope for this chart.
  • Datasvc owns the KV/object streams and now reconciles both replica count and reserved capacity:
    • datasvc.jetstreamReplicas
    • datasvc.bucketMaxBytes (profile default: 1 GiB in small)
    • datasvc.objectMaxBytes
    • datasvc.objectStoreBytes (profile default: 4 GiB in small)
  • The example HA profile intentionally shrinks those reserved capacities compared to the generic chart defaults so events can run at 3 replicas without exhausting the JetStream account's file-store budget.
  • Agent release object cleanup is enabled by default through objectStoreRetention; it keeps the most recently imported release plus any releases still referenced by active rollout state.
  • bmpCollector is scaled to 3 pods in the example profile, but its dedicated causal-overlay stream still uses bmpCollector.config.streamReplicas=1. That is an explicit sizing choice, not a pod-level HA limitation.

Key values: workload identity (spire, deprecated)

  • spire.enabled defaults to false. The chart still issues runtime mTLS certificates without SPIRE (see TLS Security).
  • Existing opt-in installs can retain spire.enabled=true during deprecation, and set spire.trustDomain to your environment's trust domain.

Key values: topology graph (dgraph, graph)

  • dgraph.enabled defaults to true. The chart installs Dgraph (Zero + Alpha) as a subchart, generates its ACL credential, mints its TLS Secrets, and applies the topology schema through a post-install Job.
  • Those Secrets are ordinary kubernetes.io/tls objects (dgraph-ca and the Alpha TLS Secret the subchart mounts). The chart creates them on install and reuses the same bytes on upgrade. A tenant cluster does not need cert-manager for Dgraph. Public HTTPS is issued separately.
  • To reuse a Dgraph cluster you already run, set dgraph.enabled=false and dgraph.external.host. Its ACL credential comes from a Secret that already exists in the namespace (dgraph.external.credentialsSecret), never from a chart value.
  • graph.backend and graph.read select where topology is written and read. Reads stay on AGE until you cut over deliberately.
  • The store layout, the migrator Job, credential handling, and the cutover/rollback procedure are in Network Topology.

Key values: sweep

The chart exposes the full sweep configuration tree (sweep.networks, sweep.ports, sweep.modes, sweep.tcp.*, sweep.icmp.*, and related tuning knobs). Rather than duplicate that reference here, see:

Inspect the current defaults for your chart version with helm show values oci://registry.carverauto.dev/serviceradar/charts/serviceradar --version <chart-version>.

Public web and edge-agent endpoints​

ServiceRadar publishes two independent paths during agent onboarding:

PathHelm valuesUsed for
Public web/APIwebNg.host, webNg.publicUrlBrowser access, Phoenix URL generation, the generated --core-url, and the API origin signed into onboarding tokens
Agent gatewaywebNg.gatewayAddress, agentGateway.publicHostnameThe gateway_addr and default TLS server name placed in the downloaded bundle, plus the public artifact endpoint

Canonical public web origin​

  • Set webNg.host to the public web DNS name and webNg.publicUrl to its bare HTTPS origin. Do this even when gatewayApi.host or ingress.host has the same value; keeping the application origin explicit prevents an exposure change from altering newly issued onboarding tokens.
  • Set webNg.publicUrl to the bare, externally reachable HTTPS origin, with no path, query, fragment, or credentials (for example, https://serviceradar.example.com). A trailing root slash is canonicalized away. Only the standard HTTPS port 443 is supported.
  • This is the canonical origin embedded in edge onboarding tokens and generated enrollment commands. It also drives Phoenix external URL generation. Never use an in-cluster Service name here.
  • When webNg.publicUrl is empty, the chart falls back through webNg.host, ingress.host, and gatewayApi.host. Set webNg.publicUrl explicitly in production so changing the exposure implementation does not change issued tokens.
webNg:
host: serviceradar.example.com
publicUrl: https://serviceradar.example.com

Edge gateway address​

  • Set webNg.gatewayAddress to the externally reachable agent-gateway host:port, normally TCP 50052. It is not an HTTP URL and it need not use the web hostname.
  • Set agentGateway.publicHostname to the same DNS name, without a port. The chart uses this hostname for the public artifact URL on agentGateway.service.artifactPort (default 50053). The bundle's TLS server name defaults to the host in webNg.gatewayAddress; ensure the certificate presented by the gateway is issued or reissued with that name.
  • If webNg.gatewayAddress is unset, the chart derives <web-host>:50052 from the public web host. That fallback is correct only when the same L4 address actually exposes the agent-gateway port. It does not make an HTTP-only Gateway listen on 50052.

Choose one exposure pattern:

  1. Dedicated agent-gateway LoadBalancer. Give the Service a dedicated DNS name, expose 50052 and 50053, and point webNg.gatewayAddress at it. This is the pattern the repository's demo overlay (helm/serviceradar/values-demo.yaml) uses.

    webNg:
    host: serviceradar.example.com
    publicUrl: https://serviceradar.example.com
    gatewayAddress: agent-gateway.example.com:50052

    agentGateway:
    publicHostname: agent-gateway.example.com
    service:
    type: LoadBalancer
    annotations:
    external-dns.alpha.kubernetes.io/hostname: agent-gateway.example.com.
  2. Shared Gateway API data plane. Route the agent-gateway ports through the same Envoy data-plane Service as the web endpoint. The parent Gateway must have TCP listeners for 50052 and 50053; an HTTPRoute on 443 cannot carry this traffic. In managed mode the chart creates the listeners. In attach mode, enable gatewayApi.agentGateway and supply parentRefs for existing listener section names.

    webNg:
    host: serviceradar.example.com
    publicUrl: https://serviceradar.example.com
    gatewayAddress: agent-gateway.example.com:50052

    agentGateway:
    publicHostname: agent-gateway.example.com
    service:
    type: ClusterIP

    # The existing Gateway must already define TCP listeners named
    # agent-grpc (50052) and agent-artifacts (50053).
    gatewayApi:
    enabled: true
    mode: attach
    host: serviceradar.example.com
    # Web HTTPS route. The existing Gateway listener must allow routes from
    # the ServiceRadar release namespace.
    parentRefs:
    - group: gateway.networking.k8s.io
    kind: Gateway
    name: serviceradar-shared-gateway
    namespace: serviceradar-system
    sectionName: https-web
    agentGateway:
    enabled: true
    grpc:
    parentRefs:
    - group: gateway.networking.k8s.io
    kind: Gateway
    name: serviceradar-shared-gateway
    namespace: serviceradar-system
    sectionName: agent-grpc
    artifacts:
    enabled: true
    parentRefs:
    - group: gateway.networking.k8s.io
    kind: Gateway
    name: serviceradar-shared-gateway
    namespace: serviceradar-system
    sectionName: agent-artifacts

    Every referenced listener must allow routes from the ServiceRadar release namespace. A cross-namespace parentRef to a Gateway is authorized by that listener's allowedRoutes; a ReferenceGrant is needed only if a route also refers to a backend in a different namespace.

For a shared public IP, both DNS names can resolve to that IP. Using a dedicated gateway DNS name remains useful because the bundle and certificate identity do not then depend on the web hostname. Confirm the chosen address accepts both ports before issuing onboarding packages.

Key values: in-cluster agent storage

  • agent.checkersStorage: PVC-backed checker config at /var/lib/serviceradar/checkers.
  • agent.cacheStorage: PVC-backed agent cache at /var/lib/serviceradar/cache.
  • agent.runtimeStorage: PVC-backed managed release runtime at /var/lib/serviceradar/agent.

Keep these enabled in production. The agent writes mutable config caches and managed release payloads under /var/lib/serviceradar; without PVC-backed storage those writes count against pod ephemeral storage and can trigger evictions under disk pressure.

Example:

agent:
checkersStorage:
enabled: true
storageClassName: fast-rwo
cacheStorage:
enabled: true
storageClassName: fast-rwo
size: 1Gi
runtimeStorage:
enabled: true
storageClassName: fast-rwo
size: 5Gi

ServiceRadar stores and distributes network credentials (for example SNMP communities and API tokens) as part of discovery, polling, and inventory sync configuration. Even though the UI does not display secrets back to users, a compromised privileged account could still try to abuse configuration to trigger unexpected outbound traffic (for example by adding attacker-controlled targets and new discovery/polling profiles).

Enable an egress NetworkPolicy to reduce blast radius and make exfiltration harder. The bundled Helm chart can install a restrictive egress policy that:

  • allows DNS (optional)
  • allows in-namespace communication (optional)
  • allows Kubernetes API server access (optional; auto-detects API endpoints via lookup)
  • allows explicit destination CIDRs you provide (recommended)

Important notes:

  • NetworkPolicy enforcement depends on your CNI (Calico, Cilium, etc). If your cluster does not enforce NetworkPolicy, enabling these values will not change runtime behavior.
  • This policy applies to pods selected by networkPolicy.podSelector (or all pods in the namespace when podSelectorMatchAll: true).
  • Edge hosts running serviceradar-agent outside Kubernetes need their own egress controls (host firewall/VPC/NACL). This policy only governs Kubernetes workloads.
  • External telemetry collectors have dedicated pod-scoped ingress policies. Use them for syslog, NetFlow, sFlow, SNMP traps, and BMP so opening a collector port does not also expose unrelated workloads. See Kubernetes External Ingestion.
  • Plugins and integrations that call public services need explicit egress. For AlienVault OTX, allow otx.alienvault.com with an FQDN-aware policy. Its CDN addresses rotate, so a static allowedCIDRs entry requires ongoing DNS resolution and CIDR maintenance.
  • Control-plane notification webhooks (Discord, Slack, Teams, generic HTTPS) egress from the core pods. webNg.obanQueues.notifications defaults to 0, so web-ng does not run dispatch and does not publish the notifications.* firehose. Kubernetes NetworkPolicy cannot match FQDNs, so add the current CDN CIDR for each destination to networkPolicy.egress.allowedCIDRs. Discord incoming webhooks currently land on Cloudflare 162.159.128.0/18 (resolved 2026-08-13); if a Discord test send times out with timeout contacting discord.com, re-resolve discord.com:443 and update that CIDR. The demo overlay (values-demo.yaml) already includes this range.

Example:

networkPolicy:
enabled: true
podSelectorMatchAll: true
ingress:
allowSameNamespace: true
allowedCIDRs:
- "10.0.0.0/8"
- "192.168.0.0/16"
flowCollectorExternal:
enabled: true
allowedCIDRs:
- "10.0.0.0/8"
logCollectorExternal:
enabled: true
allowedCIDRs:
- "10.0.0.0/8"
trapdExternal:
enabled: true
allowedCIDRs:
- "10.0.0.0/8"
bmpCollectorExternal:
enabled: true
allowedCIDRs:
- "10.0.0.0/8"
egress:
allowDNS: true
allowKubeAPIServer: true
allowDefaultNamespace: true
allowSameNamespace: true
allowedCIDRs:
- "10.0.0.0/8"
- "192.168.0.0/16"

Gateway API Syslog​

When gatewayApi.enabled=true, the chart can attach syslog to a shared Gateway API UDP listener. This is the preferred way to receive syslog in clusters that already have a shared Envoy Gateway because it avoids allocating another collector address.

Example:

gatewayApi:
enabled: true
mode: attach
syslog:
enabled: true
parentRefs:
- group: gateway.networking.k8s.io
kind: Gateway
name: serviceradar-shared-gateway
namespace: serviceradar-system
sectionName: syslog-udp

The parent Gateway must expose a UDP listener named by sectionName, and the ServiceRadar namespace must be allowed by that listener. If NetworkPolicy is enabled, allow ingress from the Gateway data-plane namespace because traffic reaches the log collector from Envoy pods.

Optional (Calico): log and deny unmatched egress

If you run Calico, you can enable a Calico NetworkPolicy that logs denied egress before denying it:

networkPolicy:
calicoLogDenied:
enabled: true
selector: "app.kubernetes.io/part-of == 'serviceradar'"
order: 1000

CNPG WAL and Checkpoint Tuning​

PostgreSQL forces a checkpoint every max_wal_size / (2 + checkpoint_completion_target). At PostgreSQL's stock max_wal_size=1GB that is 353 MB of WAL, which on a busy ServiceRadar deployment meant a checkpoint roughly every 14 seconds -- and under heavy ingest every 10 seconds, with every database backend blocked on LWLock:WALWrite. The chart therefore sizes the WAL budget instead of inheriting the defaults.

pg_wal shares the CNPG data volume (the chart declares no separate walStorage), so three parameters derive from cnpg.storageSize and are bounded relative to it:

cnpg.storageSizemax_wal_sizemin_wal_sizemax_slot_wal_keep_size
10Gi1GB256MB5GB
20Gi1GB256MB10GB
30Gi2GB256MB10GB
50Gi3GB384MB15GB
100Gi (default)7GB896MB30GB
200Gi8GB1024MB60GB
1Ti8GB1024MB307GB

max_wal_size is 7% of the volume clamped to [1GB, 8GB]; min_wal_size is an eighth of that clamped to [256MB, 1024MB]; max_slot_wal_keep_size stays at its established ~30% policy, floored at 10GB but never more than half the volume.

The 1GB floor on max_wal_size is deliberate: an install of roughly 28Gi or less keeps PostgreSQL's own default and spends no extra WAL disk. Such deployments get no checkpoint relief -- you cannot spend disk you do not have -- and should raise cnpg.maxWalSize explicitly if their volume can afford it.

Because these derive from cnpg.storageSize, that value must track the real provisioned volume. A storageSize that has drifted below the actual disk under-sizes three parameters rather than one.

Overrides​

Each has an escape hatch, and setting the PostgreSQL parameter directly always wins:

cnpg:
maxWalSize: "" # e.g. "16GB"
minWalSize: "" # e.g. "2048MB"
maxSlotWalKeepSize: "" # e.g. "250GB"
postgresqlParameters:
max_wal_size: "16GB" # takes precedence over cnpg.maxWalSize

Use PostgreSQL units (GB, MB), not Kubernetes resource units (Gi, Mi).

The same keys exist under spire.postgres.*, which hosts the application database when spire.postgres.enabled is true.

Cost and rollout​

Worst-case pg_wal use is roughly 2 x max_wal_size plus the slot cap. On the 100Gi default that moves from about 32GB to about 44GB of the volume, so check free space on the data volume before upgrading an install already near its high-water mark.

All of these parameters, plus checkpoint_timeout, checkpoint_completion_target, wal_compression and log_parameter_max_length, apply by SIGHUP reload. Changing them does not restart any CNPG pod and does not trigger a switchover.

Note that this budget does not bound WAL end-to-end: when cnpg.backup.enabled is true, WAL awaiting archival is held by neither max_wal_size nor max_slot_wal_keep_size, so a failing object-store archive can still fill the volume.

After a rollout, confirm convergence with pg_stat_checkpointer: num_timed should rise and num_requested should fall toward zero. num_requested dominating means checkpoints are still being forced by WAL volume rather than by the timer.

CNPG PgBouncer Pooler​

Kubernetes installs can enable a CNPG-managed PgBouncer pooler through the Helm chart. This deploys a postgresql.cnpg.io/v1 Pooler resource and routes PgBouncer-safe runtime database clients through the generated pooler service. Schema migrations, bootstrap jobs, and other DDL/admin paths continue to use the direct CNPG RW service.

Example:

cnpg:
pooler:
enabled: true
instances: 3
poolMode: transaction
ha:
podAntiAffinity:
enabled: true
type: preferred
monitoring:
podMonitor:
enabled: true
route:
core: true
webNg: true
parameters:
ignore_startup_parameters: "search_path"
max_client_conn: "2000"
default_pool_size: "40"
reserve_pool_size: "10"

Operational notes:

  • Transaction pooling requires clients to avoid named prepared statements. The chart sets DATABASE_PREPARE=unnamed for core and web-ng when those workloads are routed through the pooler.
  • CNPG PgBouncer presents the PostgreSQL server certificate. When verify-full is enabled, the chart connects to the pooler service but sets CNPG_TLS_SERVER_NAME to the direct CNPG RW service name for routed Elixir workloads so hostname verification remains strict.
  • Ecto sends search_path as a PostgreSQL startup parameter. The pooler defaults include ignore_startup_parameters=search_path; keep the database role search path configured server-side for routed workloads.
  • PgBouncer is deployed as an HA access layer by default with three Pooler pods and preferred pod anti-affinity. Set cnpg.pooler.ha.podAntiAffinity.type=required only when the cluster has enough nodes to satisfy strict placement.
  • Enable cnpg.pooler.monitoring.podMonitor.enabled=true when Prometheus Operator is installed. The scraper targets the CNPG PgBouncer exporter on port metrics and exposes the cnpg_pgbouncer_ metric family.
  • Keep migrations and bootstrap direct to cnpg-rw; PgBouncer transaction pooling is not appropriate for DDL, extension setup, or migration locks.

Optional StarRocks Analytics​

analytics.starrocks.* enables an opt-in telemetry warehouse for flows, scalar metrics, logs, event history, and further append-only datasets, including BMP routing events. It is off by default, and NetFlow collection does not depend on it (flowCollector.enabled is independent). BMP routing events are warehouse-only while StarRocks is enabled; the table and retention contract is in k8s/starrocks/README.md. Metric, log and event reads stay on CNPG until the dataset is named in analytics.starrocks.cutoverDatasets. Flow routing, required delivery and the historical-attribution limitation are documented in NetFlow: Flow cutover and delivery.

web-ng and core read these settings once at boot, so the chart stamps a digest of analytics.starrocks.* on both pods: a helm upgrade that changes the cut-over or shadow datasets rolls them without a manual restart. Removing metrics, logs or events from the list falls back to CNPG.

The StarRocks Frontend is never reached passwordless. With analytics.starrocks.enabled=true the chart fails to render unless analytics.starrocks.catalog.fePasswordSecretName names a Secret in the release namespace whose key analytics.starrocks.catalog.fePasswordSecretKey (default password) holds the Frontend root password:

analytics.starrocks.catalog.fePasswordSecretName is required when
analytics.starrocks.enabled=true: ...

There is no value that turns this off. core, web-ng (EventWriter Stream Load and the warehouse reader) and the catalog and storage-volume Jobs all read that one Secret as SERVICERADAR_STARROCKS_PASSWORD / FE_PASSWORD. Its value must be byte-identical to the StarRocks operator's initPassword Secret and to the live root password, because the StarRocks pods use the operator's copy to rejoin the cluster whenever they restart. Creating both Secrets for a new cluster, and the ordered procedure that converts an installation already running passwordless (generate, create both Secrets, verify the rendered Deployment, set the password by SQL, upgrade the cluster release, roll ServiceRadar, verify), are in k8s/starrocks/README.md under "Frontend root password". Check the rendered Deployment rather than the values: GitOps parameter overrides, such as an Argo CD .argocd-source-<app>.yaml in the chart directory, are applied after values files and win silently.

k8s/starrocks/network-policy.yaml restricts ingress to the StarRocks namespace to its own pods plus namespaces labelled serviceradar.carverauto.dev/starrocks-client: "true" on ports 9030, 8030 and 8040. Label the ServiceRadar namespace before applying it.

The chart does not create the warehouse schema. Cluster install, the DDL under elixir/serviceradar_core/priv/starrocks/, the CNPG JDBC catalog and its reader role are documented in k8s/starrocks/README.md. Per-key defaults, including retention, are commented in the chart's values.yaml.

Deployment Provisioning​

ServiceRadar does not provision per-customer workloads from inside the Helm chart. Each deployment is self-contained. In managed environments, a separate control plane provisions namespaces, CNPG accounts, and NATS accounts, then installs the chart for that deployment.

After the application is ready, use the supported provisioning API or Terraform provider for its bounded application resources. Follow Declarative environments for the ordered installation, identity bootstrap, configuration, and recovery workflow.

Mapper Discovery Settings​

Mapper discovery is embedded in serviceradar-agent and configured via Settings → Networks → Discovery. Discovery jobs, seeds, and credentials are stored in CNPG and delivered to agents through the GetConfig pipeline.

Configure discovery through Settings or its supported admin API, then trigger an agent config refresh. Do not seed CNPG directly. The current Terraform provider does not manage mapper discovery resources.

Device Enrichment Rule Overrides​

Core always ships with built-in enrichment rules. You can mount filesystem overrides that load from /var/lib/serviceradar/rules/device-enrichment.

Enable override mounting in values:

core:
deviceEnrichment:
rulesDir: /var/lib/serviceradar/rules/device-enrichment
filesystemOverrides:
enabled: true
existingConfigMap: serviceradar-device-enrichment-rules
# Optional alternatives:
# existingSecret: serviceradar-device-enrichment-rules
# existingClaim: serviceradar-device-enrichment-rules

ConfigMap example:

kubectl create configmap serviceradar-device-enrichment-rules \
-n serviceradar \
--from-file=ubiquiti-overrides.yaml=./ubiquiti-overrides.yaml

Apply/verify:

helm upgrade --install serviceradar oci://registry.carverauto.dev/serviceradar/charts/serviceradar \
--version <chart-version> -n serviceradar -f my-values.yaml
kubectl logs deploy/serviceradar-core -n serviceradar | rg "Device enrichment rules loaded"

Rollback to built-ins:

core:
deviceEnrichment:
filesystemOverrides:
enabled: false

UI management:

  • Open Settings → Network → Device Enrichment.
  • Use the typed rule editor to create/update/delete rules.
  • For writable UI-managed rules in Kubernetes, back the mount with a PVC (existingClaim) rather than ConfigMap/Secret.

Outbound Mail​

Configure SMTP in the Web UI: Settings -> Mail. That is the operator path. See Outbound Mail for Local vs Test vs SMTP and the field-by-field setup.

The core.mailer block below is a fallback for automation. An enabled Settings -> Mail row overrides it. Do not put a mailbox password in values; the UI stores credentials encrypted.

# Fallback only. Prefer Settings -> Mail.
core:
mailer:
# "smtp", "local", "sendgrid", ... Leave empty to infer SMTP from `relay`.
adapter: ""
relay: "smtp.example.com"
port: 587
# HELO/EHLO name this deployment announces; empty lets the relay decide.
hostname: ""
auth: "if_available" # always | never | if_available
tls: "if_available" # STARTTLS
ssl: false # implicit TLS (port 465)
from:
name: "ServiceRadar"
# Relay credentials come from an existing Secret, never from values.
existingSecret: "serviceradar-smtp"
usernameKey: "smtp-username"
passwordKey: "smtp-password"

Create the credential Secret separately:

kubectl create secret generic serviceradar-smtp \
-n serviceradar \
--from-literal=smtp-username='serviceradar' \
--from-literal=smtp-password='...'

The password is deliberately not a chart value. A password in values.yaml is a password in the rendered manifest, in helm get values, and in whatever GitOps repository holds the file.

The block renders into these container environment variables on serviceradar-core:

VariableFromMeaning
SERVICERADAR_MAILER_ADAPTERcore.mailer.adaptersmtp, local, test, or an API adapter name
SMTP_RELAY_HOSTcore.mailer.relayRelay hostname. Setting it alone selects the SMTP adapter
SMTP_RELAY_PORTcore.mailer.portDefault 587
SMTP_RELAY_HOSTNAMEcore.mailer.hostnameHELO/EHLO name
SMTP_RELAY_AUTHcore.mailer.authalways, never, if_available
SMTP_RELAY_TLScore.mailer.tlsSTARTTLS mode
SMTP_RELAY_SSLcore.mailer.ssltrue for implicit TLS
SMTP_RELAY_USERNAME / SMTP_RELAY_PASSWORDcore.mailer.existingSecretRelay credentials, by Secret reference
SERVICERADAR_MAIL_FROM_NAME / SERVICERADAR_MAIL_FROM_EMAILcore.mailer.fromDefault From:

With none of this set, the mailer resolves to a non-delivering test adapter that reports every send as successful. That is why an email notification channel refuses to validate until a relay is configured here (or in Settings > Mail) rather than looking healthy and paging nobody. An unrecognised SERVICERADAR_MAILER_ADAPTER value fails the pod's boot with the accepted list, which is a deployment that does not start rather than one that starts and mails nowhere.

Also relevant for notifications: set webNg.publicUrl to the externally reachable HTTPS origin of the web UI. Recent charts copy that value to SERVICERADAR_NOTIFICATION_ACTION_BASE_URL on web-ng and core so acknowledge / snooze / resolve links inside notifications resolve to a real address. See the Notifications Quickstart.

Prometheus Metrics Scraping and Authentication​

The chart provides Prometheus ServiceMonitor resources when observability.enabled and observability.prometheus.serviceMonitors.enabled are true.

By default, serviceradar-web-ng serves metrics on a dedicated internal port (webNg.metricsPort, default 9090). This internal listener is omitted from external Gateway and Ingress routes, allowing in-cluster Prometheus to scrape metrics without exposing them publicly.

On the public HTTP/HTTPS listener (port 4000), /metrics requires bearer token authentication:

  • Unauthenticated requests to port 4000 return 401 Unauthorized.
  • Scrapes against port 4000 require an Authorization: Bearer <token> header.
  • The bearer token is auto-generated in serviceradar-secrets under the web-ng-metrics-token key.
  • If observability.prometheus.serviceMonitors.targets.webNg.port is changed to http, the chart automatically wires bearerTokenSecret to supply the token from serviceradar-secrets.