Skip to content

LinearB On-Prem Agent v5 Upgrade Guide

This guide covers all configuration changes required when upgrading the LinearB On-Prem Agent from v4 to v5.

Before upgrading, back up your local-values.yaml and run a helm diff or Argo CD dry-run to preview the changes against your current release.


Breaking changes — action required before upgrading

1. Scheduler worker subcharts removed

v5 moves to native Kubernetes Jobs for processing work. The three Celery-based worker subcharts are removed.

Remove these keys from your local-values.yaml:

# Delete any overrides for these — the subcharts no longer exist
scheduler-worker: ...
scheduler-pm-worker: ...
scheduler-sensors-worker: ...

Any resources block under those keys sized the Celery worker Deployment itself. Those Deployments are gone and need no replacement — the v5 scheduler Deployment ships its own requests and limits, which you can override with scheduler.resources.

Job pod sizing

The WORKER_REQUEST_CPU, WORKER_REQUEST_MEMORY, WORKER_LIMIT_CPU and WORKER_LIMIT_MEMORY variables still control the resources of the Job pods that run your analysis work. In v4 each worker subchart carried its own set; in v5 the single scheduler reads them once and applies them to every Job pod it creates — git, PM, sensors and Ops alike.

Do not skip this step, even if you never tuned these values. The v4 chart set them for you: 400m CPU / 100Mi memory requests for git and PM Job pods, 200m / 100Mi for sensors, with no limits. The v5 chart ships no values for them at all, so once the worker blocks are gone they fall back to their built-in "0" and Job pods are created with zero resource requests and no limits at all. That happens whether or not you had overrides of your own — the baseline goes with the subcharts. Pods with zero requests are invisible to the Kubernetes scheduler's capacity accounting, so several can stack onto one node and exhaust it. They are also classified BestEffort — for quality-of-service purposes a request of "0" does not count as a request at all — and under node pressure the kubelet evicts pods that exceed their requests ahead of pods still within theirs, which a zero-request pod always does. Set all four explicitly.

Move your values to the scheduler key:

scheduler:
  image:
    env:
      plain:
        - name: WORKER_REQUEST_CPU
          value: "500m"
        - name: WORKER_REQUEST_MEMORY
          value: "1024Mi"
        - name: WORKER_LIMIT_CPU
          value: "2000m"
        - name: WORKER_LIMIT_MEMORY
          value: "4Gi"

"0" does not mean the same thing on both sides, so set each value deliberately:

  • On a limit — WORKER_LIMIT_CPU, WORKER_LIMIT_MEMORY — "0" omits the limit, leaving that resource unbounded. Use it when you genuinely want no ceiling.
  • On a request — WORKER_REQUEST_CPU, WORKER_REQUEST_MEMORY — "0" is a literal zero request, not "leave it unset". That is the case the warning above describes: the pod is still scheduled, but it claims no share of the node.

Because one set of values now covers every job type, sizing can no longer differ per queue. If your v4 configuration relied on PM jobs being sized differently from git jobs, choose values that suit the largest of them and contact LinearB support.

Node placement for these Job pods is configured separately, under global.workerPods — see Job pod placement below.

2. Renamed keys

Old key (v4) New key (v5)
global.versions.pods_cleaner global.versions.job_lifecycle_manager
minio.DeploymentUpdate minio.deploymentUpdate

Update your local-values.yaml wherever these appear.

3. Deprecated concurrency keys

The global job concurrency limit is replaced by per-queue limits.

Remove:

global:
  MAX_CONCURRENT_JOBS: "10"    # remove
  OPS_RESERVED_SLOTS: "2"      # remove

Add:

global:
  MAX_GIT_JOBS: "5"       # concurrent git (code analysis) jobs
  MAX_PM_JOBS: "1"        # concurrent PM (Jira/Linear/etc.) jobs
  MAX_SENSORS_JOBS: "5"   # concurrent sensors jobs

The defaults shown above match v4 behaviour for most installations. Ops jobs bypass concurrency limits in both versions.

4. Datadog chart API changes

If you have Datadog enabled, remove the following keys — they no longer exist in the upgraded Datadog subchart:

datadog-agent:
  datadog:
    kubeStateMetricsEnabled: true          # remove
    containerExcludeLogs: "image:..."      # remove
  kubeStateMetrics:
    image: ...                             # remove entire block

The upgraded Datadog subchart enables the Datadog Operator by default. The LinearB chart already disables it for you, so no action is required and you do not need to add datadog-agent.datadog.operator.enabled to your values file — overrides you set alongside it are merged with the chart default, not replacing it.

Only set it explicitly if you deliberately want the Operator installed:

datadog-agent:
  datadog:
    operator:
      enabled: true   # opt in — installs a Datadog Operator Deployment and its CRDs

Take care when overriding datadog-agent.datadog.env: it is a list, so anything you set replaces the chart's entries wholesale rather than merging with them. If you add your own variables there, repeat the entries the chart already defines, or integrations that depend on them will lose their configuration.

5. Bundled log collector (fluent-bit) removed

v5 (from v5.0.3) removes the bundled fluent-bit log collector subchart entirely. Every service logs structured JSON to stdout, so any cluster log collector picks the logs up with no agent-side configuration. If global.installFluentBit: true is still set, the chart fails the upgrade. installFluentBit: false and a fluent-bit: block on their own (for example the fluent-bit.image.registry override from the custom-registry examples) only print a warning, because the collector was never installed without installFluentBit: true.

Remove these keys from your local-values.yaml:

global:
  installFluentBit: true    # remove

fluent-bit:                  # remove entire block
  ...

See the new Log collection guide for Fluent Bit and OpenTelemetry examples. To deploy your own collector as part of this release, use the new top-level extraObjects key.

Orphaned log files after upgrade. fluent-bit wrote its buffered logs as plain files under the logs/ subdirectory of the minio-pvc PVC — the same volume that holds MinIO's object data. That PVC carries helm.sh/resource-policy: keep, so Helm never touches it, and once the fluent-bit DaemonSet is gone those files have no owner and persist indefinitely. After upgrading, inspect logs/ on that volume and delete whatever you no longer need. Do not delete anything outside logs/ — the rest of that volume is live MinIO object data.

6. Unused Redis sentinel/metrics keys removed

redis.sentinel and redis.metrics are removed from the chart, along with the redis-sentinel and redis-exporter images. Neither was ever deployed by default (both shipped with enabled: false), but if either is still explicitly set to true in your values file, the chart now fails the upgrade rather than letting it pull an image that is no longer mirrored.

Neither could have worked here anyway. Redis runs as a single standalone master, so Sentinel had no replicas to fail over between, and every service connects straight to the -master Service rather than through a Sentinel-aware client. Nothing scraped the exporter either: the agent ships no Prometheus, and the Datadog integration has no Redis check. If you run your own Prometheus, deploy an exporter via the top-level extraObjects key.

Remove:

redis:
  sentinel:
    enabled: true    # remove
  metrics:
    enabled: true    # remove

7. MinIO strategyMigrator Job removed

infra.minio.strategyMigrator is removed, along with the kubectl image and the ServiceAccount/Role/RoleBinding the Job needed. If the value is still set in your values file, the chart now fails the release and prints the replacement command.

It existed only to delete the live MinIO Deployment once, so a stale spec.strategy.rollingUpdate could not block the Recreate strategy — a one-time step that only ever affected server-side apply (Argo CD with ServerSideApply=true) over a MinIO Deployment first created by Helm or by a client-side apply. It was opt-in, so upgrades that hit this case have always needed the manual step unless it was enabled.

Remove:

infra:
  minio:
    strategyMigrator:
      enabled: true    # remove

If you sync with Argo CD + SSA and have not yet made this transition — including moving a 4.x release that was installed with helm onto Argo CD — run the check and the delete yourself before syncing — see MinIO update strategy migration below. On plain helm upgrade there is nothing to do: Helm's three-way merge nulls the stale field itself.

8. Unused volumePermissions / sysctl init-container keys removed

The os-shell image is no longer mirrored, so the image pins under rabbitmq.volumePermissions, redis.volumePermissions and redis.sysctl are gone. All three shipped enabled: false and none was ever deployed; the keys remain (still false) so the subcharts stay explicitly opted out, but the chart now fails the release if any is set to true.

volumePermissions only chowns a data directory on storage classes that ignore fsGroup — none of the supported platforms do, and redis.master.persistence is disabled anyway. sysctl needs a privileged container to raise net.core.somaxconn, which this Redis never approaches.

Remove:

rabbitmq:
  volumePermissions:
    enabled: true    # remove
redis:
  volumePermissions:
    enabled: true    # remove
  sysctl:
    enabled: true    # remove

MinIO update strategy migration (Argo CD with ServerSideApply only)

If you use plain helm upgrade, skip this section.

See README-upgrade-to-v4.md for background. Every 4.x chart ran MinIO with a RollingUpdate strategy; 5.x uses Recreate. If the sync fails with:

Deployment.apps "<release>-minio" is invalid: spec.strategy.rollingUpdate:
Forbidden: may not be specified when strategy `type` is 'Recreate'

the stale spec.strategy.rollingUpdate fields were written by a different field manager than the server-side apply now syncing the chart, so the apply cannot remove them. You are affected when:

  • your 4.x release was installed or upgraded with helm, and 5.x is the first version synced by Argo CD with ServerSideApply=true, or
  • Argo CD synced your 4.x release without ServerSideApply, and you turn it on for 5.x.

You are not affected if Argo CD has used ServerSideApply=true since its first sync, if Argo CD stays on client-side apply, or if you upgrade with helm upgrade.

Check, and delete once if needed, before syncing — or, if a sync already failed with the error above, delete and sync again. minio-pvc carries helm.sh/resource-policy: keep, so your data survives:

kubectl -n <namespace> get deploy <release>-minio -o jsonpath='{.spec.strategy.type}'

If that prints RollingUpdate:

kubectl -n <namespace> delete deploy <release>-minio

If it prints Recreate, there is nothing to do — this transition has already happened on your cluster.

The infra.minio.strategyMigrator Job that used to do this for you was removed in 5.0.3; see item 7 above.


New features

Per-queue job tuning

You can tune the job lifecycle behaviour per environment. All of the following have working defaults and only need to be set if you want to change them:

global:
  # How long before a stuck job is forcibly terminated (seconds; default 4 hours)
  JOB_ACTIVE_DEADLINE_SECONDS: "14400"

  # How long completed / failed job pods are kept before cleanup
  JOB_TTL_SUCCESS_SECONDS: "60"
  JOB_TTL_FAILED_SECONDS: "21600"

  # Retry behaviour
  MAX_JLM_RETRY_ATTEMPTS: "5"
  SCHEDULER_REQUEUE_DELAY: "600"
  MAX_SCHEDULER_RETRY_ATTEMPTS: "10"

Job pod placement

global.workerPods controls where the Job pods created by the scheduler are placed. It carries scheduling constraints only — CPU and memory for those pods come from the scheduler's WORKER_* variables described in Job pod sizing.

Both keys are optional; without them Job pods are scheduled normally.

global:
  workerPods:
    # Simple node labels
    nodeSelector:
      node-type: worker

    # Or full affinity rules
    affinity:
      nodeAffinity:
        requiredDuringSchedulingIgnoredDuringExecution:
          nodeSelectorTerms:
            - matchExpressions:
                - key: kubernetes.io/arch
                  operator: In
                  values: ["amd64"]

Argo CD namespace RBAC

If you deploy via Argo CD with the application controller running in a separate namespace, the chart can create the necessary RoleBinding for you:

global:
  argocdNamespaceRbac:
    enabled: true
    controllerServiceAccountName: "argocd-application-controller"
    controllerNamespace: "argocd"
    clusterRoleName: "admin"

Default is false. Leave disabled if you manage RBAC out-of-band.

OpenShift + Argo CD (openshift-gitops): The GitOps service account typically does not have escalation rights to bind admin. Keep this disabled and grant RBAC to the GitOps SA via your cluster's infrastructure tooling instead.

Secret management via External Secrets Operator (ESO)

v5 adds support for supplying secrets through the External Secrets Operator. Set global.infra.commonSecret.mode to choose how the chart-secrets Kubernetes Secret is populated:

Mode Description
helm Default. Helm renders the Secret from your values file.
externalSecret ESO syncs the Secret from a ClusterSecretStore. Requires ESO installed in your cluster.
existing You pre-create the Secret outside the chart (e.g. Terraform, Vault agent).

ESO example — sync from AWS Secrets Manager:

global:
  infra:
    commonSecret:
      mode: externalSecret
    externalSecrets:
      refreshInterval: "1h"
      secretStoreRef:
        name: my-cluster-secret-store   # name of your ClusterSecretStore
        kind: ClusterSecretStore
      envName: prod
      data:
        - secretKey: LINEARB_PUBLIC_API_KEY
          remoteRef:
            property: LINEARB_PUBLIC_API_KEY
        - secretKey: DD_API_KEY
          remoteRef:
            property: DD_API_KEY
        - secretKey: api-key
          remoteRef:
            property: DD_API_KEY
        - secretKey: JFROG_KEY
          remoteRef:
            property: JFROG_KEY
        - secretKey: JFROG_USER
          remoteRef:
            property: JFROG_USER

The image pull secret (regcred) can also be managed by ESO:

infra:
  createRegcred: false
  useExternalRegcred: true
  regcredExternalSecret:
    enabled: true

Custom root CA — ESO mode

If you supply a custom root certificate and want it sourced from a secret store rather than embedded in your values file:

global:
  CUSTOM_ROOT_CERTIFICATES: true
  customCa:
    secretMode: eso
    externalSecret:
      refreshInterval: 1h
      secretStoreRef:
        name: my-cluster-secret-store
        kind: ClusterSecretStore
      remoteRef:
        key: my-secret-id
        property: CUSTOM_SSL_CERT

The secret store property must be populated before the release is applied. If the ExternalSecret remains in Pending state, the CA bundle init container will fail to start.


OpenShift: security context requirements

OpenShift's restricted-v2 SCC requires pods to run with UIDs within your namespace's allocated range. v5 retains the same security context requirements as v4.

Find your namespace UID range:

oc get ns <your-namespace> -o jsonpath='{.metadata.annotations.openshift\.io/sa\.scc\.uid-range}{"\n"}'
# example: 1000900000/10000  →  use 1000900000 as your base UID

Set the global OpenShift UID

LinearB's own workloads pick up security context automatically when K8S_PLATFORM: openshift is set. Configure the namespace UID once globally:

global:
  K8S_PLATFORM: openshift
  openshiftRunAsUser: 1000900000   # replace with your namespace UID range low bound
  openshiftRunAsGroup: 1000900000

Bitnami subchart overrides (MinIO, RabbitMQ, Redis)

The MinIO, RabbitMQ, and Redis subcharts do not read the global OpenShift UID and must be overridden explicitly. Replace 1000900000 throughout with your namespace's UID range low bound.

minio:
  persistence:
    subPath: minio-export
  securityContext:
    enabled: true
    runAsNonRoot: true
    runAsUser: 1000900000
    runAsGroup: 1000900000
    fsGroup: 1000900000
    fsGroupChangePolicy: "Always"
  containerSecurityContext:
    runAsNonRoot: true
    runAsUser: 1000900000
    runAsGroup: 1000900000
    allowPrivilegeEscalation: false
    capabilities:
      drop: ["ALL"]
    seccompProfile:
      type: RuntimeDefault
  postJob:
    securityContext:
      enabled: true
      runAsNonRoot: true
      runAsUser: 1000900000
      runAsGroup: 1000900000
      fsGroup: 1000900000
  makeBucketJob:
    securityContext:
      enabled: true
    containerSecurityContext:
      runAsNonRoot: true
      runAsUser: 1000900000
      runAsGroup: 1000900000
      allowPrivilegeEscalation: false
      capabilities:
        drop: ["ALL"]
      seccompProfile:
        type: RuntimeDefault
  makeUserJob:
    securityContext:
      enabled: true
    containerSecurityContext:
      runAsNonRoot: true
      runAsUser: 1000900000
      runAsGroup: 1000900000
      allowPrivilegeEscalation: false
      capabilities:
        drop: ["ALL"]
      seccompProfile:
        type: RuntimeDefault

rabbitmq:
  podSecurityContext:
    runAsNonRoot: true
    runAsUser: 1000900000
    runAsGroup: 1000900000
    fsGroup: 1000900000
    fsGroupChangePolicy: Always
  containerSecurityContext:
    runAsNonRoot: true
    runAsUser: 1000900000
    runAsGroup: 1000900000
    allowPrivilegeEscalation: false
    capabilities:
      drop: ["ALL"]
    seccompProfile:
      type: RuntimeDefault

redis:
  master:
    podSecurityContext:
      runAsNonRoot: true
      runAsUser: 1000900000
      runAsGroup: 1000900000
      fsGroup: 1000900000
      fsGroupChangePolicy: Always
    containerSecurityContext:
      runAsNonRoot: true
      runAsUser: 1000900000
      runAsGroup: 1000900000
      allowPrivilegeEscalation: false
      capabilities:
        drop: ["ALL"]
      seccompProfile:
        type: RuntimeDefault

fsGroupChangePolicy: Always on MinIO and RabbitMQ forces kubelet to sync PVC ownership on every mount. This prevents startup failures ("Unable to write to the backend") if the UID changed between deployments.

minio.persistence.subPath: minio-export mounts MinIO's data into a subdirectory of the PVC so new data is created under the current UID. Without this, data written before a UID change may be inaccessible after upgrade.

Supplemental group for analysis workloads

The sensors, onprem-receiver, jobs-suite-worker, and jobs-suite-dispatcher containers write to paths that require membership in group 0. Add the following to your OpenShift values layer:

sensors:
  podSecurityContext:
    runAsNonRoot: true
    seccompProfile:
      type: RuntimeDefault
    supplementalGroups: [0]

onprem-receiver:
  podSecurityContext:
    runAsNonRoot: true
    seccompProfile:
      type: RuntimeDefault
    supplementalGroups: [0]

jobs-suite-worker:
  podSecurityContext:
    runAsNonRoot: true
    seccompProfile:
      type: RuntimeDefault
    supplementalGroups: [0]

jobs-suite-dispatcher:
  podSecurityContext:
    runAsNonRoot: true
    seccompProfile:
      type: RuntimeDefault
    supplementalGroups: [0]

agent-api runtime flags

None are needed. The agent-api image runs uvicorn directly and requires no OpenShift-specific environment overrides.

Earlier v5 revisions asked you to set PORT and GUNICORN_CMD_ARGS on agent-api:

# No longer required -- safe to remove from your values file.
agent-api:
  image:
    env:
      plain:
        - name: "PORT"
          value: "8080"
        - name: "GUNICORN_CMD_ARGS"
          value: "--no-control-socket --worker-tmp-dir /tmp"

The image no longer runs gunicorn, so GUNICORN_CMD_ARGS has nothing to configure, and the listening port is fixed in the image to match the chart's service.port. Leaving the block in place does no harm, but it has no effect.

Forward-proxy DNS resolver

OpenShift uses a different internal DNS service than standard Kubernetes. Set the resolver so the forward proxy can resolve upstream hostnames:

forward-proxy:
  resolver: "dns-default.openshift-dns.svc.cluster.local"

OpenShift Route for webhook receiver

On OpenShift the webhook receiver is exposed via an OpenShift Route rather than an Ingress. Configure it under infra:

infra:
  openshift:
    onpremReceiverRoute:
      annotations:
        haproxy.router.openshift.io/timeout: 120s
      # host: ""  # optional — omit to use the cluster-assigned hostname

Ensure the bundled ingress-nginx controller is not installed by setting it under global:

global:
  RECEIVER_INGRESS: false   # the Route replaces the receiver's Ingress resource
  ingress:
    installController: false

installController must be set under global.ingress. A top-level ingress.installController has no effect on whether the controller is installed — the chart decides that from global.ingress.installController, falling back to global.RECEIVER_INGRESS. If neither is set the controller is installed regardless of any top-level value, so check this key if you intend to use an ingress controller you manage yourself. That key behaves the same way on standard Kubernetes as on OpenShift.

Do not copy RECEIVER_INGRESS: false outside this OpenShift section. It also removes the receiver's own Ingress resource, which is exactly what the Route replaces here. On standard Kubernetes set only global.ingress.installController: false — that drops the controller and leaves the receiver's Ingress in place.


Upgrade steps

  1. Update your local-values.yaml with the breaking changes listed above. The chart fails the upgrade if any key the items above tell you to remove is still set.
  2. Pull the latest chart:
    helm repo update
    
  3. Run a dry-run to preview changes:
    helm diff upgrade <release-name> linearb/on-prem-agent \
      --namespace <namespace> \
      -f local-values.yaml
    
  4. Apply the upgrade:
    helm upgrade <release-name> linearb/on-prem-agent \
      --namespace <namespace> \
      -f local-values.yaml \
      --wait
    
  5. Verify all pods come up:
    kubectl get pods -n <namespace>
    

For issues during or after the upgrade, refer to the Diagnostics Guide.