Skip to content

Scaling and Tuning

Start from the defaults in Choose a size. Watch usage in HyperDX and check install health with kindo status and kindo doctor, then adjust the services that need more or less capacity.

Make every change on this page with a Helm override. Overrides persist across upgrades, so you set each one once.

Add replicas or resources to a service that runs hot, and trim a service that sits well below its requests. For example, to let the api release autoscale and give each pod more memory:

  1. Stage the settings:

    Terminal window
    kindo config helm-override set api \
    autoscaling.enabled=true \
    autoscaling.minReplicas=2 \
    autoscaling.maxReplicas=4 \
    resources.requests.memory=2Gi
  2. Preview and apply them:

    Terminal window
    kindo config helm-override diff api
    kindo config helm-override apply api
  3. Watch the service in HyperDX and adjust again if needed.

Set related values in the same set command. set rejects a path that has no effect on its own, such as autoscaling.minReplicas while autoscaling is off.

Add nodes when pods stay Pending after a change.

Each agent sandbox runs as a pod on a sandbox node, with its own 10 Gi volume. After 15 minutes without activity, a sandbox suspends: its pod stops and its volume stays. After 60 minutes without activity, the sandbox and its volume are deleted. A persistent workspace suspends the same way but keeps its volume until it goes 180 days without use, so sandbox storage grows with use.

Plan for about 60 concurrent sandboxes at Medium and 200 at Large. Each running sandbox also counts against the per-node volume attach limit in the Kubernetes cluster requirements. To support more concurrent sandboxes, add sandbox nodes.

To change the volume size or retention, set values on the sandbox release:

ValueDefaultSets
openshell.gatewayConfig.workspaceDefaultStorageSize10GiVolume size for each sandbox created after the change
openshell.gatewayConfig.workspaceStorageClass(empty)StorageClass for sandbox volumes; empty uses the cluster default
api.sandboxLifecycle.deleteAfterMinutes60Minutes without activity before a sandbox and its volume are deleted
api.sandboxLifecycle.persistentDeleteAfterMinutes259200Minutes without use before a persistent workspace and its volume are deleted; at least deleteAfterMinutes

A retention change applies to existing persistent workspaces.

For example, to give each sandbox a 20 Gi volume and keep persistent workspaces for 90 days:

Terminal window
kindo config helm-override set sandbox \
openshell.gatewayConfig.workspaceDefaultStorageSize=20Gi \
api.sandboxLifecycle.persistentDeleteAfterMinutes=129600
kindo config helm-override apply sandbox

A larger volume size also raises the block storage you need, so check that your StorageClass has room.

The PII analyzer runs one replica with 4 workers, requesting 2.5 CPU and 4 GiB of memory. Its memory use scales with the number of workers, so change workers and memory together.

To shrink it on a smaller install:

Terminal window
kindo config helm-override set presidio \
analyzer.container.env.WORKERS=2 \
analyzer.container.resources.requests.memory=2Gi \
analyzer.container.resources.limits.memory=2Gi
kindo config helm-override apply presidio

To handle more PII detection traffic, raise workers and memory in the same proportion.

ClickHouse stores telemetry on two 300 Gi volumes and keeps it for 7 days. To keep more or less history, change the retention period in Signals and retention.

To add space, expand the ClickHouse volumes. Your StorageClass must allow volume expansion. Volumes can grow but cannot shrink.

  1. List the ClickHouse data volumes:

    Terminal window
    kubectl -n clickhouse get pvc
  2. Expand each data- volume to the new size, for example:

    Terminal window
    kubectl -n clickhouse patch pvc <pvc-name> \
    -p '{"spec":{"resources":{"requests":{"storage":"500Gi"}}}}'
  3. Confirm the new capacity:

    Terminal window
    kubectl -n clickhouse get pvc