Skip to content

Plan & Provision Infrastructure

Provision the cluster and services on this page before you install Kindo. Kindo runs on any conformant Kubernetes cluster, with backing services that are managed or self-hosted. When everything is in place, confirm it with the prerequisites checklist.

Kindo requires every item on this page unless it is marked recommended or optional.

Choose the size that matches your expected usage.

SizeActive usersAgent runs per day
Medium50 to 200About 1,500
Large200 to 1,000About 5,000
ComponentMediumLarge
Cluster nodes (total)12 nodes: 96 vCPU, 192 GB RAM20 nodes: 320 vCPU, 640 GB RAM
 Application nodes6 × 8 vCPU, 16 GB RAM10 × 16 vCPU, 32 GB RAM
 Sandbox nodes6 × 8 vCPU, 16 GB RAM10 × 16 vCPU, 32 GB RAM
Block storage1.5 TiB3 TiB
ComponentMediumLarge
PostgreSQL: main2 vCPU, 8 GB RAM4 vCPU, 16 GB RAM
PostgreSQL: auxiliary2 vCPU, 8 GB RAM2 vCPU, 8 GB RAM
Redis3 GB RAM6 GB RAM
RabbitMQ2 vCPU, 8 GB RAM4 vCPU, 16 GB RAM
  • Node counts are the maximum each node group needs. With a cluster autoscaler, set them as the node group maximums.
  • Sandbox nodes run agents’ tools and scripts. Add sandbox nodes if you expect more agent runs.
  • Block storage comes from your StorageClass and grows as people keep persistent agent workspaces. Volumes can grow, where the StorageClass supports it, but never shrink.
  • Required services you run in the Kindo cluster add to the cluster capacity. Add their resources, for every replica, to the totals.

Start from these sizes, monitor usage, and adjust as it grows; see Scaling and Tuning.

AttributeRequirement
Version1.32 or higher
NetworkingAny conformant CNI
StorageDynamic volume provisioning, with a default StorageClass that provisions ReadWriteOnce volumes
IngressAn ingress controller (NGINX, Traefik, AWS Load Balancer Controller, or equivalent); kindo-cli asks for its class name
TLScert-manager, or certificates you provision yourself; see TLS certificates
RBACEnabled
Metricsmetrics-server installed
External Secrets OperatorKindo installs its own. If the cluster already runs one, it must serve external-secrets.io/v1 (ESO 0.17 or later), or the install stops
Volume attach limitYour CSI driver must allow at least 10 attached volumes per sandbox node at Medium and 20 at Large. If your instance type allows fewer, add sandbox nodes
Availability zonesRun nodes in every zone your StorageClass provisions volumes in. A pod with a volume restarts only in its volume’s zone
Sandbox node group (optional)A dedicated, tainted node group keeps agent activity off the application nodes. Set a matching sandbox node selector and toleration at install
  • AWS Load Balancer Controller with a non-VPC CNI (such as Calico or Cilium): configure services behind the controller as NodePort and use target-type=instance, or use an in-cluster ingress controller such as NGINX or Traefik.

Kindo connects to these services. Run them as managed services, on your own servers, or in the cluster.

AttributeRequirement
Version17.0 or higher (17.4 recommended)
Concurrent connections200 or more

Provision a main instance, an auxiliary instance, and, recommended, a dedicated instance for Hatchet, Kindo’s task scheduler. Set the Hatchet instance as postgres.adminHatchet in environment-bindings.yaml; without it, Hatchet uses the auxiliary instance. The installer creates the databases.

Apply these settings to the instance that hosts the hatchet database, at every size:

ParameterRecommended
autovacuum_max_workers10
autovacuum_vacuum_scale_factor0.1
autovacuum_analyze_scale_factor0.05
autovacuum_vacuum_threshold25
autovacuum_analyze_threshold25
autovacuum_vacuum_cost_delay10
autovacuum_vacuum_cost_limit1000

Preflight warns about Hatchet settings that differ from these. Fix them before you continue the install.

To apply the Hatchet settings:

  • AWS RDS / Aurora PostgreSQL: create a custom DB parameter group, set the parameters, and attach it to the Hatchet instance.
  • Google Cloud SQL: set them as database flags on the Hatchet instance. Changing autovacuum_max_workers requires an instance restart.
  • Azure Database for PostgreSQL: set them as server parameters on the Hatchet server.
  • Self-managed PostgreSQL: set them in postgresql.conf and reload. Changing autovacuum_max_workers requires a restart.
AttributeRequirement
Version7.0 or higher (7.2 recommended)
Concurrent connections100 or more
ModeStandalone: a single primary, no shards or replicas

To confirm the mode, check that this command returns redis_mode:standalone:

Terminal window
redis-cli INFO server | grep redis_mode
  • AWS ElastiCache: Cluster Mode Disabled, with num_node_groups = 1 and replicas_per_node_group = 0.
  • Azure Cache for Redis: Basic or Standard tier (single node, non-clustered).
  • Google Cloud Memorystore: Basic tier.
AttributeRequirement
Version3.13 or higher
Disk20 GB minimum
Management pluginEnabled

Use any store that serves the S3 API, managed or self-hosted. Create a private bucket for user file uploads, such as kindo-uploads, with server-side encryption and public access blocked.

If the store runs inside the cluster, users’ browsers cannot reach its cluster-local address to download files. Publish it on its own hostname:

  1. Expose the store at a dedicated HTTPS hostname with a valid certificate and no path prefix, such as https://s3.kindo.company.com. It must serve path-style requests (https://s3.kindo.company.com/<bucket>/<key>), and any proxy in front of it must forward the Host header unchanged.
  2. In environment-bindings.yaml, set storage.endpointUrlPublic to that hostname, and keep storage.endpointUrl on the cluster-local address.
  • Azure Blob Storage: it does not serve the S3 API, so put an S3-compatible gateway in front of it. If the gateway is s3proxy, use 4.0.0 or later; earlier versions break multipart uploads.

Kindo needs an authenticated SMTP relay to send sign-in codes and other email. Use your own SMTP server or a provider that supplies SMTP credentials, such as Amazon SES, SendGrid, or Mailgun. The relay must accept username and password authentication.

Have its host, username, password, and from address ready for the install; the port defaults to 587.

Kindo needs at least one model endpoint: a cloud provider API, a model server you host, or both. A self-hosted model runs on GPU servers, in this cluster or anywhere on your network the cluster can reach. Choose your models and size any GPUs with Prepare AI Models before you provision.

Pick a parent domain you control; the examples use kindo.company.com. Your ingress load balancer serves every hostname below. Create a DNS record for each, or one wildcard record for *.kindo.company.com. Tools such as external-dns can create the records for you.

SubdomainServes
app.The Kindo app
superadmin.The superadmin dashboard
api.The Kindo API and inbound webhooks
integrations-api.The integrations API
integrations-connect.Sign-in pages for connecting integrations
hatchet.The task scheduler’s API and dashboard
unleash.The feature flag service
unleash-edge.Feature flag delivery to the app
hyperdx.The observability dashboard

If you run object storage inside the cluster, also publish its public hostname; see S3-compatible object storage.

When kindo-cli manages ingress, it creates the routes for every host. If you manage ingress yourself, print the routes and configure them as described in Configure & Validate:

Terminal window
kindo ingress manifest

Webhooks, including agent trigger URLs such as https://api.<domain>/webhook/agent/<agentId>/<token>, use api.<domain> by default. If api.<domain> is behind an allowlist, keep its /webhook path reachable from the internet, or give webhooks their own hostname:

  1. Set endpoints.webhooks: https://webhooks.<domain> in the install contract.
  2. Create a DNS record and a TLS certificate for that hostname. When kindo-cli manages ingress, it publishes the hostname with a route for /webhook only.

Use cert-manager with an ACME issuer (such as Let’s Encrypt) or a private CA, or a pre-issued wildcard certificate for *.kindo.company.com stored as a Kubernetes Secret. kindo-cli references the certificate by name; it does not issue or rotate it.

All Kindo traffic is TCP. Open only the ingress entry points, to the internet or to your corporate network if Kindo is internal-only.

PortPurpose
443HTTPS: all user traffic
80HTTP redirect to 443 (optional, recommended)

The install and Kindo itself fail if the cluster cannot reach these destinations. Confirm them with your network team before you provision.

PortDestinationPurpose
443registry.kindo.ai and other container registries, AI model providers, external object storageImage pulls, model inference, file storage
5432PostgreSQLDatabase traffic
6379RedisCache and streaming
5672RabbitMQMessage queue
25, 587, 465SMTP relayEmail

Access controls in your environment are the most common cause of a blocked install. Arrange these with your security team early.

  • Install identity. The identity that runs the install needs rights to create namespaces, service accounts, role bindings, CRDs, and other cluster-scoped resources.
  • Organization guardrails. Review policies that could block the install or the resources it creates: SCPs (AWS), organization policies (GCP), or management group policies (Azure).
  • Workload identity. Services that call your cloud provider, such as the model gateway calling Amazon Bedrock, use workload identity: IRSA (AWS), Workload Identity (GCP), or Microsoft Entra Workload ID (Azure).
  • Container images. Kindo pulls images from registry.kindo.ai with credentials Kindo provides. Allow the registry, or your internal mirror if air-gapped, in your registry allowlist and admission controllers. If you require scanned images, ask Kindo for SBOMs or a mirror-and-scan workflow.
  • Secrets store. Choose where Kindo keeps its secrets: Kubernetes Secrets (the default), AWS Secrets Manager, Azure Key Vault, or HashiCorp Vault. See External Secrets Store. Decide who can read the kindo-secrets-config Secret.
  • Encryption. For production, encrypt PostgreSQL, Redis, object storage, and Kubernetes Secrets at rest, and PostgreSQL connections in transit.

For a production install, also plan for the following:

  • On-demand nodes. Run production on on-demand nodes. Spot nodes can be reclaimed at any time, which interrupts running agents and restarts stateful services.
  • High availability. Run PostgreSQL with streaming replication or as a managed service with automatic failover. Run RabbitMQ as a cluster of 3 or more nodes with quorum queues, or as a managed service.
  • Backups. Take daily automated PostgreSQL backups with 7 or more days of retention. Back up Kindo’s configuration as described in Back up and restore.
  • Disk alerts. Alert on disk usage on every PostgreSQL instance, especially Hatchet, and on the cluster’s persistent volumes, before a volume fills.

Run the install from a bastion, jump box, laptop, or CI runner that can reach the cluster’s API server. Install these tools on it:

ToolVersion
python3.11 or higher
kindo-cliLatest release, provided by Kindo
kubectl1.32 or higher
helm4.0 or higher
helmfile1.2 or higher
yq4.0 or higher (the Go implementation)
psqlOptional: needed for the kindo db commands after install

Then confirm that the kubectl context for the target cluster works and has the permissions the install needs:

Terminal window
kubectl cluster-info
kubectl auth can-i create namespace --all-namespaces

Each item names the team that typically owns it, so you can split the list across teams.

OwnerItem
☐PlatformSize chosen (Choose a size)
☐Cloud OpsCluster provisioned with the required add-ons, volume attach limits, and zone coverage (Kubernetes cluster)
☐DBA / Cloud OpsPostgreSQL instances provisioned and Hatchet settings applied (PostgreSQL)
☐DBA / Cloud OpsRedis provisioned and confirmed standalone (Redis)
☐DBA / Cloud OpsRabbitMQ provisioned (RabbitMQ)
☐Cloud OpsObject storage bucket created, plus a public hostname if the store runs in the cluster (S3-compatible object storage)
☐PlatformSMTP relay details in hand (Email)
☐Security / AppDevAt least one model endpoint ready, with GPUs sized for any self-hosted model (AI models)
☐NetworkingParent domain chosen and DNS records planned (DNS subdomains)
☐Security / NetworkingTLS approach decided (TLS certificates)
☐Networking / SecurityInbound and outbound firewall rules approved (Inbound, Outbound)
☐Cloud Ops / SecurityInstall identity, organization guardrails, and workload identity arranged (Access and security)
☐SecurityRegistry access and credentials, secrets store, and encryption arranged (Access and security)
☐Platform / DBAOn-demand nodes, high availability, backups, and disk alerts planned for production (Production recommendations)
☐PlatformInstall host ready (Install host)
☐Platform, Networking, SecurityCapacity, network, and security plans approved