Plan & Provision Infrastructure
Provision the cluster and services on this page before you install Kindo. Kindo runs on any conformant Kubernetes cluster, with backing services that are managed or self-hosted. When everything is in place, confirm it with the prerequisites checklist.
Kindo requires every item on this page unless it is marked recommended or optional.
1. Choose a size
Section titled “1. Choose a size”Choose the size that matches your expected usage.
| Size | Active users | Agent runs per day |
|---|---|---|
| Medium | 50 to 200 | About 1,500 |
| Large | 200 to 1,000 | About 5,000 |
Cluster capacity
Section titled “Cluster capacity”| Component | Medium | Large |
|---|---|---|
| Cluster nodes (total) | 12 nodes: 96 vCPU, 192 GB RAM | 20 nodes: 320 vCPU, 640 GB RAM |
| Application nodes | 6 × 8 vCPU, 16 GB RAM | 10 × 16 vCPU, 32 GB RAM |
| Sandbox nodes | 6 × 8 vCPU, 16 GB RAM | 10 × 16 vCPU, 32 GB RAM |
| Block storage | 1.5 TiB | 3 TiB |
Required service sizes
Section titled “Required service sizes”| Component | Medium | Large |
|---|---|---|
| PostgreSQL: main | 2 vCPU, 8 GB RAM | 4 vCPU, 16 GB RAM |
| PostgreSQL: auxiliary | 2 vCPU, 8 GB RAM | 2 vCPU, 8 GB RAM |
| Redis | 3 GB RAM | 6 GB RAM |
| RabbitMQ | 2 vCPU, 8 GB RAM | 4 vCPU, 16 GB RAM |
Sizing notes
Section titled “Sizing notes”- Node counts are the maximum each node group needs. With a cluster autoscaler, set them as the node group maximums.
- Sandbox nodes run agents’ tools and scripts. Add sandbox nodes if you expect more agent runs.
- Block storage comes from your StorageClass and grows as people keep persistent agent workspaces. Volumes can grow, where the StorageClass supports it, but never shrink.
- Required services you run in the Kindo cluster add to the cluster capacity. Add their resources, for every replica, to the totals.
Start from these sizes, monitor usage, and adjust as it grows; see Scaling and Tuning.
2. Kubernetes cluster
Section titled “2. Kubernetes cluster”| Attribute | Requirement |
|---|---|
| Version | 1.32 or higher |
| Networking | Any conformant CNI |
| Storage | Dynamic volume provisioning, with a default StorageClass that provisions ReadWriteOnce volumes |
| Ingress | An ingress controller (NGINX, Traefik, AWS Load Balancer Controller, or equivalent); kindo-cli asks for its class name |
| TLS | cert-manager, or certificates you provision yourself; see TLS certificates |
| RBAC | Enabled |
| Metrics | metrics-server installed |
| External Secrets Operator | Kindo installs its own. If the cluster already runs one, it must serve external-secrets.io/v1 (ESO 0.17 or later), or the install stops |
| Volume attach limit | Your CSI driver must allow at least 10 attached volumes per sandbox node at Medium and 20 at Large. If your instance type allows fewer, add sandbox nodes |
| Availability zones | Run nodes in every zone your StorageClass provisions volumes in. A pod with a volume restarts only in its volume’s zone |
| Sandbox node group (optional) | A dedicated, tainted node group keeps agent activity off the application nodes. Set a matching sandbox node selector and toleration at install |
Provider notes
Section titled “Provider notes”- AWS Load Balancer Controller with a non-VPC CNI (such as Calico or Cilium): configure services behind the controller as
NodePortand usetarget-type=instance, or use an in-cluster ingress controller such as NGINX or Traefik.
3. Required services
Section titled “3. Required services”Kindo connects to these services. Run them as managed services, on your own servers, or in the cluster.
PostgreSQL
Section titled “PostgreSQL”| Attribute | Requirement |
|---|---|
| Version | 17.0 or higher (17.4 recommended) |
| Concurrent connections | 200 or more |
Provision a main instance, an auxiliary instance, and, recommended, a dedicated instance for Hatchet, Kindo’s task scheduler. Set the Hatchet instance as postgres.adminHatchet in environment-bindings.yaml; without it, Hatchet uses the auxiliary instance. The installer creates the databases.
PostgreSQL settings for Hatchet
Section titled “PostgreSQL settings for Hatchet”Apply these settings to the instance that hosts the hatchet database, at every size:
| Parameter | Recommended |
|---|---|
autovacuum_max_workers | 10 |
autovacuum_vacuum_scale_factor | 0.1 |
autovacuum_analyze_scale_factor | 0.05 |
autovacuum_vacuum_threshold | 25 |
autovacuum_analyze_threshold | 25 |
autovacuum_vacuum_cost_delay | 10 |
autovacuum_vacuum_cost_limit | 1000 |
Preflight warns about Hatchet settings that differ from these. Fix them before you continue the install.
Provider notes
Section titled “Provider notes”To apply the Hatchet settings:
- AWS RDS / Aurora PostgreSQL: create a custom DB parameter group, set the parameters, and attach it to the Hatchet instance.
- Google Cloud SQL: set them as database flags on the Hatchet instance. Changing
autovacuum_max_workersrequires an instance restart. - Azure Database for PostgreSQL: set them as server parameters on the Hatchet server.
- Self-managed PostgreSQL: set them in
postgresql.confand reload. Changingautovacuum_max_workersrequires a restart.
| Attribute | Requirement |
|---|---|
| Version | 7.0 or higher (7.2 recommended) |
| Concurrent connections | 100 or more |
| Mode | Standalone: a single primary, no shards or replicas |
To confirm the mode, check that this command returns redis_mode:standalone:
redis-cli INFO server | grep redis_modeProvider notes
Section titled “Provider notes”- AWS ElastiCache: Cluster Mode Disabled, with
num_node_groups = 1andreplicas_per_node_group = 0. - Azure Cache for Redis: Basic or Standard tier (single node, non-clustered).
- Google Cloud Memorystore: Basic tier.
RabbitMQ
Section titled “RabbitMQ”| Attribute | Requirement |
|---|---|
| Version | 3.13 or higher |
| Disk | 20 GB minimum |
| Management plugin | Enabled |
S3-compatible object storage
Section titled “S3-compatible object storage”Use any store that serves the S3 API, managed or self-hosted. Create a private bucket for user file uploads, such as kindo-uploads, with server-side encryption and public access blocked.
If the store runs inside the cluster, users’ browsers cannot reach its cluster-local address to download files. Publish it on its own hostname:
- Expose the store at a dedicated HTTPS hostname with a valid certificate and no path prefix, such as
https://s3.kindo.company.com. It must serve path-style requests (https://s3.kindo.company.com/<bucket>/<key>), and any proxy in front of it must forward theHostheader unchanged. - In
environment-bindings.yaml, setstorage.endpointUrlPublicto that hostname, and keepstorage.endpointUrlon the cluster-local address.
Provider notes
Section titled “Provider notes”- Azure Blob Storage: it does not serve the S3 API, so put an S3-compatible gateway in front of it. If the gateway is s3proxy, use 4.0.0 or later; earlier versions break multipart uploads.
Kindo needs an authenticated SMTP relay to send sign-in codes and other email. Use your own SMTP server or a provider that supplies SMTP credentials, such as Amazon SES, SendGrid, or Mailgun. The relay must accept username and password authentication.
Have its host, username, password, and from address ready for the install; the port defaults to 587.
AI models
Section titled “AI models”Kindo needs at least one model endpoint: a cloud provider API, a model server you host, or both. A self-hosted model runs on GPU servers, in this cluster or anywhere on your network the cluster can reach. Choose your models and size any GPUs with Prepare AI Models before you provision.
4. Network
Section titled “4. Network”DNS subdomains
Section titled “DNS subdomains”Pick a parent domain you control; the examples use kindo.company.com. Your ingress load balancer serves every hostname below. Create a DNS record for each, or one wildcard record for *.kindo.company.com. Tools such as external-dns can create the records for you.
| Subdomain | Serves |
|---|---|
app. | The Kindo app |
superadmin. | The superadmin dashboard |
api. | The Kindo API and inbound webhooks |
integrations-api. | The integrations API |
integrations-connect. | Sign-in pages for connecting integrations |
hatchet. | The task scheduler’s API and dashboard |
unleash. | The feature flag service |
unleash-edge. | Feature flag delivery to the app |
hyperdx. | The observability dashboard |
If you run object storage inside the cluster, also publish its public hostname; see S3-compatible object storage.
When kindo-cli manages ingress, it creates the routes for every host. If you manage ingress yourself, print the routes and configure them as described in Configure & Validate:
kindo ingress manifestInbound webhooks
Section titled “Inbound webhooks”Webhooks, including agent trigger URLs such as https://api.<domain>/webhook/agent/<agentId>/<token>, use api.<domain> by default. If api.<domain> is behind an allowlist, keep its /webhook path reachable from the internet, or give webhooks their own hostname:
- Set
endpoints.webhooks: https://webhooks.<domain>in the install contract. - Create a DNS record and a TLS certificate for that hostname. When
kindo-climanages ingress, it publishes the hostname with a route for/webhookonly.
TLS certificates
Section titled “TLS certificates”Use cert-manager with an ACME issuer (such as Let’s Encrypt) or a private CA, or a pre-issued wildcard certificate for *.kindo.company.com stored as a Kubernetes Secret. kindo-cli references the certificate by name; it does not issue or rotate it.
Inbound
Section titled “Inbound”All Kindo traffic is TCP. Open only the ingress entry points, to the internet or to your corporate network if Kindo is internal-only.
| Port | Purpose |
|---|---|
| 443 | HTTPS: all user traffic |
| 80 | HTTP redirect to 443 (optional, recommended) |
Outbound
Section titled “Outbound”The install and Kindo itself fail if the cluster cannot reach these destinations. Confirm them with your network team before you provision.
| Port | Destination | Purpose |
|---|---|---|
| 443 | registry.kindo.ai and other container registries, AI model providers, external object storage | Image pulls, model inference, file storage |
| 5432 | PostgreSQL | Database traffic |
| 6379 | Redis | Cache and streaming |
| 5672 | RabbitMQ | Message queue |
| 25, 587, 465 | SMTP relay |
5. Access and security
Section titled “5. Access and security”Access controls in your environment are the most common cause of a blocked install. Arrange these with your security team early.
- Install identity. The identity that runs the install needs rights to create namespaces, service accounts, role bindings, CRDs, and other cluster-scoped resources.
- Organization guardrails. Review policies that could block the install or the resources it creates: SCPs (AWS), organization policies (GCP), or management group policies (Azure).
- Workload identity. Services that call your cloud provider, such as the model gateway calling Amazon Bedrock, use workload identity: IRSA (AWS), Workload Identity (GCP), or Microsoft Entra Workload ID (Azure).
- Container images. Kindo pulls images from
registry.kindo.aiwith credentials Kindo provides. Allow the registry, or your internal mirror if air-gapped, in your registry allowlist and admission controllers. If you require scanned images, ask Kindo for SBOMs or a mirror-and-scan workflow. - Secrets store. Choose where Kindo keeps its secrets: Kubernetes Secrets (the default), AWS Secrets Manager, Azure Key Vault, or HashiCorp Vault. See External Secrets Store. Decide who can read the
kindo-secrets-configSecret. - Encryption. For production, encrypt PostgreSQL, Redis, object storage, and Kubernetes Secrets at rest, and PostgreSQL connections in transit.
6. Production recommendations
Section titled “6. Production recommendations”For a production install, also plan for the following:
- On-demand nodes. Run production on on-demand nodes. Spot nodes can be reclaimed at any time, which interrupts running agents and restarts stateful services.
- High availability. Run PostgreSQL with streaming replication or as a managed service with automatic failover. Run RabbitMQ as a cluster of 3 or more nodes with quorum queues, or as a managed service.
- Backups. Take daily automated PostgreSQL backups with 7 or more days of retention. Back up Kindo’s configuration as described in Back up and restore.
- Disk alerts. Alert on disk usage on every PostgreSQL instance, especially Hatchet, and on the cluster’s persistent volumes, before a volume fills.
7. Install host
Section titled “7. Install host”Run the install from a bastion, jump box, laptop, or CI runner that can reach the cluster’s API server. Install these tools on it:
| Tool | Version |
|---|---|
python | 3.11 or higher |
kindo-cli | Latest release, provided by Kindo |
kubectl | 1.32 or higher |
helm | 4.0 or higher |
helmfile | 1.2 or higher |
yq | 4.0 or higher (the Go implementation) |
psql | Optional: needed for the kindo db commands after install |
Then confirm that the kubectl context for the target cluster works and has the permissions the install needs:
kubectl cluster-infokubectl auth can-i create namespace --all-namespaces8. Prerequisites checklist
Section titled “8. Prerequisites checklist”Each item names the team that typically owns it, so you can split the list across teams.
| Owner | Item | |
|---|---|---|
| ☐ | Platform | Size chosen (Choose a size) |
| ☐ | Cloud Ops | Cluster provisioned with the required add-ons, volume attach limits, and zone coverage (Kubernetes cluster) |
| ☐ | DBA / Cloud Ops | PostgreSQL instances provisioned and Hatchet settings applied (PostgreSQL) |
| ☐ | DBA / Cloud Ops | Redis provisioned and confirmed standalone (Redis) |
| ☐ | DBA / Cloud Ops | RabbitMQ provisioned (RabbitMQ) |
| ☐ | Cloud Ops | Object storage bucket created, plus a public hostname if the store runs in the cluster (S3-compatible object storage) |
| ☐ | Platform | SMTP relay details in hand (Email) |
| ☐ | Security / AppDev | At least one model endpoint ready, with GPUs sized for any self-hosted model (AI models) |
| ☐ | Networking | Parent domain chosen and DNS records planned (DNS subdomains) |
| ☐ | Security / Networking | TLS approach decided (TLS certificates) |
| ☐ | Networking / Security | Inbound and outbound firewall rules approved (Inbound, Outbound) |
| ☐ | Cloud Ops / Security | Install identity, organization guardrails, and workload identity arranged (Access and security) |
| ☐ | Security | Registry access and credentials, secrets store, and encryption arranged (Access and security) |
| ☐ | Platform / DBA | On-demand nodes, high availability, backups, and disk alerts planned for production (Production recommendations) |
| ☐ | Platform | Install host ready (Install host) |
| ☐ | Platform, Networking, Security | Capacity, network, and security plans approved |
