Runs the agent on Cloud Run with workspace state in Cloud Storage, secrets in Secret Manager, and the periodic sweep as a scheduled Cloud Run Job.
No Kubernetes. Cloud Run gives us a managed HTTPS endpoint, scale-to-zero, per-revision rollback and workload identity, and the two things that actually need a long-lived process — webhooks and the schedule — are handled by a request-driven service and a cron-triggered job respectively.
┌──────────────────────────────┐
Slack / Teams ─────▶│ Cloud Run service │
Console / API ─────▶│ cisoexpress-prod │
│ request-driven, 1 instance │
└───────────┬──────────────────┘
│
┌───────────▼──────────────────┐
│ gs://…/prod/platform.json │ tenant registry (auth, limits, cadence)
│ gs://…/prod/tenants/<slug>/ │ one workspace document per customer
└───────────▲──────────────────┘
│
Cloud Scheduler ──▶ ┌───────────┴──────────────────┐
every 15 min │ Cloud Run job: sweep │
│ one pass, then exit │
└──────────────────────────────┘
Why a job for the sweep
A timer inside the service would need CPU allocated around the clock, would be lost on every redeploy, and would run once per instance. Moving the sweep into a job means:
- the service scales to zero and runs with CPU throttled between requests;
- each workspace's cadence lives in the registry (
schedule.lastSweepAt), so it survives restarts and a job run knows exactly what is due; - a failed sweep is a failed execution — visible, retryable and alertable — instead of a log line.
Two writers, one document
The service and the sweep job both write the registry and the workspace documents. This is deliberate and safe:
- writes are version-checked (
ifGenerationMatch), so a stale writer finds out instead of overwriting; - a losing writer replays its pending changes onto the reloaded document, so both sets of changes survive;
- counters are incremented, never replaced, which is why usage from the service and schedule marks from the job can land in the same document seconds apart.
This is not theoretical. Before it was fixed, the service and the job each wrote their in-memory copy, and the bucket
version history showed platform.json alternating between 975 and 1803 bytes — a provisioned workspace was erased by
a process that had loaded the registry before it existed. services/agent/src/__tests__/concurrency.test.ts pins the
behaviour now.
Operational rule that came out of it: when you ship a fix to the write path, delete the old revisions
(gcloud run revisions delete). A lingering pre-fix revision can still write with the old semantics.
Prerequisites
gcloud auth login
gcloud auth application-default login # Terraform uses ADC, or pass a short-lived token (see below)
gcloud config set project <project-id>
# Terraform (Homebrew's core formula moved; use HashiCorp's tap)
brew install hashicorp/tap/terraform
Billing must be enabled on the project. It is not optional: Cloud Run, Artifact Registry and Secret Manager all require it.
Deploy
Three phases, because the first apply needs an image and the image needs a registry.
cd deploy/terraform
export TF_VAR_access_token="$(gcloud auth print-access-token)" # only if you have no ADC
# TF_VAR_access_token is optional; ADC is the normal path.
# 1. Enable the APIs and create the image repository.
terraform init
terraform apply -var project_id=<project-id> \
-target=google_project_service.enabled -target=google_artifact_registry_repository.images
# 2. Build and push the image (Cloud Run needs linux/amd64, so buildx, not plain docker build).
cd ../..
REPO="$(cd deploy/terraform && terraform output -raw image_repository)"
TAG=v0.2.2
docker buildx build --platform linux/amd64 -f services/agent/Dockerfile -t "$REPO:$TAG" --push .
# 3. Everything else.
cd deploy/terraform
terraform apply -var project_id=<project-id> -var image_tag=$TAG
terraform output next_steps
Grab the operator key and open the console:
gcloud secrets versions access latest --project <project-id> --secret=cisoexpress-prod-admin-api-key
open "$(terraform output -raw console_url)"
Provision the first customer
ADMIN=$(gcloud secrets versions access latest --project <project-id> --secret=cisoexpress-prod-admin-api-key)
curl -X POST "$(terraform output -raw api_url)/api/platform/tenants" \
-H "X-API-Key: $ADMIN" -H 'Content-Type: application/json' \
-d '{"slug":"acme","name":"Acme Freight","domain":"acmefreight.com","plan":"growth","frameworks":["soc2","iso27001"]}'
The response contains the workspace API key once. Hand it over as
https://<console>/console/?tenant=acme&key=ce_live_…, which the console stores locally and strips from the URL.
Releases
TAG=$(git rev-parse --short HEAD)
docker buildx build --platform linux/amd64 -f services/agent/Dockerfile \
-t "$(terraform output -raw image_repository):$TAG" --push .
gcloud run deploy cisoexpress-prod --region <region> --image "<image>:$TAG"
gcloud run jobs update cisoexpress-prod-sweep --region <region> --image "<image>:$TAG"
Terraform creates the service from image_tag but ignores later image drift on purpose, so releases roll a
revision without Terraform trying to roll it back. Infrastructure changes stay a deliberate terraform apply.
enable_ci = true creates a Workload Identity Federation pool so .github/workflows/deploy.yml can do the above
without a service account key; the outputs give you the two repository variables to set.
Subdomains (acme.ciso.express)
Cloud Run domain mappings cannot be wildcards, so per-tenant subdomains need a load balancer:
platform_domain = "ciso.express"
enable_custom_domain = true
That adds a global external HTTPS load balancer with a Google-managed certificate for ciso.express and
*.ciso.express — roughly $18/month for the forwarding rule, on top of Cloud Run. The custom_domain output
gives you the two A records to create (apex and wildcard); the certificate takes 15–60 minutes after DNS resolves.
Without it, workspaces are addressed with X-Tenant and ?tenant= links, which the console and API both support.
Connectors
Verification connectors (GitHub, Okta, AWS) are configured per workspace, by the customer,
through the console or PATCH /api/connectors/:kind. Their credentials are sealed with
TENANT_SECRET_KEY and stored on the tenant record, and the API only ever reports which secrets
are present.
That is the intended path for a SaaS deployment: the customer hands over a read-only token and nothing long-lived sits in the platform's environment.
For a self-hosted deployment that wants its own workspace connected without anyone typing a
token into a browser, .env.example lists platform-level defaults (GITHUB_TOKEN, OKTA_API_TOKEN,
AWS_ROLE_ARN, …). They apply to the platform workspace only — a platform GitHub token can never be
used to scan a customer's organisation.
# What the installed connectors cover, and what they do not:
npm run coverage
npm run coverage -- --gaps
Connector syncs run as part of the sweep job's daily evidence pass, so they are covered by the Cloud Scheduler trigger above and need no extra infrastructure.
What it costs
| Item | Roughly | Notes |
|---|---|---|
| Cloud Run service | $0–5 /month | Scales to zero; you pay for requests |
| Cloud Run sweep job | < $1 /month | Four short runs an hour |
| Cloud Scheduler | free | 3 jobs free per month |
| Cloud Storage | < $1 /month | A few KB per workspace, plus versions |
| Secret Manager | < $1 /month | Two secrets, negligible reads |
| Artifact Registry | ~$1 /month | Keep the cleanup policies enabled |
| Load balancer | ~$18 /month | Only with enable_custom_domain |
Secrets
Three secrets, all created here and stored in Secret Manager:
| Secret | Purpose | Rotation |
|---|---|---|
…-admin-api-key |
Platform operator: provisions and manages workspaces | terraform apply after changing the value; the service picks it up on the next revision |
…-tenant-secret-key |
Encrypts customer credentials at rest (AES-256-GCM) | Changing this makes existing tenant credentials unreadable — they must be re-entered |
…-llm-api-key |
Model provider key, when llm_provider is set |
terraform apply after changing the value |
Secret values live in Terraform state. That is the normal trade-off for managing them here; keep the state bucket
private (versions.tf has the backend block) and never commit terraform.tfvars.
Scaling
max_instance_count = 1 on the service is intentional. Workspace documents are cached in memory per instance, so a
second instance would serve a customer stale reads and — before the version-checking above — could lose writes.
Raising it needs a shared cache or a store that can transact per document, not just a Terraform change.
The sweep job is safe to run repeatedly and safe to overlap for different workspaces, which is why a bounded
SCHEDULER_MAX_TENANTS_PER_TICK exists: one run should finish comfortably inside the schedule.
Troubleshooting
| Symptom | Cause |
|---|---|
403 Forbidden from Google's frontend |
The service has no public invoker binding. Set public_service = true, or keep it private and put the load balancer in front. |
must support amd64/linux |
The image was built on Apple Silicon without --platform linux/amd64. |
A workspace key returns 401 Unknown API key |
The key's workspace is missing from the registry. Check gcloud storage cat gs://…/prod/platform.json. The workspace document survives independently — re-provision the same slug to restore access with a new key. |
| Revision unhealthy right after apply | Usually a secret reference the runtime identity cannot read; the IAM bindings are ordered before the service for this reason. |
cannot destroy service without setting deletion_protection=false |
Terraform refuses to delete a protected service. Set deletion_protection = false, apply, then flip it back if you want the guard rail. |
First port of call for anything else:
gcloud run services describe cisoexpress-prod --region <region> --format="value(status.conditions[0].message)"
gcloud logging read 'resource.type="cloud_run_revision" AND severity>=WARNING' --limit=20 --freshness=1h