The one decision everything else follows from
The agent's behaviour is rule-first. The language model is a writer, not a decision-maker.
SLAs, risk scores, escalation order, ownership resolution and follow-up windows are plain arithmetic over typed
data in packages/shared and services/agent/src/agent/triage.ts. The model is called in exactly two places:
- Composing a message when no playbook fits (
services/agent/src/agent/outreach.ts). - Interpreting an ambiguous reply from a human, mapped onto a fixed set of intents (
agent/inbound.ts).
This is not an anti-LLM stance, it is what makes the product sellable. A CISO has to answer "why did you page my CTO at 23:40?" and "why did this close without evidence?" with a reproducible rule, in front of an auditor. A model that improvises escalation decisions cannot be defended, cannot be tuned, and cannot be tested. Everything deterministic is covered by the test suite; the model's job is to sound like a competent human.
The second consequence: the agent degrades instead of failing. With no API key, an expired key, or a model outage, every capability still works in template-driven form. The boot probe reports model health to the console so an operator sees the truth rather than a silent downgrade.
Layers
packages/shared Pure domain. No I/O, no runtime dependencies.
domain/types.ts Person, Issue, Control, ActivityEvent, AgentTask, Organization...
domain/roles.ts 14 roles, their accountabilities, and the escalation ladder per issue category
domain/catalog.ts 90 control definitions with owner role, automation level, evidence hint, agent question
domain/scoring.ts SLA burn, explainable risk score, framework readiness, posture summary
agent/playbooks.ts 16 message templates plus the rules that choose between them
state/store.ts WorkspaceState + pure reducers (assignIssue, updateIssueStatus, acceptRisk, addEvidence...)
fixtures/seed.ts Helios Robotics: a mid-flight tenant, not a happy path
api/client.ts Typed REST client
services/agent All I/O and side effects.
agent/triage.ts planIssue / planControl / planSweep - the decision engine
agent/outreach.ts compose, send, log, schedule the follow-up ladder
agent/inbound.ts interpret a human reply, then mutate through the reducers
agent/tools.ts the only surface the LLM can use to change anything
agent/chat.ts console conversation (LLM tool loop, deterministic fallback)
agent/digest.ts weekly executive digest
agent/runtime.ts one bounded pass of the agent, recorded as an AgentRun
integrations/ Slack (Web API), Teams (Bot Framework + Graph), email (HTTP relay)
store.ts StoreDriver interface + JSON file driver
scheduler.ts interval sweep, evidence pass, weekly digest
server.ts Fastify routes, webhook verification, static console
apps/app Vite + React console. Reads one bootstrap snapshot, derives everything.
apps/web Next.js marketing site.
Why the console talks to a single bootstrap endpoint
GET /api/bootstrap returns the whole workspace plus a computed posture summary. Every panel derives from that one
snapshot, which means the risk score and the issue list can never disagree, and a mutation has exactly one cache key
to invalidate. The workspace is a few hundred kilobytes for a mid-size company, so this is far cheaper than the
consistency bugs of a dozen endpoints.
When the agent is unreachable, apps/app/src/lib/api.ts swaps in a client that runs the same reducers in the
browser over the bundled tenant. Closing an issue, filing evidence or accepting risk behaves identically and the
activity feed stays coherent. That is what lets the console be demoed before any infrastructure exists.
Escalation, precisely
The most subtle part of the product, and the one that took the most iteration:
- Escalation level is computed from unanswered outreach since the last human reply. A person who answers is never punished by having the ladder advance underneath them.
- A human reply puts the ball back in their court: the next contact is a nudge to them, not a climb upward.
- An owner who has never been contacted is asked before anyone is escalated past them.
- Blocked controls escalate to the owner's manager, because a blocked control stalls an entire framework.
- Windows shrink with severity and with each level: critical 4h, high 12h, medium 24h, low 72h, divided by the level.
- Four levels maximum, ending at a named human with two honest options: fund the fix, or record a risk acceptance with an expiry date.
The test suite asserts all five behaviours, including the regression that a reply must reset the ladder rather than advance it.
Evidence reuse
One control's evidence frequently satisfies another, and asking a human twice for the same artifact is the fastest
way to lose their cooperation. Overlap between frameworks is modelled explicitly: SOC 2 CC6.7 and GDPR Art. 32
map to the same encryption evidence, and CC1.4 and ISO A.6.3 share the training completion report. When an
issue closes, updateIssueStatus credits every linked control with a confidence bump, which is how closing an issue
visibly moves the readiness number.
Extending it
Add a framework. Append ControlDefinition entries to packages/shared/src/domain/catalog.ts with an id,
framework-native code, owner role, automation level, evidence hint and the plain-language question the agent asks.
Add the FrameworkId to taxonomy.ts. Nothing else changes: readiness, the console, the digest and the agent's
evidence chasing all pick it up.
Add a directory or evidence source. Implement MessagingAdapter (integrations/types.ts) for messaging, or
write a collector that produces Evidence and files it through addEvidence. Scanners, MDM, IdP and cloud posture
tools all reduce to "produce evidence, open an issue when something is wrong".
Swap the store. Implement StoreDriver (load, save, describe). Postgres fits in about 200 lines; the agent
and the HTTP layer never touch persistence directly.
Add a role. Add it to ROLES, to ISSUE_ESCALATION_LADDER for the categories it owns, and to
CONTROL_OWNER_ROLE for the domains it covers. Routing, the people page and the escalation math follow.
Multi-tenancy
One process, many customers. The design goal was to make cross-tenant leakage a compile error rather than a code-review question, so:
- No ambient current tenant.
AgentContextis passed explicitly into every agent function. There is no module-level "current workspace" to accidentally read, and no repository function that reads a tenant implicitly. - Every handler resolves its own context through
contextFor, which goes through the tenant manager's cache. A route physically cannot read another workspace's state without asking for it by slug. - The key decides the workspace. A subdomain or
X-Tenantheader is an assertion that is enforced against the key; the absence of either is not an assertion, so a workspace key simply means "my own workspace". Getting this backwards is subtle: the first version defaulted to the platform workspace and rejected every customer key that did not send a header. The test suite now pins both directions. - Inbound events are attributed before they are trusted. Slack
team_idand the Azure tenant id select the workspace, then the signature is verified with that workspace's secret. An unattributable event gets a 202 and is dropped. Falling back to "the only tenant" happens only when there is exactly one, which is the self-hosted case.
Registry vs workspace: why the split
| Tenant registry | Workspace document | |
|---|---|---|
| Holds | slugs, plan, limits, credentials, usage, API key hashes | people, issues, controls, activity, tasks |
| Size | kilobytes for thousands of customers | hundreds of kilobytes per customer |
| Read | on every request | on demand, cached |
| Write | rarely | constantly |
Keeping them apart is what makes the cache in tenancy/manager.ts possible: the hot,
per-customer document is loaded lazily into an LRU (TENANT_CACHE_SIZE, default 25) and
flushed to disk before eviction, so a cold workspace never loses its last mutation. The
small registry stays resident because every request needs it for authorization.
Concurrent loads of the same cold workspace share one in-flight promise, so a burst of traffic on a freshly evicted tenant causes exactly one disk read.
Credentials
Customers can bring their own Slack app, Azure Bot and model provider. Those values are
sealed with AES-256-GCM under TENANT_SECRET_KEY and are never returned by the API — the
console sees hasBotToken: true, never the token. Three behaviours worth stating:
- production without
TENANT_SECRET_KEYrefuses to store credentials rather than writing them in plaintext; - non-production generates an ephemeral key and warns loudly, so local development is not blocked by ceremony;
unsealpasses through values that were configured rather than sealed, so a token can be injected by hand in an emergency without re-encrypting anything.
Webhook verification and the raw body
Slack signs the exact bytes it sent. Fastify's default JSON parser throws those bytes away, so the server installs its own parser that keeps the raw string and parses it itself. Re-serialising a parsed object produces a different byte sequence and would silently break every signature check — a bug that only shows up in production, against real Slack.
Two starting states
Workspace.open decides between them, and the distinction is a safety property, not a convenience:
| Condition | Result |
|---|---|
| Store has data | Load it |
Store empty, SEED_ON_FIRST_BOOT true (default outside production) |
The demo tenant — Helios Robotics, 17 issues, 14 past SLA |
Store empty, SEED_ON_FIRST_BOOT false (default in production) |
A clean tenant — your org, the full catalog with an owner per control, no issues, suggest only autonomy, one onboarding message |
The clean tenant is why packages/shared/src/state/fresh.ts exists. Without it, a production boot on an empty volume
would either crash or serve another company's incidents, and both are unacceptable in a security product. The
starter org chart deliberately covers every role the catalog references, so a new customer does not open the console
to a wall of "nobody owns this".
Autonomy and the release queue
In suggest_only — the default for a new tenant — every message is composed and written to the workspace as a
queued outreach record, then listed in the console's queue for a human to release. That has two consequences worth
being precise about:
- Queued asks never advance the escalation ladder. Only delivered outreach (sent, no response, failed) counts toward escalation levels. Escalating past a message that was never delivered would be meaningless.
- Queued asks do suppress duplicates. The planner treats an unreleased draft as "in flight", so repeated sweeps do not stack copies of the same message in the queue. Without this, a suggest-only tenant grows one duplicate draft per issue per sweep — which is exactly the bug the test suite now guards against.
Concurrency between processes
The deployment shape puts two processes against the same documents: the API service (webhooks, console, chat) and the sweep job (the agent's own runs). Both mutate the tenant registry and the workspace documents.
The first implementation lost data because of it. Each process wrote its whole in-memory copy, so whichever wrote last
won. In a live GCP deployment the bucket's version history showed platform.json alternating between 975 and 1803
bytes — one process erasing the other's tenant — and a freshly provisioned customer could not authenticate because
their record had been rolled back out of existence.
Three mechanisms fix it, and they are asserted in services/agent/src/__tests__/concurrency.test.ts:
- Version-checked writes.
ObjectStore.puttakesifGenerationMatch; on GCS that is the object's generation, and the file and in-memory drivers use the document's own revision counter (rev:<n>) as the token. Creating a document is conditional on it not existing ('0'), so two concurrent provisions cannot both "succeed" and leave one of them orphaned. - Reducer replay.
Workspacekeeps the reducers that produced the pending changes. On a conflict it reloads the current document and re-applies them on top, so the other writer's work is preserved and ours is layered on, rather than one of the two being chosen. - Increments, not replacements. Usage counters are folded in as deltas (
recordUsage) and coalesced for a few seconds. A process that wrote its own stale totals would reset another process's numbers; a delta cannot.
Two smaller consequences worth knowing:
getByApiKeytreats a miss as a reason to re-read the registry before rejecting a key, so a workspace created by another process cannot be locked out by a stale cache. Reads otherwise stay in-memory, because they are on the hot path and this document is small.- The in-memory drivers clone on read and write. Returning the stored object meant the "persisted" document aliased the caller's working copy, which made every version check look like a conflict — a bug the concurrency tests caught immediately.
Operational lesson
Shipping a fix to the write path is not enough on its own: an old revision that is still serving keeps writing with
the old semantics. After deploying the concurrency fix, the stale revisions had to be deleted
(gcloud run revisions delete) before the registry stopped oscillating. This is in the deployment README's
troubleshooting table for that reason.
Known limits
- No per-user accounts inside a workspace. The workspace API key is the credential, so there is no SSO, no RBAC and no per-operator attribution inside a customer's workspace; console actions are attributed to the workspace's security owner. Adding user auth is the natural next layer, and the auth hook is the single place it would plug into.
- The registry is a JSON document, not a database. Fine to thousands of workspaces on one node; Postgres is the right next step and the driver interface is already keyed by tenant.
- The service runs a single instance. Workspace documents are cached in memory per instance, so a second
instance would serve stale reads. Writes merge correctly (see above), but reads would not be fresh. Raising
max_instance_countneeds a shared cache, not a Terraform change. - Sweep cadence is per tenant, but the trigger is global. Cloud Scheduler wakes the job every 15 minutes and each workspace's own interval decides whether it is actually swept. That is correct but means the finest granularity is the schedule, not the tenant's interval — a workspace with a 15-minute cadence and a 15-minute trigger has no slack.
- Teams cannot initiate a chat with someone who has never messaged the bot (a Bot Framework constraint). The agent falls back to Slack or email for that person until a conversation reference exists.
- Teams inbound authentication uses a shared secret in this build rather than full Bot Framework JWT validation. Slack inbound is fully verified (HMAC signature plus replay window).
- No object storage for uploaded evidence:
Evidence.urlis a link, not an upload. Attachments are the next increment. - Scanner connectors are seeded, not live. The AWS/GitHub/Okta integrations carry realistic findings and modeled capabilities; the API clients are the next integration task.