Architecture
TensorCost is a multi-tenant SaaS control plane for AI cost governance. The platform attributes every dollar across GPU fleets, managed inference, and agent workloads, surfaces recommendations, and (with customer opt-in) takes enforcement actions. This page describes the system at the level a platform-engineering reviewer cares about: what runs where, how tenants are isolated, how data flows in, and how events flow out.High-level topology
Backend — gateway plus 14 microservices
Every service is NestJS + TypeScript, communicates over gRPC (proto contracts inpackages/proto), persists to its own Postgres schema, and emits events to Redis-backed pub/sub.
Services run on AWS Fargate behind an Application Load Balancer (REST + WebSocket) and a Network Load Balancer (gRPC). Postgres is RDS Multi-AZ; Redis is ElastiCache cluster mode with replicas.
Frontend — shell plus 14 microfrontends
Module Federation (Vite) loads each microfrontend on demand into a single shell. Each MF owns a domain.
Stack: React 18, MUI 5, Vite, Module Federation, RTK Query (migrating to TanStack Query). Real-time updates via
socket.io joined to a tenant-scoped room.
Shared packages
Twenty-one packages inpackages/ carry the cross-cutting concerns. The ones an integrator most often interacts with:
@tensorcost/proto— gRPC contracts.@tensorcost/agent-sdk— IMDS auto-detect, HMAC signing, retry/backoff.@tensorcost/db-utils— Sequelize defaults, RLS hook (runAsBypass), tenant-scoped repository.@tensorcost/rbac— role + scope checks; the same primitive guards REST, gRPC, and MCP tools.@tensorcost/audit-trail— append-only audit ledger writes.@tensorcost/observability— OpenTelemetry init, structured logger, trace-correlated logs.@tensorcost/feature-flags— LaunchDarkly client +useFeature()for the frontend.@tensorcost/jobs— scheduled-job framework with per-tenant fairness.@tensorcost-internal/mcp-framework— MCP tool-registration framework.
Multi-tenancy and Row-Level Security
Every customer is atenant. Every row in every table that holds tenant data carries a tenant_id and is protected by Postgres Row Level Security.
- Current rollout: 21 of 50 target tables enabled (cost, enforcement, integration, gpu, identity, monitoring schemas in flight).
- Bypass discipline: the only path that bypasses RLS is
runAsBypass(tenantId, fn)from@tensorcost/db-utils, which setsapp.tenant_idfor the duration of the transaction. gRPC handlers verify the agent’stenantIdvia HMAC before any bypass call. - Linting: any
sequelize.query()that touches tenant data must be wrapped in a transaction orrunAsBypass. ARUN-AS-BYPASS-LINTrule is on the way; until then, code review enforces the discipline.
Agent ingest — HTTPS by default, gRPC + HMAC + IMDS as an opt-in transport
The unified GPU agent runs in customer environments. By default it reports over plain HTTPS using an API key. Customers who opt into the gRPC transport get a long-lived bidirectional stream on TCP/50051 instead, with a stronger per-message handshake:- Agent boots, auto-detects EC2 metadata via IMDSv2 where available (with a
±300sskew guard and a Redis-backed nonce replay defense). - First gRPC message is
AgentHellowith tenant ID, agent ID, key ID, nonce,timestamp_unix, and an HMAC-SHA256 signature over those fields. gpu-serviceverifies the HMAC withtimingSafeEqual, rejects nonce reuse viaSETNX nonce:{tenantId}:{keyId}:{hex(nonce)}with 600s TTL.- Verified
tenantIdis bound to the stream session; every subsequent DB write runs insiderunAsBypasswith thattenantId. - The NLB listener accepts TLS only (wildcard ACM cert); plaintext gRPC is rejected before reaching the handler.
Managed-inference ingest — read-only, redacted
Bedrock and friends never see an outbound call from us. The data flows in over four read-only paths:
Raw prompts and responses are never stored. Hashes only. Customers concerned about content can disable request-metadata capture entirely and still get model/cost/latency attribution. Raw logs live in customer-owned S3 buckets; we read with read-only IAM scoped to the tenant.
Real-time events
socket.io runs on the same host as the console app, routed to the gateway over a dedicated /rt path so the upgrade request doesn’t fall through to the static frontend. Clients are joined to a tenant-scoped room at handshake; cross-tenant events are unreachable by construction.
The full event catalog and reliability model is in real-time events.
Observability
OpenTelemetry across the backend (Node.js auto-instrumentation) and the agent (Pythongrpc + requests instrumentation). Logs are JSON, tagged with trace_id / span_id. Default deployment shape is the AWS Distro for OpenTelemetry collector as an ECS sidecar forwarding to X-Ray + CloudWatch, but any OTLP-compatible backend works. See observability.
MCP — agents query TensorCost
A built-in MCP server exposes scope-guarded tools for cost queries, fleet inspection, workload attribution, inference analytics, and (where the tenant grants the scope) write tools. Claude Desktop, internal agents, and partner integrations all consume the same surface. Tool grants are RBAC-checked at every call.Where to read more
- Agent installation — the customer-side install flow.
- Bedrock integration — the lead managed-inference adapter.
- SOC 2 readiness — the security and compliance posture.
- API reference — gateway REST surface.