Skip to content

Security and Governance

  • Least privilege by agent: Sentinel and Detective only need read-scoped MCP access (metrics/logs/traces). Producer needs write access to Incidents. Responder needs write access to Alerting/Annotations — but every Responder write for a high risk action is gated behind request_human_approval (see agents.md), which cannot be bypassed by the LLM.
  • Grafana credential scope note: every agent shares one self-hosted mcp-grafana server and, through it, one Grafana service-account token's write scope, with the approval gate as the compensating control (see above). For stricter separation, deploy a second mcp-grafana instance with a read-only service account for Sentinel/Detective and point only Producer/Responder at a write-scoped one.
  • mcp-grafana is reachable, but token-gated: the self-hosted MCP server (infra/scripts/deploy-mcp-grafana.sh) runs as its own public Cloud Run service, since Cloud Run's IAM-based service-to-service auth doesn't fit the MCP client's one-shot header model. It's gated instead by a caller-auth token (--server-auth-token / MCP_GRAFANA_SERVER_TOKEN) that mcp-grafana itself supports for exactly this case — generated automatically, stored in Secret Manager, and never exposed outside the two Cloud Run services that need it.
  • Auditability: every agent action is persisted as a Firestore agent_events document (see agents.md) and mirrored into a Grafana Incident timeline via add_activity_to_incident / create_annotation — the incident record in Grafana and the Firestore incident document should always agree.
  • Secrets: Grafana service-account tokens, the mcp-grafana caller-auth token, and any Gemini API key (if used instead of Vertex AI) live in Secret Manager, never in source or client-side code.
  • Dedicated service accounts: on Cloud Run, the backend, frontend, and self-hosted mcp-grafana server each run as their own service account (premiere-backend, premiere-frontend, premiere-mcp-grafana) instead of the shared, broadly-privileged Compute Engine default service account. The backend's account is granted only roles/aiplatform.user (Gemini via Vertex AI), roles/secretmanager.secretAccessor, roles/cloudsql.client, roles/datastore.user (Firestore), and logging/monitoring writers; the frontend's and mcp-grafana's accounts get only logging/monitoring writers plus (for mcp-grafana) roles/secretmanager.secretAccessor for its own two secrets. See infra/scripts/00-setup.sh and infra/scripts/README.md.
  • Gemini access: the backend authenticates to Gemini via Vertex AI using its own service account's Application Default Credentials by default -- no API key is generated, transmitted, or stored anywhere in this path. A Gemini Developer API key is only used (and only stored, in Secret Manager) if explicitly opted into via GOOGLE_API_KEY.
  • Access control: JWT-based auth (app/auth.py) with a viewer < operator < admin role hierarchy. The control room and history views are intentionally unauthenticated (read-only), but approving/rejecting a remediation, injecting a demo anomaly, and all user/workspace management require signing in as operator/admin respectively -- enforced server-side via require_role() FastAPI dependencies, not just hidden in the UI. Passwords are stored as salted PBKDF2-HMAC-SHA256 hashes (app/auth.py:hash_password), never plaintext or reversibly encrypted.
  • Bootstrap admin: one admin account is created on first startup from ADMIN_EMAIL/ADMIN_PASSWORD; if ADMIN_PASSWORD is left unset, a random one is generated and logged once at WARNING level rather than a fixed default shipping in source -- rotate it via the admin API once you've signed in with it.
  • JWT secret: JWT_SECRET should be set explicitly for any deployment meant to survive a restart or run multiple instances; left unset, a random secret is generated per-process, which just means issued tokens stop validating across a restart (fails safe, doesn't fail open).
  • Audit trail: every sensitive action (login, approve/reject, inject-anomaly, chaos, user/workspace management) is written to AuditLogRow with the real actor's email, visible at GET /api/audit-log and the frontend's /audit page. On an approved remediation, the persisted approved_by is the authenticated actor who called the approve endpoint (see approval.take_resolved_by), not whatever value an LLM's structured output happened to fill in for that field -- an LLM has no way to actually know who clicked approve, so it is never trusted as the source of truth for that.