The platform

One system. Three layers. Five specialist agents.

Cloud telemetry flows in, gets reasoned over by specialist agents sharing one knowledge base, and comes out the other side as either a proposal or — once policy and a human allow it — a safely executed change.

01 — Architecture

How data becomes a decision.

Three layers, one continuous loop: observe, reason, act.

Layer 1 — Cloud environment
AWSAzureGCPKubernetes Telemetry & logsConfig & IaC stateEvents
↓
Layer 2 — Agentic orchestration (shared RAG knowledge base)
FinOps
SRE / SysOps
SecOps
Compliance
Knowledge
↓
Layer 3 — Governance & actions
Policy-as-Code guardrailsIaC integration Event-driven automationHuman approval gateSafe execution
Every recommendation is traceable to a source. Every high-impact action passes through the same approval path.

That's the design principle behind the whole platform — trusted by the engineers who have to answer for what it does, not just what it flags.

02 — Where this stands today

Two agents operational. Proven architecture. Measurable quality.

Not slides, not wireframes — working code. 334 tests pass with zero failures. 31 MCP tools exposed. Architecture scored 80/100 against PROOF, Microsoft multi-agent, and Gartner production-readiness standards.

334
tests passing, zero failures
31
MCP tools across 2 agents
80/100
architecture evaluation score
Architecture evaluation — scored against industry standards
Agent Intelligence 20/25 Tool Integration 23/25 Safety & Governance 24/25 Observability 16/25 Production 17/25
03 — The agents

Five specialists. One brain.

Each agent owns a domain an SME would otherwise need a dedicated hire for. All of them read from the same knowledge base, so a cost decision knows about an open vulnerability, and an incident knows what changed in the last deploy.

FinOps agent · operational

Cost & rightsizing

255 tests · 23 MCP tools · 16 detectors
  • Detects idle compute, orphaned resources, GPU waste, data-transfer egress
  • 23% avg GPU utilization — flags AI workloads bleeding cash
  • Terraform sync pipeline: discover 28 resource types (AWS 14 + Azure 7 + GCP 7) → import → drift detect → human-gated approval. Provider threaded end-to-end — Azure/GCP run the same full loop as AWS.
  • Existing-IaC mode adds plan-based attribute drift: terraform plan in your own workspace catches changes made outside TF
  • Multi-sub-agent architecture: Scanner → Terraform → Remediation → Reporter
  • Terraform apply NEVER auto-executed — all mutations through human approval queue
SRE / SysOps agent · operational

Health & incidents

79 tests · 8 MCP tools · 16 EKS/AKS tools
  • 16 ReadTool/ActionTool implementations for AWS EKS and Azure AKS
  • Z-score anomaly detection with 6 pattern-matched root causes (OOM, CrashLoop, etc.)
  • Deployment validation, RCA with confidence scores, runbook-driven remediation
  • Policy engine: environment allowlisting + per-resource rate limiting
SecOps agent

Posture & vulnerabilities

Checks the fence line, all the time.
  • Checks posture against CIS, SOC 2, ISO 27001, NIST
  • Detects vulnerabilities and misconfigurations
  • Monitors IAM risk across accounts and roles
  • Prioritizes remediation by actual exposure
Compliance agent

Audit & evidence

Architecture ready — shares SecOps findings.
  • Continuously maps evidence to compliance controls
  • Cuts the manual effort auditors used to demand
  • Shares findings directly with the SecOps agent
  • Turns audit season into a non-event
Knowledge agent

Natural language over your whole estate

The interface to everything the other four know.

It learns your architecture and operational history, so anyone on the team can ask a plain-language question and get a grounded answer — not a dashboard to go interpret themselves.

Why did costs increase last week?
Which security findings need attention right now?
What changed before last night's incident?
04 — Safety by design

96% of enterprises deploy agents. Only 12% can govern them.

OutSystems surveyed 1,900 IT leaders in 2026. PwC reports 78% plan to increase agent autonomy — but only 21% have governance models ready. CloudSentri's governance is the competitive moat.

risk tiers
Four tiers: READ_ONLY → LOW → HIGH → DESTRUCTIVE. Every action is classified before execution.
approval queue
SQLite-backed human approval gate. Persists across restarts. Never auto-executes HIGH or DESTRUCTIVE actions.
audit trail
Full audit log: who proposed, who checked policy, who approved, when. Every remediation action is traceable.
never auto-apply
Terraform apply is NEVER auto-executed. The tool generates shell commands — a human runs them.
read/write split
Only the RemediationAgent has write capabilities. All other agents are read-only — enforced at the architecture level.
rate limiting
Per-resource action caps prevent flapping loops. SysOps policy engine enforces environment allowlisting.
05 — Under the hood

Built on principles the best teams converge on.

The 2026 State of AI Agents industry report confirms: the dominant production pattern is single tool-use with human review — exactly CloudSentri's propose→policy-check→approve→generate pipeline.

deterministic core + agentic edge
16 detectors are pure code — the agent only reads their output and never invents a number
multi-sub-agent orchestration
FinOpsOrchestrator coordinates Scanner → Terraform → Remediation → Reporter via async event bus
terraform sync pipeline
Discover 28 resource types (AWS 14 + Azure 7 + GCP 7) → import blocks → plan -generate-config-out → drift detection → human-gated approval. Provider threaded through the whole loop, plus plan-based attribute drift for existing Terraform workspaces.
MCP-native tools
31 tools exposed over stdio protocol. Compatible with any MCP client.
async-safe job queue
ThreadPoolExecutor for long ops (terraform plan, boto3 scans) — never freezes the MCP event loop
4 LLM providers
Direct API, AWS Bedrock, GCP Vertex AI, Azure — pick for data residency, not capability
90+ enrichment metrics
AWS/Azure/GCP/K8s metrics matching native optimization tools: Compute Optimizer, Advisor, Recommender, VPA
reliability-aware cost
CloudSentri uniquely combines FinOps + SRE/SysOps — the category Komodor's 2026 analysis says SRE teams co-sign
See it applied to your team

Explore solutions by role.