v0.3 — policy-gated remediation, causal root-cause, MCP integrations, learned auto-resolve

Watches your services. Reduces your logs. Fixes what it's sure about.

Point Sentryops at any backend or deployment — a plain URL or a Kubernetes namespace — and it discovers what's actually running behind it, probes it Postman-style, tracks its health continuously, and turns a wall of logs into a root cause and a fix. A dependency graph ranks the real cause over its symptoms, and an approved fix executes for real behind a policy gate. Known-safe errors resolve themselves, and the list of "known-safe" keeps growing on its own. Everything else waits for you, with the full evidence trail attached.

Tier-1 resolves itself. Tier-2 waits for a person. Every decision — either way — is written to an audit log you can read.

sentryops query — "is payments healthy?"

Why we built this

Every on-call rotation hits the same three walls. Sentryops exists to knock all three down at once.

You don't actually know what's running

A "health check" is usually one URL, one convention, one guess. Real deployments are dozens of endpoints and pods, and nobody's kept the list current.

The logs are the crime scene, but there's too much of them

The answer is in there somewhere, in among 100,000 lines nobody has time to read. By the time someone greps it out, the incident's been open for an hour.

Every alert pages a human, even the boring ones

A connection reset that would've cleared itself on retry still wakes someone up at 3am, because nothing in the pipeline is trusted to tell "boring" from "dangerous."

Six steps, every query, no exceptions

The same pipeline runs whether the answer turns out to be "nothing's wrong" or "restart the payments pod" — nothing about how carefully it checks depends on how it thinks the story ends.

01

Resolve

Matches your query to the specific service(s) it's actually about.

02

Plan

Drafts an investigation: which checks, in what order, against which service.

03

Investigate

Runs each step for real — service registry, tickets, runbooks, live endpoint probes, Kubernetes pods, reduced & diagnosed logs.

04

Reason

Weighs the evidence and drafts a decision, with its reasoning attached.

05

Safety guard

Non-LLM backstops double-check the target is unambiguous, and that a known-safe error signature — not just a hopeful guess — is what's letting anything skip human approval.

06

Act or ask

Executes immediately, or opens an approval and waits.

Watch it decide

Pick a real query and step through the same six-stage pipeline every request runs — Resolve, Plan, Investigate, Reason, Safety guard, Act or ask. No account, no login, nothing to install.

This replay uses real historical decisions, not a live cluster — every field below is the real shape of an actual response, just broken into steps you control.

Real evidence, replayed — not a live cluster. Point Sentryops at your own deployment to watch it decide for real.

Create a free account

Everything it actually does

Not a chatbot bolted onto your infra — a planner with real tools, a real memory, and a real opinion about what's safe to touch.

CapabilityWhat it doesWhy it matters
Multi-endpoint discovery Register a small Postman-style collection of paths (or let it fall back to conventional health paths) and bulk-probe every one, concurrently, with a real pass/fail per endpoint. One URL was never the whole picture. Now the whole picture is one click.
Real Kubernetes pod discovery Give it a namespace and it asks kubectl directly — which pods exist, their phase, container readiness, restart counts. Ground truth from the cluster, not a cached label someone forgot to update.
Continuous health tracking A background scheduler re-checks every registered service on an interval and keeps a full history — and opens a ticket automatically the moment something goes unhealthy. You find out from Sentryops, not from a customer.
Deterministic log reduction Paste or upload raw logs — thousands of lines is fine. They're deduplicated and clustered by normalized signature, then ranked by severity and frequency, before anything touches a model. 100,000 lines become a dozen clusters you can actually read, in milliseconds, for free.
LLM root-cause diagnosis The reduced clusters are handed to a language model that returns a root cause, the exact implicated pattern, and concrete remediation steps. Not "here are your logs" — "here's what's wrong and here's how to fix it."
Curated + learned auto-resolve A small, explicit whitelist of known-transient signatures (connection resets, timeouts, DNS blips) resolves automatically — plus a per-deployment learned memory that starts empty and promotes a new signature only after an unbroken run of human approvals. One rejection resets it to zero. The boring 3am pages stop paging anyone, and the list of "boring" keeps growing on its own.
Causal root-cause (Production World Model) Declare which services depend on each other, and a persisted dependency graph ranks which currently-unhealthy service actually explains the others — not just which one happened to page first. Diagnosis stops being per-service guesswork and starts being graph-aware.
Real, policy-gated remediation An approved restart, scale, or rollback issues a real kubectl mutation — gated end-to-end by an OPA-style policy engine (namespace allowlist, denied services, a dry-run switch) that's independent of both the LLM and the approving human. Closes the loop from "here's the root cause" to "here's the fix, executed" — with hard boundaries a buyer can actually audit.
Tiered human approval Anything ambiguous, multi-service, or explicitly requested by name always stops for a person to approve or reject. Autonomy where it's earned, a human in the loop everywhere else.
MCP integrations A protocol-correct Model Context Protocol client — not a hardcoded wrapper for one vendor. Point it at the Slack, PagerDuty, GitHub, or Datadog/Prometheus MCP server (or any other) and an approval being opened or decided notifies it automatically. Plugs into the tools you already run instead of asking you to adopt new ones.
Full audit trail Every query, plan, step result, approval, policy decision, MCP notification, and executed action is logged with a correlation id you can trace end to end. Nothing the agent does is a black box after the fact.
Benchmarked accuracy A benchmark harness runs the real auto-resolve and root-cause engines against a labeled synthetic corpus and reports precision/recall/F1 — with a hard safety gate that fails the build if a dangerous scenario is ever auto-resolved. An eval you can rerun and audit, not a number in a slide deck.
Embed it or call it An embeddable Rust SDK with zero HTTP overhead, a dependency-free Go client + agentctl CLI, or plain REST from anywhere. Fits into what you've already built instead of demanding a rewrite.

Sentryops vs. doing it by hand

Same underlying problems. Very different Tuesday night.

Without SentryopsWith Sentryops
Health checksOne URL, pinged manually, trusted blindlyMulti-endpoint, Postman-style, plus real Kubernetes pod discovery
Log triageGrep through tens of thousands of lines by handDeterministic clustering + LLM root cause in seconds
Root causeGuess from whichever service happened to page firstA dependency graph ranks the actual cause over its symptoms
RemediationEvery alert pages a human, alwaysKnown-safe errors auto-resolve — and approved fixes execute for real, gated by policy
ApprovalTrust the on-call engineer's judgment aloneNon-LLM safety guards force review on anything ambiguous
NotificationsCustom webhook code per tool you useAny MCP server (Slack, PagerDuty, GitHub) wired in via config
IntegrationBespoke scripts glued to one stackEmbeddable Rust SDK, Go client/CLI, or plain REST
HistoryTribal knowledge and Slack scrollbackA queryable, correlation-id'd audit log of every decision

Talk to the tools you already run

A real Model Context Protocol client: JSON‑RPC over stdio, the actual handshake, not a wrapper hardcoded to one vendor's API. Point it at any MCP server and an approval opening or getting decided reaches it automatically.

Slack

Post to a channel

An approval request or decision lands wherever your team already watches for incidents.

PagerDuty

Fire an on‑call trigger

The same event opens or resolves a page without a separate integration to maintain.

GitHub

Drop it in an issue

Route the decision into a repo your engineers already have open.

ServiceNow

Update a ticket

Keep the change record in sync with what the agent actually did.

mcp.json
{
  "servers": {
    "slack": {
      "command": "npx",
      "args": ["-y", "@modelcontextprotocol/server-slack"],
      "env": { "SLACK_BOT_TOKEN": "xoxb-..." }
    }
  },
  "on_approval_requested": {
    "server": "slack",
    "tool": "post_message",
    "extra_arguments": { "channel": "#incidents" }
  }
}

No file at all means every hook is a silent no‑op, the same secure default the policy engine uses. Datadog and Prometheus fit the same protocol; pulling live metrics into an investigation as evidence, not just a notification target, is the natural next step and isn't built yet.

Measured, not asserted

A benchmark harness (agent-bench) runs the exact deterministic decision engines agent-server calls on a live request against a labeled, disclosed-synthetic corpus — not a simulation, not marketing math.

Accuracy
100%
Precision
100%
Recall
100%
F1 score
1.000
Dangerous false-positives
0 / 20
World Model root cause
10/10

By failure class, 27 scenarios

ClassScenariosShould auto-resolveResult
Transient network5Yes5/5 correct
Transient DNS2Yes2/2 correct
Auth & permission4No4/4 correct
Data integrity3No3/3 correct
Capacity exhaustion3No3/3 correct
Deploy regression3No3/3 correct
Security incident3No3/3 correct
Ambiguous / novel4No4/4 correct

Plus 10 dependency-graph topologies for root-cause ranking: fan-ins, chains, a diamond, a symmetric cycle, and a disconnected node that must never get blamed just for being unhealthy nearby. This deliberately does not publish an MTTR-reduction percentage. That needs real incident timelines from real usage, which is what the pilot below is for. Results are versioned in agent-rs/bench-results/ so accuracy stays visible over time, and the corpus itself is in source control to audit, not a black box.

terminal
# from agent-rs/
cargo run -p agent-bench

Packaging & pricing

An open SDK for teams who'd rather run it themselves, and a hosted product for teams who'd rather not, priced per monitored service instead of per host or per gigabyte.

Two ways to take it

Free & open

Embedded SDK

Run agent-server yourself, or embed agent-sdk directly into an existing Rust service with zero HTTP overhead. The full decision engine ships with it: policy, World Model, and learned auto-resolve are never gated behind a license check.

  • MIT-licensed, source in this repo
  • No billing relationship at all
  • A Go client + agentctl CLI, or plain REST, if Rust isn't your stack
Paid

Hosted, isolated per customer

We run agent-server and the dashboard for you — a dedicated deployment per customer, its own database and config, not a shared multi-tenant platform. Managed upgrades, backups, and uptime, with the full audit trail hosted alongside it.

  • The dashboard, accounts, and audit trail, hosted
  • Your data never shares a database with anyone else's
  • Same decision engine as the open SDK — nothing held back for paying customers

Pricing, per monitored service

Pilot
$0
up to 3 services, 60–90 days
  • Full feature set, no cap on approvals
  • In exchange for a case study + reference
  • Direct input into the roadmap
Team
$30/service/mo
11–50 services
  • Unlimited MCP integrations
  • SSO, role-based dashboard access
  • Priority support, shared Slack channel
Enterprise
Custom
50+ services, or self-hosted & supported
  • Dedicated deployment, custom SLA
  • SOC2 report access (see below)
  • Custom policy review with your security team

Sized against what teams already pay: Datadog charges $15/host/month for infrastructure monitoring alone and $31/host/month for APM, and a typical mid-market observability stack (Datadog + PagerDuty + a handful of smaller tools) runs about $10,300/month. Flat per-service pricing here is deliberate — no per-GB or per-event surprises on top of a number you already agreed to.

The pilot

Step 1

Recruit 2–3 teams

Real on-call pain, already running Kubernetes.

Step 2

Free for 60–90 days

Full feature set, dedicated deployment, direct access to the team building it.

Step 3

Track real approvals

The same audit trail already built — against real incidents this time.

Step 4

Publish the real number

A real MTTR/auto-resolve-accuracy result, with permission, replacing "stays an open item."

SOC2: what's true today

Already true, in the codebase

  • Every decision (query, policy verdict, MCP notification, executed action) writes to a correlation-id'd audit log
  • The policy engine is default-deny: nothing mutates without an explicit namespace allowlist
  • Secrets are gitignored and never echoed back by any status route
  • Session auth is server-side, bcrypt-hashed, httpOnly cookies

Still needed for certification

  • A named security policy, incident response plan, and vendor review process
  • Access reviews and offboarding procedures once there's a team, not just a founder
  • An actual Type I audit engagement, timed to just after the first paid contracts
  • Encryption-at-rest attestation for wherever customer databases end up hosted

Test it live, against a real deployment

The live-test dashboard is a real client of the API above — register a backend, bulk-probe it, discover its pods, paste in logs, and ask it free-text questions. It sits behind a free account so we know who's kicking the tires (and so your test data isn't world-readable).

Two ways to integrate

Embed it directly in a Rust service with zero HTTP overhead, or call a running instance from anywhere with the Go SDK — or just speak REST.

main.rs — embedded, no server
use agent_sdk::{Agent, AgentConfig};

#[tokio::main]
async fn main() -> anyhow::Result<()> {
    let agent = Agent::bootstrap(
        AgentConfig::from_env()
    ).await?;

    let res = agent
        .submit_query("is payments healthy?")
        .await?;

    println!("{:?}", res.tier);
    Ok(())
}
main.go — talking to a running instance
import agent "github.com/devops-helpdesk/agent-go"

client := agent.NewClient("http://localhost:8000")

resp, err := client.SubmitQuery(ctx,
    "is payments healthy?")

if resp.ApprovalID != nil {
    approver := "alice"
    client.DecideApproval(ctx,
        *resp.ApprovalID, "approve", &approver)
}

What actually changed under the hood

This started as a rewrite of an existing Python service — same routes, same JSON, same decision logic. What's different is underneath, and what's been added since.

BeforeNow
RuntimePython interpreter, per-request overheadCompiled Rust binary
Deployment artifactApp code + a venv of dependenciesOne static binary
Type safetyRuntime errors on malformed dynamic dictsChecked at compile time
Health checksSingle URL, GET-onlyPostman-style multi-endpoint + real k8s pod discovery
Logs"No log aggregation source connected."Deterministic reduction + LLM diagnosis + auto-resolve
IntegrationHTTP onlyEmbeddable Rust SDK, Go SDK/CLI, or REST

Up and running in four commands

The full manual — configuration, every route, deployment options — lives on the docs page. Here's the short version.

terminal
# 1. set your model provider key
export GROQ_API_KEY=your-key-here

# 2. build and run the server
cd agent-rs && cargo run -p agent-server

# 3. from anywhere else, talk to it
cd agent-go && go run ./cmd/agentctl query "is payments healthy?"

# 4. approve or reject anything it flags
go run ./cmd/agentctl approvals decide 1 --decision approve

Tell us what's missing

Tried the demo, read the docs, poked at the API — good or bad, we want to hear it. No account needed.

Get in touch

Questions about deploying this, integrating it, or extending it — happy to talk.

Name
Harshit Rai
Phone