v0.3 — policy-gated remediation, causal root-cause, MCP integrations, learned auto-resolve
Point Sentryops at any backend or deployment — a plain URL or a Kubernetes namespace — and it discovers what's actually running behind it, probes it Postman-style, tracks its health continuously, and turns a wall of logs into a root cause and a fix. A dependency graph ranks the real cause over its symptoms, and an approved fix executes for real behind a policy gate. Known-safe errors resolve themselves, and the list of "known-safe" keeps growing on its own. Everything else waits for you, with the full evidence trail attached.
Tier-1 resolves itself. Tier-2 waits for a person. Every decision — either way — is written to an audit log you can read.
Every on-call rotation hits the same three walls. Sentryops exists to knock all three down at once.
A "health check" is usually one URL, one convention, one guess. Real deployments are dozens of endpoints and pods, and nobody's kept the list current.
The answer is in there somewhere, in among 100,000 lines nobody has time to read. By the time someone greps it out, the incident's been open for an hour.
A connection reset that would've cleared itself on retry still wakes someone up at 3am, because nothing in the pipeline is trusted to tell "boring" from "dangerous."
The same pipeline runs whether the answer turns out to be "nothing's wrong" or "restart the payments pod" — nothing about how carefully it checks depends on how it thinks the story ends.
Matches your query to the specific service(s) it's actually about.
Drafts an investigation: which checks, in what order, against which service.
Runs each step for real — service registry, tickets, runbooks, live endpoint probes, Kubernetes pods, reduced & diagnosed logs.
Weighs the evidence and drafts a decision, with its reasoning attached.
Non-LLM backstops double-check the target is unambiguous, and that a known-safe error signature — not just a hopeful guess — is what's letting anything skip human approval.
Executes immediately, or opens an approval and waits.
Pick a real query and step through the same six-stage pipeline every request runs — Resolve, Plan, Investigate, Reason, Safety guard, Act or ask. No account, no login, nothing to install.
Real evidence, replayed — not a live cluster. Point Sentryops at your own deployment to watch it decide for real.
Create a free accountNot a chatbot bolted onto your infra — a planner with real tools, a real memory, and a real opinion about what's safe to touch.
| Capability | What it does | Why it matters |
|---|---|---|
| Multi-endpoint discovery | Register a small Postman-style collection of paths (or let it fall back to conventional health paths) and bulk-probe every one, concurrently, with a real pass/fail per endpoint. | One URL was never the whole picture. Now the whole picture is one click. |
| Real Kubernetes pod discovery | Give it a namespace and it asks kubectl directly — which pods exist, their phase, container readiness, restart counts. |
Ground truth from the cluster, not a cached label someone forgot to update. |
| Continuous health tracking | A background scheduler re-checks every registered service on an interval and keeps a full history — and opens a ticket automatically the moment something goes unhealthy. | You find out from Sentryops, not from a customer. |
| Deterministic log reduction | Paste or upload raw logs — thousands of lines is fine. They're deduplicated and clustered by normalized signature, then ranked by severity and frequency, before anything touches a model. | 100,000 lines become a dozen clusters you can actually read, in milliseconds, for free. |
| LLM root-cause diagnosis | The reduced clusters are handed to a language model that returns a root cause, the exact implicated pattern, and concrete remediation steps. | Not "here are your logs" — "here's what's wrong and here's how to fix it." |
| Curated + learned auto-resolve | A small, explicit whitelist of known-transient signatures (connection resets, timeouts, DNS blips) resolves automatically — plus a per-deployment learned memory that starts empty and promotes a new signature only after an unbroken run of human approvals. One rejection resets it to zero. | The boring 3am pages stop paging anyone, and the list of "boring" keeps growing on its own. |
| Causal root-cause (Production World Model) | Declare which services depend on each other, and a persisted dependency graph ranks which currently-unhealthy service actually explains the others — not just which one happened to page first. | Diagnosis stops being per-service guesswork and starts being graph-aware. |
| Real, policy-gated remediation | An approved restart, scale, or rollback issues a real kubectl mutation — gated end-to-end by an OPA-style policy engine (namespace allowlist, denied services, a dry-run switch) that's independent of both the LLM and the approving human. |
Closes the loop from "here's the root cause" to "here's the fix, executed" — with hard boundaries a buyer can actually audit. |
| Tiered human approval | Anything ambiguous, multi-service, or explicitly requested by name always stops for a person to approve or reject. | Autonomy where it's earned, a human in the loop everywhere else. |
| MCP integrations | A protocol-correct Model Context Protocol client — not a hardcoded wrapper for one vendor. Point it at the Slack, PagerDuty, GitHub, or Datadog/Prometheus MCP server (or any other) and an approval being opened or decided notifies it automatically. | Plugs into the tools you already run instead of asking you to adopt new ones. |
| Full audit trail | Every query, plan, step result, approval, policy decision, MCP notification, and executed action is logged with a correlation id you can trace end to end. | Nothing the agent does is a black box after the fact. |
| Benchmarked accuracy | A benchmark harness runs the real auto-resolve and root-cause engines against a labeled synthetic corpus and reports precision/recall/F1 — with a hard safety gate that fails the build if a dangerous scenario is ever auto-resolved. | An eval you can rerun and audit, not a number in a slide deck. |
| Embed it or call it | An embeddable Rust SDK with zero HTTP overhead, a dependency-free Go client + agentctl CLI, or plain REST from anywhere. |
Fits into what you've already built instead of demanding a rewrite. |
Same underlying problems. Very different Tuesday night.
| Without Sentryops | With Sentryops | |
|---|---|---|
| Health checks | One URL, pinged manually, trusted blindly | Multi-endpoint, Postman-style, plus real Kubernetes pod discovery |
| Log triage | Grep through tens of thousands of lines by hand | Deterministic clustering + LLM root cause in seconds |
| Root cause | Guess from whichever service happened to page first | A dependency graph ranks the actual cause over its symptoms |
| Remediation | Every alert pages a human, always | Known-safe errors auto-resolve — and approved fixes execute for real, gated by policy |
| Approval | Trust the on-call engineer's judgment alone | Non-LLM safety guards force review on anything ambiguous |
| Notifications | Custom webhook code per tool you use | Any MCP server (Slack, PagerDuty, GitHub) wired in via config |
| Integration | Bespoke scripts glued to one stack | Embeddable Rust SDK, Go client/CLI, or plain REST |
| History | Tribal knowledge and Slack scrollback | A queryable, correlation-id'd audit log of every decision |
A real Model Context Protocol client: JSON‑RPC over stdio, the actual handshake, not a wrapper hardcoded to one vendor's API. Point it at any MCP server and an approval opening or getting decided reaches it automatically.
An approval request or decision lands wherever your team already watches for incidents.
The same event opens or resolves a page without a separate integration to maintain.
Route the decision into a repo your engineers already have open.
Keep the change record in sync with what the agent actually did.
{
"servers": {
"slack": {
"command": "npx",
"args": ["-y", "@modelcontextprotocol/server-slack"],
"env": { "SLACK_BOT_TOKEN": "xoxb-..." }
}
},
"on_approval_requested": {
"server": "slack",
"tool": "post_message",
"extra_arguments": { "channel": "#incidents" }
}
}
No file at all means every hook is a silent no‑op, the same secure default the policy engine uses. Datadog and Prometheus fit the same protocol; pulling live metrics into an investigation as evidence, not just a notification target, is the natural next step and isn't built yet.
A benchmark harness (agent-bench) runs the exact deterministic decision engines agent-server calls on a live request against a labeled, disclosed-synthetic corpus — not a simulation, not marketing math.
By failure class, 27 scenarios
| Class | Scenarios | Should auto-resolve | Result |
|---|---|---|---|
| Transient network | 5 | Yes | 5/5 correct |
| Transient DNS | 2 | Yes | 2/2 correct |
| Auth & permission | 4 | No | 4/4 correct |
| Data integrity | 3 | No | 3/3 correct |
| Capacity exhaustion | 3 | No | 3/3 correct |
| Deploy regression | 3 | No | 3/3 correct |
| Security incident | 3 | No | 3/3 correct |
| Ambiguous / novel | 4 | No | 4/4 correct |
Plus 10 dependency-graph topologies for root-cause ranking: fan-ins, chains, a diamond, a symmetric cycle, and a disconnected node that must never get blamed just for being unhealthy nearby. This deliberately does not publish an MTTR-reduction percentage. That needs real incident timelines from real usage, which is what the pilot below is for. Results are versioned in agent-rs/bench-results/ so accuracy stays visible over time, and the corpus itself is in source control to audit, not a black box.
# from agent-rs/
cargo run -p agent-bench
An open SDK for teams who'd rather run it themselves, and a hosted product for teams who'd rather not, priced per monitored service instead of per host or per gigabyte.
Two ways to take it
Run agent-server yourself, or embed agent-sdk directly into an existing Rust service with zero HTTP overhead. The full decision engine ships with it: policy, World Model, and learned auto-resolve are never gated behind a license check.
agentctl CLI, or plain REST, if Rust isn't your stackWe run agent-server and the dashboard for you — a dedicated deployment per customer, its own database and config, not a shared multi-tenant platform. Managed upgrades, backups, and uptime, with the full audit trail hosted alongside it.
Pricing, per monitored service
Sized against what teams already pay: Datadog charges $15/host/month for infrastructure monitoring alone and $31/host/month for APM, and a typical mid-market observability stack (Datadog + PagerDuty + a handful of smaller tools) runs about $10,300/month. Flat per-service pricing here is deliberate — no per-GB or per-event surprises on top of a number you already agreed to.
The pilot
Real on-call pain, already running Kubernetes.
Full feature set, dedicated deployment, direct access to the team building it.
The same audit trail already built — against real incidents this time.
A real MTTR/auto-resolve-accuracy result, with permission, replacing "stays an open item."
SOC2: what's true today
The live-test dashboard is a real client of the API above — register a backend, bulk-probe it, discover its pods, paste in logs, and ask it free-text questions. It sits behind a free account so we know who's kicking the tires (and so your test data isn't world-readable).
Embed it directly in a Rust service with zero HTTP overhead, or call a running instance from anywhere with the Go SDK — or just speak REST.
use agent_sdk::{Agent, AgentConfig}; #[tokio::main] async fn main() -> anyhow::Result<()> { let agent = Agent::bootstrap( AgentConfig::from_env() ).await?; let res = agent .submit_query("is payments healthy?") .await?; println!("{:?}", res.tier); Ok(()) }
import agent "github.com/devops-helpdesk/agent-go" client := agent.NewClient("http://localhost:8000") resp, err := client.SubmitQuery(ctx, "is payments healthy?") if resp.ApprovalID != nil { approver := "alice" client.DecideApproval(ctx, *resp.ApprovalID, "approve", &approver) }
This started as a rewrite of an existing Python service — same routes, same JSON, same decision logic. What's different is underneath, and what's been added since.
| Before | Now | |
|---|---|---|
| Runtime | Python interpreter, per-request overhead | Compiled Rust binary |
| Deployment artifact | App code + a venv of dependencies | One static binary |
| Type safety | Runtime errors on malformed dynamic dicts | Checked at compile time |
| Health checks | Single URL, GET-only | Postman-style multi-endpoint + real k8s pod discovery |
| Logs | "No log aggregation source connected." | Deterministic reduction + LLM diagnosis + auto-resolve |
| Integration | HTTP only | Embeddable Rust SDK, Go SDK/CLI, or REST |
The full manual — configuration, every route, deployment options — lives on the docs page. Here's the short version.
# 1. set your model provider key export GROQ_API_KEY=your-key-here # 2. build and run the server cd agent-rs && cargo run -p agent-server # 3. from anywhere else, talk to it cd agent-go && go run ./cmd/agentctl query "is payments healthy?" # 4. approve or reject anything it flags go run ./cmd/agentctl approvals decide 1 --decision approve
Tried the demo, read the docs, poked at the API — good or bad, we want to hear it. No account needed.
Questions about deploying this, integrating it, or extending it — happy to talk.