Incident investigation

Investigate production incidents with the evidence already in your systems.

AetherOps groups related log failures into one incident, cites every claim to a specific log line, and lists what it can't determine instead of guessing. Built for backend teams running Postgres-backed services.

We're onboarding backend teams directly — no self-serve signup yet.


Evidence first

Every claim traces back to one log line

Sample incident — demo data, not a live system
Raw log line received 2026-09-14T08:12:03.441Z ERROR payment-svc HikariPool-1 - Connection is not available, request timed out after 30000ms.
Evidence ID log:4b9eb6c0
Claim in the analysis below "HikariCP pool saturation — maximum-pool-size (10) reached under load"

This is a real screenshot of the running app — not a mockup. Seeded log events went through the actual ingestion → grouping → analysis pipeline shown below.

Incident detail — payment-svc
AetherOps dashboard: sidebar listing 8 incidents of varying severity and status, with a CRITICAL open payment-svc incident selected showing AI analysis at 55% confidence, a likely cause cited to specific log evidence IDs, listed unknowns, and read-only recommended next checks
Incident detail — api-gateway same pipeline, different service
AetherOps incident detail screen showing a HIGH severity api-gateway incident, 5xx rate elevated on /v2/orders, with evidence-cited AI analysis at 55% confidence

The rest of the incident lifecycle

From first alert to resolved, not just the analysis screen

War rooms and on-call paging are live in the backend — Slack-linked, Twilio-paged, calendar-aware — they just don't have their own dashboard page yet, so they're shown here in the same design system rather than as live screenshots.

War room — auto-opened
Active
payment-svc: HikariCP connection pool exhausted
#inc-2026-09-14-payment-svc opened automatically at CRITICAL
Participants
  • Incident Commander · joined 1m ago
  • Responder · joined 1m ago
  • Responder · joined just now
Every join, leave, and ack is written to the incident timeline  ·  closes automatically when the incident resolves
On-call — payment-svc
Payment Platform Primary
Current on-call: •••• 4471 12h shifts · handoff in 3h 40m
Escalation policy
  • 1 Page primary — SMS + phone call
  • 2 No ack in 5 min — escalate to secondary
  • 3 No ack in 10 min — page the incident commander
Skips anyone marked unavailable via connected calendar (PTO-aware)  ·  delivered over Twilio

How it works

One log event, followed all the way through

Same eight stages every time. Here's real data going in and what comes out the other side.

Log event in
{
  "serviceId": "payment-svc",
  "level": "ERROR",
  "message": "HikariPool-1 - Connection is
    not available, request timed out
    after 30000ms.",
  "timestamp": "2026-09-14T08:12:03.441Z"
}
  1. ingest
  2. correlate
  3. group
  4. analyze
  5. evidence
  6. explain
  7. unknowns
  8. next checks
Incident out
{
  "title": "payment-svc: HikariCP
    connection pool exhausted",
  "severity": "CRITICAL",
  "status": "OPEN",
  "eventCount": 3,
  "confidenceScore": 0.58
}

Integrations

Where incidents get reported and resolved

What's built and shipped today.

Slack (incoming webhook)

Paste your own Slack incoming webhook URL to post incident alerts — no app install required.

Jira

Create and link a Jira ticket directly from an incident, using your own Jira credentials.

ServiceNow

Create and link a ServiceNow ticket directly from an incident, using your own ServiceNow credentials.

SSO / SAML

Configure your own SAML identity provider for tenant login.


Trust & security

What's actually enforced

Every claim below is something you can find and read in the source, not a compliance badge.

Raw log line received POST /api/v1/checkout Authorization: Bearer eyJhbGciOiJIUzI1NiJ9.eyJzdWIiOiJ1c2VyXzQ0MiJ9.k3f9x7 user=442
What's actually stored POST /api/v1/checkout Authorization: Bearer [REDACTED] user=442

SHA-256 API keys

Raw keys are shown once at creation. Only the hash is stored.

Postgres row-level security

Every tenant-scoped table is RLS-enforced, including against the app's own database role — not just application-layer filtering.

Per-key rate limiting

Ingestion is rate-limited per API key, configurable per deployment.

Constant-time comparisons

2FA verification uses constant-time comparison, not a timing-vulnerable equality check.

Real-database test suite

Tenant isolation is tested against a real Postgres instance with RLS enforced, not mocked out.


Pricing

Usage limits, not guesswork

Limits are enforced per calendar month and reset automatically. What's below is what the system actually enforces.

Configurable log retention (up to 365 days) is available on every plan and purged automatically — it isn't a paid tier feature.

Free
$0/mo
For solo developers and side projects.
  • 5 AI analyses / month
  • 10 runbook generations / month
  • 5 postmortem drafts / month
  • Incident grouping
Starter
$29/mo
For small teams shipping to production.
  • 50 AI analyses / month
  • 50 runbook generations / month
  • 25 postmortem drafts / month
  • 50 incidents / month
Team
Custom
For teams running multiple services in production.
  • Everything in Starter
  • Meaningfully higher monthly limits, scoped to your usage
  • Priority support
  • Pricing set in a direct conversation, not a fixed tier

See it against your own logs

We're onboarding backend teams directly and iterating on real feedback.

Read the API reference