Autonomous SRE Agent

CI failures, fixed before anyone wakes up.

A test fails on your GitHub Actions. Nightingale receives the webhook, loads the failing file, calls GPT-4o for a typed fix plan, runs the full test suite in an isolated sandbox, and opens a pull request. Under 10 seconds.

"Most CI failures are not interesting."

A wrong assertion. A class imported with the wrong casing. An off-by-one that only shows in tests. These take a senior engineer two minutes to fix and cost an hour of their day once you add the notification, context switch, branch, PR, and review wait. Nightingale handles the two-minute fixes so engineers only see the ones that actually need them.

The real cost

Two minutes of work. One hour of overhead.

The fix itself is almost never the slow part. Every routine CI failure pulls a senior engineer out of deep work, forces them to context-switch, open a branch, wait for review, and wait for CI to re-run.

That overhead compounds. Ten incidents a week at 77 minutes each is 770 minutes of senior engineering time spent on work a machine can do.

Without NightingaleTime
Notification arrives, context switch15 min
Read logs, find root cause20 min
Write the actual fix2 min
Branch, commit, open PR10 min
Wait for review + CI re-run30+ min
Total per incident77+ min
Without Nightingale
Engineer paged at 3am
Senior time on routine work
Blocked until human acts
77+ minutes per failure
With Nightingale
PR waiting in the morning
Engineer reviews, not triages
Pipeline unblocked in seconds
Under 10 seconds per failure
Live pipeline demo

Watch the pipeline run.

0.0s pipeline elapsed
Push
CI Fails
Webhook
GPT-4o
Sandbox
Score
PR Opened
nightingale -- live pipeline idle -- press Run Pipeline to start
Nightingale SRE v1.0.0 -- ready
Listening on /webhook/github
Press "Run Pipeline" to simulate a CI failure and watch the full autonomous repair loop.
Confidence breakdown
test_pass_ratio35%--
blast_radius25%--
attempt_penalty15%--
risk_modifier15%--
self_consistency10%--
Weighted total --
Pull request opened
#42
Fix test_calculate_total assertion value
nightingale/fix/test-calculate-total-3f9a2c1 → main
1 file changed    1 insertion    1 deletion
Confidence: 0.97 (threshold: 0.85)
Resolved in: 8.3s
Nightingale opened this PR and will not merge it. Human review and approval required.
Scoring

Every resolution decision has a number attached.

Five weighted factors determine whether Nightingale acts or escalates. The formula is transparent. Every factor is logged. There is no black box.

Below 85%, nothing is written. The fix plan and full incident report are generated anyway and sent to the engineer, so even escalations are useful.

Resolve threshold: 0.85  |  adjustable in config.yaml
Confidence formula -- weighted 5-factor score
35%test_pass_ratioAll tests pass in sandbox?
25%blast_radiusFewer files changed = safer
15%attempt_penaltyFirst attempt scores higher
15%risk_modifiergpt-4o-mini classifies file risk
10%self_consistencyModel's stated confidence
Confidence breakdown in terminal
Confidence breakdown printed inline. The decision, score, and per-factor values are visible on every run.
Live dashboard

Everything logged, nothing hidden.

Every incident, confidence score, root cause, and pipeline trace persists in SQLite and survives restarts. The dashboard polls every 3 seconds with DOM-diffed updates -- no page flash, no visual disruption.

Nightingale live incident feed
Live incident feed. Green = resolved. Amber = active. Red = escalated.
Nightingale incident detail
Incident detail. Root cause, pipeline trace, and per-factor confidence breakdown.

Seven components.
One pipeline.

Each component has a single job. Nothing is shared across concerns. The full pipeline runs in under 10 seconds for most failures -- context load, reasoning, sandbox verification, confidence scoring, and PR creation.

>_

MarathonAgent

GPT-4o fix generation with a reflective loop. Up to 3 attempts. Failure output from each attempt feeds back into the next prompt.

▶

VerificationAgent

Applies the fix in a sandbox, runs the full pytest suite, and returns a typed VerificationResult with pass ratio and output.

#

ConfidenceScorer

Five-factor weighted formula. Uses gpt-4o-mini for independent risk classification of changed files. Transparent, logged per-factor.

✓

Sandbox

Copies the repository to a temporary directory. SHA-256 hashes the original before and verifies it after. Any mismatch rejects the result.

⇧

GitHubPRCreator

Creates a branch, commits the fix, and opens a PR via the GitHub REST API. Falls back to direct file write if no token is configured.

●

IncidentDatabase

SQLite-backed store for incident history, live status updates, metrics, and config. Survives process restarts.

~

Orchestrator

Coordinates the full pipeline. Registers incidents, drives the reflective loop, calls the scorer, and hands off to the resolution engine.

☍

Webhook + Dashboard

FastAPI server. Receives GitHub webhooks, serves the live dashboard, and exposes a REST API for incident history and metrics.

Webhook received --> Context loaded --> GPT-4o fix plan (up to 3 attempts) --> Sandbox verification
                   --> Confidence score (5-factor) --> >= 0.85: GitHub PR opened / < 0.85: escalate with full report
<10s
Typical resolution time
0.85
Confidence threshold to act
3
Max GPT-4o reasoning attempts
0
PRs merged without human approval
Quick start

Running in three commands.

Nightingale runs locally against your own repository. Set an API key, run the demo to verify everything works, then start the server and point a webhook at it.

01

Install and configure

Clone the repository, install dependencies, and set your OpenAI API key.

shellgit clone https://github.com/maxoutlabs/nightingale-sre.git cd nightingale-sre pip install -r requirements.txt # Linux/macOS: export OPENAI_API_KEY='sk-...' # Windows PowerShell: $env:OPENAI_API_KEY = 'sk-...'
02

Run the self-check and demo

Nine-point diagnostic confirms everything is wired up. The demo runs all three scenarios against the included test repo.

shellpython main.py --self-check # 9-point diagnostic python main.py --demo # interactive scenario picker python main.py --scenario 1 # wrong assertion python main.py --scenario 2 # broken import python main.py --scenario 3 # off-by-one bug
03

Start the server and dashboard

The server runs on port 8000. Open the dashboard at localhost:8000 to see the live incident feed.

shellpython main.py --server # dashboard: http://localhost:8000 # webhook: http://localhost:8000/webhook/github # api: http://localhost:8000/api/v1/incidents
Production

Wire it into any GitHub repo in four steps.

Nightingale needs to run on a machine with network access and a local copy of your repository. A small VPS, an EC2 instance, or the CI server itself all work. The webhook does the rest.

Step 01
Run the server
Start Nightingale on any machine that has your repo cloned. Expose it with a public URL or an ngrok tunnel for testing.
python main.py --server --port 8000
Step 02
Add the webhook
In your GitHub repo settings, add a webhook pointing at your server URL. Select workflow run and check run events.
Settings > Webhooks > Add webhook Payload URL: https://your-host:8000 /webhook/github Events: workflow runs, check runs
Step 03
Set credentials
Set a GitHub token for PR creation and an optional Slack webhook for escalation notifications. Both are optional.
export GITHUB_TOKEN='ghp_...' export GITHUB_REPO='owner/repo' export SLACK_WEBHOOK_URL='https://...'
Step 04
Push a failing commit
That is it. The next CI failure triggers the full pipeline. A PR will be open before anyone is notified.
# config.yaml: confidence: resolve_threshold: 0.85 # raise to 0.90 for more # conservative auto-resolve

The next failure is
already handled.

Open source. Self-hosted. No merge without human review.

View on GitHub → Read the docs