Documentation
Technical reference for Nightingale SRE -- from webhook setup to confidence scoring internals.
Architecture
Nightingale is a stateless webhook service. A failed CI run triggers a GitHub webhook, Nightingale diagnoses the failure, generates a fix, verifies it in an isolated sandbox, and opens a pull request. No persistent agent loop, no background polling.
Components
| Module | File | Responsibility |
|---|---|---|
| WebhookReceiver | nightingale/webhook.py | Receive, verify, and dispatch GitHub webhook events |
| Diagnoser | nightingale/diagnoser.py | Call GPT-4o with failure context, extract structured root cause |
| Patcher | nightingale/patcher.py | Generate unified diff patch from diagnosis and source context |
| SandboxRunner | nightingale/sandbox.py | Apply patch to isolated copy, run tests, capture output |
| ConfidenceScorer | nightingale/scorer.py | Score fix quality across 5 dimensions, decide whether to open PR |
| GitHubClient | nightingale/github_client.py | Create branches, commit patches, open pull requests via REST API |
| Dashboard | nightingale/dashboard.py | Bloomberg-style terminal UI (Rich library) |
Quick Start
Get Nightingale running locally in under 5 minutes.
1. Clone and install
git clone https://github.com/maxoutlabs/nightingale-sre cd nightingale-sre pip install -r requirements.txt
2. Configure credentials
cp .env.example .env
# Edit .env with your values:
OPENAI_API_KEY=sk-...
GITHUB_TOKEN=ghp_...
GITHUB_WEBHOOK_SECRET=your-webhook-secret
TARGET_REPO=owner/repo3. Run the server
python -m nightingale.main INFO Nightingale SRE starting... INFO Webhook server on http://0.0.0.0:8765 INFO Target repo: owner/repo INFO Confidence threshold: 70
4. Expose via ngrok (for GitHub webhook)
ngrok http 8765 # Copy the Forwarding URL, e.g. https://abc123.ngrok.io # Add it as your webhook URL in GitHub: /webhook/github
Core Concepts
Fix Pipeline
A single end-to-end repair takes three LLM-free steps plus one GPT-4o call. The LLM call happens only at the diagnosis stage -- every other step is deterministic code.
Confidence Threshold
Nightingale will not open a PR unless the fix scores above the configured confidence_threshold (default: 70). Below that, the event is logged and flagged for human review. This prevents low-quality patches from landing in your codebase.
Sandbox Isolation
Every patch is tested in an isolated copy of the repository. Nightingale clones the repo into a temp directory, applies the unified diff, runs the test suite, and captures stdout/stderr. The original repo is never modified until the PR is opened.
Idempotency
Nightingale tracks which check_run IDs have already been processed. Re-delivered webhooks are ignored. If an open PR for a given branch already exists, Nightingale updates it rather than opening a duplicate.
Webhook Setup
Nightingale listens for GitHub check_run events. Configure a webhook in your repository settings.
GitHub repository settings
| Field | Value |
|---|---|
| Payload URL | https://your-host/webhook/github |
| Content type | application/json |
| Secret | Any strong random string (matches GITHUB_WEBHOOK_SECRET) |
| Events | Check runs (or "Send me everything" for testing) |
Payload verification
Every incoming request is verified using HMAC-SHA256 before any processing. The X-Hub-Signature-256 header is checked against the raw request body and your webhook secret. Requests with missing or invalid signatures are rejected with HTTP 401.
# nightingale/webhook.py -- signature check import hmac, hashlib def verify_signature(payload: bytes, secret: str, sig_header: str) -> bool: expected = "sha256=" + hmac.new( secret.encode(), payload, hashlib.sha256 ).hexdigest() return hmac.compare_digest(expected, sig_header)
Event filtering
Nightingale only acts on check_run events where:
- Action is
completed - Conclusion is
failure - The check run is on the default branch (or a configured branch)
Configuration
Nightingale reads configuration from nightingale.yml in the project root. All keys are optional -- defaults are shown below.
confidence_threshold: 70 # 0-100, PR only opens above this max_attempts: 3 # retries per failure before giving up sandbox_timeout: 60 # seconds to wait for test run model: gpt-4o # OpenAI model for diagnosis test_command: pytest # command to run tests in sandbox pr_draft: false # open PRs as draft pr_label: nightingale # label to apply to opened PRs branch_prefix: nightingale/fix # prefix for repair branches notify_slack: false # post to Slack on each outcome dashboard: true # show terminal dashboard
| Key | Type | Default | Description |
|---|---|---|---|
| confidence_threshold | int | 70 | Minimum score to open a PR. Below this, fix is logged but not submitted. |
| max_attempts | int | 3 | How many times to retry patch generation if the sandbox run fails. |
| sandbox_timeout | int | 60 | Seconds before the sandbox test run is killed. |
| model | str | gpt-4o | OpenAI model used for root cause analysis and patch generation. |
| test_command | str | pytest | Command run inside the sandbox to verify the fix. |
| pr_draft | bool | false | Open pull requests as drafts for human review before merge. |
| branch_prefix | str | nightingale/fix | Prefix for repair branches. Full branch: {prefix}-{run_id}. |
Environment Variables
Required secrets go in .env (never commit this file). Nightingale reads them via python-dotenv at startup.
| Variable | Required | Description |
|---|---|---|
| OPENAI_API_KEY | required | OpenAI API key for GPT-4o calls |
| GITHUB_TOKEN | required | Personal access token or GitHub App token with repo scope |
| GITHUB_WEBHOOK_SECRET | required | HMAC secret matching the webhook configuration in GitHub |
| TARGET_REPO | required | Repository to monitor, in owner/repo format |
| NIGHTINGALE_PORT | optional | Webhook server port (default: 8765) |
| SLACK_WEBHOOK_URL | optional | Incoming webhook URL for Slack notifications |
| LOG_LEVEL | optional | Log verbosity: DEBUG, INFO, WARNING, ERROR (default: INFO) |
repo (read code, create branches, open PRs) and read:checks (read check run logs). Use a fine-grained token scoped to only the target repository.
REST API
Nightingale exposes a small HTTP API alongside the webhook endpoint. All responses are JSON.
Endpoints
| Method | Path | Description |
|---|---|---|
| POST | /webhook/github | Receive GitHub webhook events |
| GET | /health | Health check -- returns {"status":"ok"} |
| GET | /status | Current agent status and last repair details |
| GET | /repairs | List of recent repair attempts with outcomes |
| GET | /repairs/{id} | Full details of a specific repair: diagnosis, patch, confidence, PR URL |
| POST | /trigger | Manually trigger a repair for a given commit SHA |
GET /status
{
"status": "idle",
"total_repairs": 14,
"success_rate": 0.857,
"last_repair": {
"id": "rep_a1b2c3",
"timestamp": "2026-05-17T09:14:22Z",
"outcome": "pr_opened",
"confidence": 84,
"pr_url": "https://github.com/owner/repo/pull/47"
}
}POST /trigger
// Request body { "repo": "owner/repo", "sha": "abc123def456", "check_run_id": 9876543210 }
Confidence Scoring
Before opening a PR, Nightingale scores the fix across five dimensions. The composite score must exceed confidence_threshold (default 70) or the patch is discarded.
Scoring dimensions
| Dimension | Weight | How it's measured |
|---|---|---|
| test_pass_rate | 35% | Fraction of tests passing after patch is applied in the sandbox |
| patch_precision | 25% | Lines changed / lines in failing test file -- smaller targeted patches score higher |
| syntax_validity | 20% | Whether the patched file parses without syntax errors |
| diagnosis_confidence | 10% | GPT-4o self-reported confidence from structured output |
| scope_containment | 10% | Whether the patch is limited to the identified failure file |
# nightingale/scorer.py -- composite scoring def score(self, result: SandboxResult, diagnosis: Diagnosis) -> ConfidenceScore: components = { "test_pass_rate": result.pass_rate * 35, "patch_precision": self._precision(result) * 25, "syntax_validity": result.syntax_ok * 20, "diagnosis_confidence": diagnosis.confidence / 100 * 10, "scope_containment": self._scope(result) * 10, } total = sum(components.values()) return ConfidenceScore(total=total, components=components)
confidence_threshold for repositories with flaky tests where even partial fixes are valuable. Raise it for production-critical repos where every merged fix must be clean.
Sandbox
The sandbox is how Nightingale verifies a fix before submitting it. It runs entirely locally with no external calls.
Sandbox modes
| Mode | How it works | When to use |
|---|---|---|
| venv | Creates a fresh virtualenv, installs dependencies, runs test command | Default -- works without Docker |
| docker | Spins up a container from the repo's Dockerfile, runs tests inside | When parity with CI environment matters |
| subprocess | Runs tests in the current Python environment | Quick local development, no isolation |
Configure the mode via sandbox_mode in nightingale.yml (default: venv).
What the sandbox runs
# Steps for venv mode
git clone --depth=1 {repo} .sandbox/run_{id}
cd .sandbox/run_{id}
git apply patch.diff
python -m venv .venv
.venv/bin/pip install -r requirements.txt -q
.venv/bin/pytest --tb=short --timeout=30 -qSafety Controls
Nightingale includes several guardrails to prevent bad patches from reaching your codebase.
Patch scope limits
By default, Nightingale will refuse to apply a patch that:
- Modifies more than 3 files
- Changes more than 50 lines across all files
- Touches files outside the test directory or the identified failure file
- Deletes any test functions
LLM output validation
The GPT-4o response is parsed as structured JSON via function calling. If the model returns a malformed response or a patch that fails JSON schema validation, the attempt is retried up to max_attempts times before the event is abandoned.
No auto-merge
Nightingale only opens pull requests. It never merges. A human (or your existing branch protection rules and CI checks) decides whether a fix is acceptable.
Audit trail
Every event is written to nightingale.db (SQLite). Each record includes the original failure, the diagnosis, the patch diff, sandbox output, confidence score, and the final outcome. Query it with any SQLite client to audit Nightingale's decisions.
# View last 10 repair attempts sqlite3 nightingale.db " SELECT id, timestamp, outcome, confidence FROM repairs ORDER BY timestamp DESC LIMIT 10; "