Ghost Agent Factory SRE
Checks the health of the Ghost Agent Factory that every agent runs on, and finds the failed runs the platform caused.
What this agent does
This read-only agent checks the health of the Ghost Agent Factory in one workspace. It reads connector test results, worker pools, token budgets, proxy traffic, webhook deliveries, approvals, the run queue, and the platform version. It decides which failed runs the platform caused and which the agents caused. It ranks each finding by the number of workflows it affects.
The challenge
When an agent run fails, the first question is whether the agent broke or the platform under it did. A failed connector, a worker pool with no workers, or a blocked host can fail many workflows at once. Each failure looks like an agent problem until someone checks the platform. For the same reason, nobody notices an outdated platform version or a stalled upgrade.
The solution
The agent reads every platform signal in the workspace in one pass. It compares each failed run with a fixed list of platform causes and shows the evidence for each match. It ranks findings by the number of workflows they affect, so the fix that unblocks the most work comes first. It reads saved results only and changes nothing, so it runs on a read-only key.
Workflow
- 01
Read platform signals
Read connector test results, worker pools, token budgets, proxy traffic, webhook deliveries, approvals, the run queue, and the platform version.
- 02
Find platform problems
Turn each signal past its threshold into a finding, and list the workflows it affects.
- 03
Attribute failed runs
Compare each failed run with the platform causes, and set aside the failures the agents caused.
- 04
Rank and report
Rank findings by affected workflows, and open the report with a one-line verdict.
Agent template
# Ghost Agent Factory SRE
## Measurable outcomes
Every platform problem in the workspace is in the report, with the workflows it affects. Every failed run the platform caused is in the report, with its evidence. Track the open findings and the platform-caused failed runs on every run.
## Procedure
For a given workspace, review the last 24 hours, unless I set other values. Report each connector whose saved last test failed. Report each worker pool with fewer live workers than its configured replicas. Report each capability that a bound skill requires and no pool provides. Report each workspace or workflow above 80% of its token rate budget. Report proxy denials and DNS failures, grouped by host. Report failed webhook deliveries, grouped by webhook. Report approvals waiting more than 24 hours. Report runs queued more than 15 minutes without a worker. Report an available platform update, and an upgrade still pending after 1 hour. For each finding, list the workflows it affects. For each failed run, read its error and event summary and compare them with these platform causes: no worker or no capability for the run, a queue timeout, a proxy denial or DNS failure, a credential or connector failure, a token rate budget limit, an approval timeout, or a lost worker. Record a platform fault that matches none of them as other. Show the evidence for every failed run the report counts against the platform. Leave out the failures the agents caused. Rank findings by the number of workflows they affect, and put platform-caused failures first within a tie. Open the report with a one-line verdict: healthy, or the count of findings and affected workflows. When any read fails, report the result as incomplete, never as healthy.
## Requirements
It reads the workspace through the Ghost Agent Factory MCP with a read-only API key, and needs nothing more. Platform-wide signals such as worker pools and the platform version appear in every workspace's report. It never runs connector tests, changes a resource, or starts a run. Related templates
-
AWS Resource Logging and Delivery
Identifies the AWS log sources in an account that are not enabled or not delivering logs.
Reporting and Compliance / Infrastructure Operations 1 tools -
Datadog Service SRE
Reviews a service's Datadog metrics, traces, error tracking, and deployments each day, traces each error to its source code, and proposes the fix, on any runtime that reports to Datadog.
Infrastructure Operations 4 tools -
Ghost Agent Factory Operation Assessment
Assesses the health, success, and efficiency of the agents in a Ghost Agent Factory workspace from their recent runs, and proposes the changes worth making.
Ghost Agent Factory / Reporting and Compliance 1 tools -
GitHub Dependabot Triage Agent Report
Provides a weekly compliance report on the agent runs inside the Ghost Agent Factory for Dependabot triage and recheck agents, with alerts dismissed, alerts reopened, and analyst time saved.
Vulnerability Management / Reporting and Compliance / Ghost Agent Factory 2 tools