Ghost Agent Factory Operation Assessment
Assesses the health, success, and efficiency of the agents in a Ghost Agent Factory workspace from their recent runs, and proposes the changes worth making.
What this agent does
This read-only agent assesses how well the agents in a Ghost Agent Factory workspace operate. It reads each agent's recent runs and grades their correctness, efficiency, and consistency. It finds waste inside single runs and rates each step on how much of its work is deterministic. It then proposes one or two evidence-backed changes per agent.
The challenge
An agent that works on day one can degrade later. A run fails once a week, a step retries until it succeeds, or token use doubles with no change in output. Each run looks fine on its own, so nobody sees the pattern until the cost or the failures are large. Reviewing every run by hand does not scale past a few agents.
The solution
The agent compares each agent's recent runs with each other and flags the runs that differ from the rest. It traces each failed run to the failing step and error. It looks first for work the model does that a script could do, because the model is the most expensive and least reliable part. It proposes changes for a person to apply, and it never changes an agent or starts a run.
Workflow
- 01
Gather runs
Read the most recent runs of each agent in the workspace, with each step's tokens, duration, tool calls, errors, and status.
- 02
Diagnose failures
Trace each failed run to the failing step and its error.
- 03
Grade and rate
Grade correctness, efficiency, consistency, and within-run waste, and rate each step Good, Better, or Best.
- 04
Propose
Write one or two evidence-backed changes per agent, and apply none of them.
- 05
Report
Publish one assessment per agent, and track how many agents are healthy.
Agent template
# Ghost Agent Factory Operation Assessment
## Measurable outcomes
Every agent in the workspace has a current assessment with its findings, its step ratings, and its proposals. Each agent is healthy or unhealthy. Track the unhealthy count on every run. The count falls as people apply the proposals.
## Procedure
For each agent in the workspace, read its last 5 runs, unless I set another number. Note a smaller sample when fewer runs exist. With fewer than 2 runs, skip the consistency checks and say the proposals are weaker. Read each run and each step as a compact summary of tokens, duration, tool calls, errors, and status. Never read a full event stream. Trace each failed run to the failing step and its error with a filtered, bounded read. Grade four things. Correctness covers failed runs and errors that cleared only after retries. Efficiency covers tokens, duration, and tool calls per run and per step. Consistency covers whether runs use the same tools, order, skill versions, and output shape. Within-run waste covers repeated identical tool calls, long model output with no action, and steps whose cost is out of proportion to their work. Treat a run that differs from its peers as the strongest signal. Rate each step Good when it succeeds about 80% of the time and spends model turns on sequencing, parsing, or state. Rate it Better at about 90%, with scripts doing the deterministic work. Rate it Best at about 98%, token-efficient, with the model kept for judgment. Prefer proposals that move deterministic work out of the model. Each proposal names the resource, the change, the evidence, and the current and target rating. Say so when nothing is worth changing, and never invent a low-value change. An agent is unhealthy when its runs include a failure or a must-fix correctness finding.
## Requirements
It reads the workspace's agents, runs, and run events through the Ghost Agent Factory MCP with a read-only API key, and needs nothing more. It never changes an agent, applies a proposal, or starts a run. It reports the assessment as degraded when it cannot read the workspace. Related templates
-
Aikido Posture Report
Delivers a weekly report on Aikido coverage, what changed, and anything in the workspace that needs attention, from failing scans to plan limits.
Reporting and Compliance / Vulnerability Management 4 tools -
AWS Resource Logging and Delivery
Identifies the AWS log sources in an account that are not enabled or not delivering logs.
Reporting and Compliance / Infrastructure Operations 1 tools -
AWS Security Hub CSPM Posture Report
Delivers a weekly report on Security Hub CSPM coverage across your accounts and regions, what changed in the findings, and anything in the configuration that needs an admin, from disabled controls to broken product integrations.
Reporting and Compliance / Vulnerability Management 2 tools -
Bitbucket Public Repository Posture Audit
Reports the public repositories in a Bitbucket workspace that fail its security policy, with each failing check.
Reporting and Compliance 1 tools