Google Cloud Run SRE
Reviews a Cloud Run service's metrics, logs, and errors each day, traces each error to its source code, and proposes the fix.
What this agent does
This read-only agent checks the health of a Google Cloud Run service. It reads the service's traffic, latency, and resource metrics, and it groups its errors from Error Reporting and Cloud Logging. It ties each error group to the revision that produced it. It then traces the errors that code can fix to a file and line in the service's repository and proposes the fix.
The challenge
A Cloud Run service can fail a small share of its requests for days before anyone looks. The errors are spread across logs, Error Reporting, and metrics, and no single view shows all of them. A new revision can introduce the errors, or an upstream API can cause them. Engineers spend the first hour of every investigation finding out which one it is and where the code is.
The solution
The agent reads every signal for the service in one pass and groups errors by cause, not by log line. It separates errors that code can fix from upstream failures and client errors. It flags the errors that start with a recent deploy and lists the commits in that deploy. It flags saturation only, never spare capacity, so every finding is something an engineer can act on.
Workflow
- 01
Read the service
Read the service's limits, its current and recent revisions, and the image each revision runs.
- 02
Read metrics
Read request counts by status class, latency percentiles, CPU and memory use, and instance counts for the window.
- 03
Group errors
Group the window's errors from Error Reporting and from error logs that Error Reporting missed.
- 04
Correlate revisions
Tie each error group to its revisions, and flag the groups that started with a recent deploy.
- 05
Trace to code
Find each fixable error in the service's repository and propose the fix at its file and line.
- 06
Report
Publish the health headline and the fixable issues, or only the headline when the service is healthy.
Agent template
# Google Cloud Run SRE
## Measurable outcomes
Every error group in the service's window is either traced to a file and line with a proposed fix or marked as external. Every recent deploy that introduced errors is flagged with its commits. Track the 5xx rate and the count of fixable issues on every run.
## Procedure
For a given Cloud Run service, review the last 24 hours, unless I set another window. Read the service's limits and its recent revisions, with the git commit each revision's image was built from. Compute the error rate from 5xx responses. Show 4xx responses separately as client errors. Read the p50, p95, and p99 latency, the p95 CPU and memory use, and the peak instance count. Flag CPU above 80%, memory above 85%, and instances held at the maximum scale. State the direction of a fix, never a new value. Never flag spare capacity. Group errors from Error Reporting and from error logs without stack traces, and say when the log read was truncated. For each group, check its revisions. A group concentrated on a current revision deployed within the last 6 hours is a bad-deploy suspect, so list the commits between the prior revision and it. A group on a revision with no traffic is likely fixed. A group spread across revisions is long-standing. Flag any revision whose image comes from a different image repository than the service's own. For the top 10 groups, find the error in the service's repository from its reported location or its message text. Upstream API failures, network timeouts, expired tokens, and client errors are external. For each other group, name the file and line, what the code does there, and a one-line fix. When the platform retries failed work, check whether the retries succeeded. Failures that recovered on retry are a delay, not lost work. The issues list contains only fixable issues and saturation flags. It is empty when the service is healthy.
## Requirements
It needs read access to the service in Cloud Run, Cloud Monitoring, Cloud Logging, Error Reporting, and Artifact Registry, and read access to the service's repository in GitHub, GitLab, or Bitbucket, and nothing more. The source code provider and repository for the service are set when the agent is built. It never deploys, changes the service's configuration, or changes code. Related templates
-
AWS Resource Logging and Delivery
Identifies the AWS log sources in an account that are not enabled or not delivering logs.
Reporting and Compliance / Infrastructure Operations 1 tools -
Datadog Service SRE
Reviews a service's Datadog metrics, traces, error tracking, and deployments each day, traces each error to its source code, and proposes the fix, on any runtime that reports to Datadog.
Infrastructure Operations 4 tools -
Ghost Agent Factory SRE
Checks the health of the Ghost Agent Factory that every agent runs on, and finds the failed runs the platform caused.
Ghost Agent Factory / Infrastructure Operations 1 tools