One canary rule, three systems to reconcile

6 minute read     Updated:

Kate Gooch
Kate Gooch

An engineering standard only works at scale when it reaches every service it applies to, draws on the right evidence, and keeps working as those services change.

Take the sample rule: “Every production release of a critical service must use a canary rollout.”

A canary setting can pass inspection while the next release skips the canary. The application changes in one repo, rollout configuration lives in another, and the service’s criticality is recorded in a catalog. The engineer making the change may never open the other two.

Can the platform team connect those records, verify what happened during deployment, and spot a critical service missing from the check altogether?

We’ll follow that example rule across all three systems and show how Earthly Lunar helps turn scattered evidence into a repeatable guardrail. Platform teams can thus spot coverage gaps, give developers clear feedback, and block releases when they trust the evidence.

Start with the service, not the repo count

A dashboard can tell you how many repositories pass a check. That number gets a little fuzzy when a repository is not the same thing as a service.

One repo might contain several services. Another service might have application code in one repo and deployment configuration in another. Criticality may live in a catalog neither repo controls.

For the canary requirement, the platform team needs a list of critical services and a way to connect each one to its Helm configuration. Then, it needs to keep those relationships current as services are added, renamed, and retired. What sounded like a one-line rule has become an ongoing inventory problem.

Blank results deserve attention, too. A service with an approved exception, one whose configuration could not be checked, and one that has never been mapped are different problems. None should count as a pass.

One service, three sources: application code, criticality in the service catalog, and Helm configuration are mapped to payments-api. Lunar collects and normalizes the evidence, evaluates the canary check, and provides PR feedback or configured blocking. Configuration evidence alone does not prove deployment behavior.

Be precise about what the check proves

Once a service is mapped, the next question is what the evidence actually says.

Finding a canary setting in its Helm repository shows only what that configuration declares. It does not prove that the deployment used the expected rollout behavior. If that distinction matters to the requirement, the team needs evidence from deployment as well.

Now the recurring job is clear: gather evidence from the right places, connect it to the right service, evaluate the rule when things change, and give developers a result they can understand. Doing that once for a handful of services is manageable. Keeping it working as teams and pipelines change is where platform time disappears.

This is where Earthly Lunar comes in to help. Its GitHub/GitLab integration and CI agent collect repository and build signals into a component view. Guardrails evaluate that evidence and can report findings in pull requests or block a deployment when the rule is ready. For this split-repository rule, the platform team still has to define how the application repo, Helm repo, and criticality record map to the same service.

With those decisions made, the platform team has a repeatable way to run the check, and a sound basis for deciding when a finding should warn or block.

Turn up enforcement when the evidence is ready

A new guardrail often exposes a mislabeled service or a configuration file the check cannot read. It may also uncover an exception that was never documented.

Blocking every change on day one would turn those discoveries into a queue of frustrated engineers.

A more workable rollout starts by showing the results to the platform team. Findings can then appear in the pull request where a developer can act on them. In a large engineering organization, that direct feedback is more likely to reach the right person than another broad announcement. Once the team trusts the service mapping and the check’s meaning, it can decide which failures should block a merge or a release. A missing canary on a critical production service may deserve a different response from the same finding on an experimental service.

Exceptions need design too. Who can approve one? Does it apply to a particular change or an entire service? When does it expire? Who reviews what was waived? At 2 a.m., engineers need a path through an urgent release, and the platform team needs a record of the decision.

Lunar supports gradual enforcement, from visibility and PR feedback through blocking checks. The platform team still owns the judgment about when to move between those stages.

Turn a requirement into a guardrail

A second standard might require an SBOM for every release.

A workflow file that names an SBOM tool shows intent. Proving the requirement means finding a published SBOM tied to the artifact being shipped. Lunar’s CI agent can collect build signals from instrumented self-hosted runners, while the platform team defines what counts as publication and tracks where evidence is missing.

Canary configuration shows intent; deployment evidence must show canary behavior for the service, environment, and release. Naming an SBOM tool shows intent; a published SBOM must be tied to the shipped artifact. Missing evidence is a coverage gap, not a pass.

Lunar’s 200+ built-in guardrails cover common checks. For the canary rule or any unique rules or standards within a given organization, Earthly’s AI agent skills can draft a collector and policy from plain language. The team still defines the service map and what counts as proof, then reviews the code before enforcement.

When the next service appears

The first successful canary check is cool. But the more exciting moment comes a few weeks later. A team creates a new critical service. Its application code lands in one repo and its Helm configuration in another. Will the service enter the canary standard? If the result is missing, will anyone know why?

The test is whether the new service enters the process: map it to its code and deployment configuration, collect the relevant evidence, and show its team the result. If it has not been mapped, that gap should be visible rather than disappear from the results.

If a standard depends on information scattered across repositories and internal systems, bring it to a Lunar demo. We can walk through the actual workflow and what it would take to enforce it.

Kate Gooch
Kate Gooch is on Earthly’s GTM team and has a soft spot for developer tooling. She’s worked across DevOps, IDPs, and AppSec. Engineering guardrails is the first category she’s worked in that doesn’t need an abbreviation or acronym, though she regrets to report that EGGs is available. Engineering Governance Guardrails. Finally, something you can enforce over easy.

Updated:

Published:

Get notified about new articles!
We won't send you spam. Unsubscribe at any time.