AI Code Reviewers vs. Guardrail Engines: Where each fits

8 minute read     Updated:

Kate Gooch
Kate Gooch

An AI reviewer catches a standards violation. The author fixes it, and the PR merges. Next week, another team introduces the same problem in a different repository.

For the platform team, each repeat raises a familiar question: how many more repositories need the same fix? The standard is already agreed. Yet engineers keep spending review time explaining it, correcting violations, and trying to work out whether other teams are following it.

Take a company standard that prohibits third-party CI dependencies from referencing a moving branch such as main. The reviewer spots docker/build-push-action@main, and the author pins the action. Across the organization, teams use custom workflows, override shared templates, and follow different conventions. How does the platform team know that requirement is being checked everywhere it applies?

AI reviewers help engineers keep up with PR volume and uncover problems they hadn’t anticipated. Guardrail engines complement that work by turning agreed standards into repeatable checks.

This article explores what it takes to move from an individual review finding to consistent enforcement across repositories, and why collecting evidence from different workflows is central to making that work.

TL;DR

  • Use AI reviewers and guardrails together. Reviewers help uncover problems, including risks the team hasn’t anticipated. Guardrails provide repeatable enforcement of defined engineering standards.

  • Enforcing a standard across different workflows requires comparable evidence. Data from repositories and pipelines must be collected and normalized before a platform team can check adherence across the organization.

  • Enforce throughout the development lifecycle. Run checks during AI authoring, on pull requests, and before deployment, as the evidence each check needs becomes available.

Was the standard checked?

As PR volume grows, AI reviewers help engineers keep up. They can catch problems a busy human reviewer might miss, problems that are hard to describe as a rule, or issues that should become a rule, but aren’t yet.

For example, an AI reviewer might flag a retry path that could charge a customer twice. Investigating that risk means examining timeout behavior, the payment provider, and the service’s idempotency handling. The reviewer has surfaced something worth investigating.

Other requirements are already settled: which CI references are allowed, which tests must run, or which services need a runbook. The platform team needs those requirements checked on every applicable change.

Giving a reviewer the right instructions takes ongoing work. A company standard changes, but a service’s AGENTS.md still describes the old convention. The AI reviewer follows outdated guidance. Even an up-to-date file may leave out details that determine which requirements apply, such as the service’s criticality or compliance scope.

Putting 100+ requirements in REVIEW.md or AGENTS.md adds another challenge: the model still has to interpret and apply them consistently. GitHub’s documentation on customizing Copilot code review cautions that Copilot may not follow every instruction every time, and that long instruction files can cause instructions to be overlooked. When a review is silent, did the requirement pass, get missed, or lack the data needed to evaluate it?

Cost and access also shape how deeply an automated reviewer can investigate. Running a frontier model on every PR adds cost, while deeper investigation may require broader access than the team wants to grant.

Individual review comments leave the platform team with another job: determining whether the same standards hold across repositories. Freeform findings are difficult to compare across teams, especially when their pipelines and conventions differ. Consistent enforcement requires comparable evidence and a defined result for each check, including what happens when evidence is missing.

Some review products also include static analyzers or quality gates. Assess those checks separately from their model-generated comments: which requirement does each check evaluate, and what does a passing result establish?

Make a pass mean something specific

In the opening example, an author replaces docker/build-push-action@main after a review comment. The next repository needs the same check. Suppose the organization prohibits third-party CI dependencies from using reference names such as main, master, and latest. A policy can check those names directly wherever the relevant dependency data is collected.

A guardrail engine can check collected CI dependency data against a list of prohibited reference names (for example, Earthly Lunar has the no-mutable-refs guardrail). With the same policy and inputs, the check produces the same result. The author can fix the reference and rerun it.

A passing result means none of the collected third-party references matched the prohibited names. That is a specific claim the team can verify; proving every dependency is immutable would require a broader check. If dependency data is absent, this policy skips the check, so the team needs to distinguish missing evidence from a pass.

Before making the policy block PRs, the platform team can test which references it accepts and rejects, validate the collected data, and decide how to handle missing evidence. Once the check is defined, the next challenge is applying it across the organization.

The difficult part is coverage

Writing the reference rule is straightforward. Applying it across an organization depends on collecting comparable evidence from the workflows teams actually use.

Even repositories using the same CI platform can differ. One team declares third-party actions directly in its workflow. Another calls a reusable workflow that contains those dependencies. Others add wrappers, override shared templates, or maintain custom build steps. The standard applies across those repositories, but the evidence needed to check it may live in different places.

A check placed in a shared template reaches the repositories that adopted that template and kept it current. The platform team still needs to account for the repositories outside that path. Otherwise, a successful rollout can leave gaps: the standard is enforced for some teams, while others continue introducing the same problem.

That makes data collection and normalization part of the enforcement problem. A guardrail engine needs to gather the relevant signals from different workflows and represent them consistently, so the same policy can evaluate the same requirement across services. Defining a common rule is only useful if the system can supply the evidence that rule needs.

With comparable evidence, the platform team can see which services meet the standard, which violate it, and where collection needs attention. Teams keep their own workflows, while the organization gains a consistent way to check the outcome.

At every step of the SDLC

Engineering standards need checks at different points in development. A prohibited workflow reference can be caught while someone edits a configuration file. A test-result requirement has to wait until the tests run. Each check needs to run when the evidence it depends on becomes available.

Guardrail engines evaluate those signals against defined requirements. Earthly Lunar, for example, uses agent hooks to check an AI agent’s work during authoring. With the relevant collector configured, a hook can flag a prohibited workflow reference and feed the result back to the agent, giving it a chance to fix the problem before opening a PR. Checks at PR and deployment gates can then report or block violations using the evidence available at those stages.

The requirements apply to human-written and AI-generated changes alike. An AGENTS.md file helps explain conventions to an agent. Guardrails check whether the work meets the applicable standards, while AI reviewers investigate implementation risks such as the payment retry path described earlier.

Together, they cover requirements the organization has already defined and problems it hasn’t yet thought to check for.

Turn the repeated comment into a maintained check

Earthly Lunar is a guardrail engine that collects and normalizes data from repositories and CI/CD, then evaluates it against centrally managed engineering standards. Those checks run during AI authoring, on pull requests, and before deployment, as the evidence they need becomes available.

Its no-mutable-refs guardrail is one example: it checks collected third-party CI references against a defined list of prohibited names. Lunar also provides 200+ prebuilt guardrails and AI skills for building collectors and policies.

For an organization-specific standard, the platform team can define the evidence it needs, test the check on a limited scope, and review the results before enabling blocking. The team maintains the policy centrally as standards and workflows change. The next time a repository adds the prohibited workflow reference, an established guardrail can catch it.

AI reviewers continue investigating unexpected risks, while Lunar gives the platform team consistent enforcement of agreed standards across repositories and throughout the SDLC.

Bring a standard that’s hard to verify across your repositories and internal systems to a Lunar demo, and we’ll walk through the evidence needed to check it and how enforcement would fit into your workflows. To start identifying guardrails for your own organization, share an engineering standards doc, AGENTS.md, or postmortem to explore which requirements could become automated checks.

Kate Gooch
Kate Gooch is on Earthly’s GTM team and has a soft spot for developer tooling. She’s worked across DevOps, IDPs, and AppSec. Engineering guardrails is the first category she’s worked in that doesn’t need an abbreviation or acronym, though she regrets to report that EGGs is available. Engineering Governance Guardrails. Finally, something you can enforce over easy.

Updated:

Published:

Get notified about new articles!
We won't send you spam. Unsubscribe at any time.