Findings and verdicts
Two verdicts and one fixed shape. Every finding names the value, then the token or component it should be, then the line it comes from.
Crocotaste comments only on what it can point at in your repository. If it cannot show you the line that makes something a rule, it says nothing.
Verdicts
Violation · a value that is exactly a token of yours written raw, a component of yours hand-rolled, or a recorded convention or decision broken.
Drift · a near miss. A color close to one of your tokens, a spacing value just off your scale, a token pair that does not reach the AA contrast floor.
The difference is certainty, not severity. A violation turns the check neutral and drift leaves it green, and neither blocks anything: findings are posted as a comment review, never request-changes, and a neutral check passes branch protection, so what ships stays your call.
What a finding looks like
**Violation** · `#4ade80` is `--color-success` — use the token (tokens.css:18)
Source: `tokens.css:18`
```suggestion
<div className="bg-success" />
```
<details>
<summary>🤖 Prompt for your coding agent</summary>
```
In src/app/page.tsx around line 12:
Replace bg-[#4ade80] with the design token --color-success (tokens.css:18).
```
</details>
Four parts, every time: the verdict and the message, the source line, a suggestion you can commit in one click, and a collapsed instruction for the agent that wrote the code.
A suggestion is always exactly one line. A suggestion block is a one-click commit, so a multi-line one would be a way for a reviewed repository to get code committed through us; anything that is not one safe line is dropped and the finding explains itself in words instead. When a raw colour matches several of your tokens at once (white is often the page, the inverse text and the text on the brand), the finding names each of them and offers no suggestion, because which one you mean is the role the colour plays, and a one-click guess would commit the wrong one. A colour in an SVG fill or stroke attribute, or an HTML bgcolor, is flagged the same way without a suggestion, since var() there does not work in every client, and so is a raw value inside a Sass or Less function such as darken(), math.div() or one your team wrote, which runs when the stylesheet compiles, before any var() exists. A plain CSS function such as calc() or a gradient keeps its suggestion. Raw values are only ever reported by these checks: an AI finding that cites a rule against raw values in your DESIGN.md is removed before it posts, so a hex written as text in a code sample is never reported as a colour. A raw value that is only the fallback inside var(--token, …) is never flagged, because the token is already there, and the same value repeated on one line is one finding, whose suggestion fixes every copy. A spacing or radius value exactly halfway between two steps of your scale is named with both and gets no suggestion either, for the same reason. A spacing value your DESIGN.md defines has no class of its own, so in a Tailwind repository the suggestion is Tailwind's default class for it when the value is one of its steps (16px is p-4), and none when your theme sets its own spacing unit or steps, or the repository has a Tailwind config that could.
A suggestion from the AI review has a second test: it has to be a bounded edit of the line it replaces. At least half of that line survives, it adds no URL the line did not already carry, and it is never offered on an import line. One that fails the test is dropped, and the finding keeps its message and its prompt for your agent.
Where findings come from
Two sources, and you can tell them apart by what they cite.
Checks that need no model compare a literal your pull request added against a value already in your repository: a raw hex where you have a color token, a Tailwind palette class (bg-green-500, bg-white) whose colour is a token of yours, a spacing value that is not on your scale, a font size that is not on your type scale (your --text-* or --font-size-* tokens), a bg-/text- pair from your own palette below 3:1. They cost nothing and they stay quiet unless they are certain.
A var(--token) with no fallback that nothing in your repository defines or uses, one or two letters from a token your stylesheets define, is a violation naming the real token, with it as the suggestion when only one is that close. A name that differs from yours only in its digits (--gray-12 beside --gray-1) is left alone, since a package's scale often sits beside your own, and the check stands down when a stylesheet builds names at compile time.
A box-shadow or z-index value exactly equal to one of your tokens (a --shadow-* or --z-*) is flagged with the token as the suggestion. Nothing close but different is, and a value two tokens share is left alone, since the role it plays decides which one is meant.
3:1 is the WCAG AA floor for large text and interface components, and it is the level this check can stand behind: it reads two class names and never the type size, so holding every pair to the 4.5:1 body-text level would flag correct buttons. Both the message and its agent prompt name 4.5:1 for body text, so the last part of that call stays yours.
These checks refuse a place before they refuse a value. A token's own definition is not a use of it, so a declaration line and a tailwind.config.* are not scanned for raw values at all. A value inside a comment or a content: string is not code. top, right, bottom, left and inset position one element against another rather than sitting on your spacing scale. A color matches a token only at the same alpha, so an 8% tint is never told to become the fill it is a tint of. And a variant like hover:bg-… is a different state, so it is never paired against the base text color.
A palette class is measured the way a hex is: exactly a token is a violation, close to one is drift, and the suggestion is the token's own utility (bg-red-500 becomes bg-danger, a hover: prefix kept). One your theme defines itself (--color-red-500 in your @theme) is your token and is never flagged, and a tint (bg-red-500/20) is left alone, since no token shares its alpha. A white or black class is measured only as a background; as text, a border or a fill it plays a role a value check cannot see. The check runs only in a repository that uses Tailwind (a Tailwind config, or a stylesheet that imports it or declares an @theme block); a Bootstrap bg-white is never measured.
The AI review judges what is left against your DESIGN.md and the parsed system, and it may report only three kinds of thing: a semantic token used for the wrong role, a hand-rolled element where a listed component of yours exists, or a violation of a listed convention or decision. Aesthetics, naming and wording are out of scope. It holds no design opinions; the inventory of your system is all it works from. It reads the lines you removed beside the ones you added, so a rule a removed line satisfied and its replacement does not is a finding on the replacement.
Every AI finding must cite a line that defines something in your references, or it is deleted before it reaches the pull request. That is enforced in code after the model answers, not asked of it in a prompt.
At most three AI findings per pull request, ever, not three per push. When more survive, the summary says how many were withheld.
How many arrive at once
At most 12 findings are listed per push, violations first, and both the summary and that push's review say how many are held back. They post on later pushes as the listed ones are fixed or dismissed, and the review that posts them says they were held back from an earlier push. A stylesheet that adds 200 raw hex values should get a readable review, not a wall of comments.
A finding is never posted twice. Each carries a stable marker that survives force-pushes and line shifts, so an unrelated edit above it does not turn one finding into two.
A finding you fix is marked fixed once the review can see the line is gone: its comment opens with ✅ Fixed in and a link to the commit, and collapses as resolved; the check drops it and the summary counts it as fixed. A line you edit that still breaks the same rule gets a new comment, and the old one collapses as outdated. So does an AI finding the review lets go without proof of a fix (it read the whole change and no longer reports it): outdated says nothing was fixed, only that the finding no longer stands. If the code brings a fixed finding back, its comment opens again. GitHub's "Resolve conversation" button stays yours to press: resolving a thread takes write access to your code, which Crocotaste never asks for. If your branch rules require conversation resolution before merging, resolve our collapsed threads with the rest. A finding you dismiss with @crocotaste ignore is marked resolved and stays in the summary as dismissed, counted by nothing the check decides.
To dismiss one instead of fixing it, reply with @crocotaste ignore <reason>. The same findings are published as JSON on the check run for agents: the check run and JSON.