ACTIVEagent
perf-sentinel
Measures whether a change moved performance, using an alternating A/B rather than a before-and-after.
- Model
- sonnet
- Tools
- ReadGrepGlobBash
- Definition
- .claude/agents/perf-sentinel.md
- Project
- zckr-dev
- Cost
- Sonnet, but many Lighthouse runs — wall-clock expensive rather than token expensive.
WHEN IT RUNS
Measures web performance changes with a controlled A/B rather than guessing. Use when a change might affect Lighthouse, Core Web Vitals, or payload size, or when a performance number needs to be stated with confidence.
CONTEXT
The change under test, the repo's own Lighthouse and build tooling, and both arms of the comparison built from source.
MEMORY
None, and that is the point: a baseline taken twenty minutes ago is not a baseline, because machine load drifts over a session.
EVALUATION
Both arms, at least three alternating runs each, all numbers reported with no cherry-picking. Overlapping ranges are reported as noise rather than as a finding.
WHERE IT FAILS · 3
- Confounded by a stray dev server or browser left running — the most common cause of a fake regression.
- Compares uncompressed payload sizes, where repetitive markup shrinks ten to one and a large delta turns out to be about a kilobyte on the wire.
- Misreads a scoring cliff: near a threshold, a small timing shift swings the score a lot and means much less than it looks like.
LESSONS · 2
- A regression stated plainly is worth more than a green number nobody trusts.
- If a speculative fix does not help, revert it rather than leaving the complexity behind.