ACTIVEagent

perf-sentinel

Measures whether a change moved performance, using an alternating A/B rather than a before-and-after.

Model
sonnet
Tools
ReadGrepGlobBash
Project
zckr-dev
Cost
Sonnet, but many Lighthouse runs — wall-clock expensive rather than token expensive.

WHEN IT RUNS

Measures web performance changes with a controlled A/B rather than guessing. Use when a change might affect Lighthouse, Core Web Vitals, or payload size, or when a performance number needs to be stated with confidence.

CONTEXT

The change under test, the repo's own Lighthouse and build tooling, and both arms of the comparison built from source.

MEMORY

None, and that is the point: a baseline taken twenty minutes ago is not a baseline, because machine load drifts over a session.

EVALUATION

Both arms, at least three alternating runs each, all numbers reported with no cherry-picking. Overlapping ranges are reported as noise rather than as a finding.

WHERE IT FAILS · 3

  • Confounded by a stray dev server or browser left running — the most common cause of a fake regression.
  • Compares uncompressed payload sizes, where repetitive markup shrinks ten to one and a large delta turns out to be about a kilobyte on the wire.
  • Misreads a scoring cliff: near a threshold, a small timing shift swings the score a lot and means much less than it looks like.

LESSONS · 2

  • A regression stated plainly is worth more than a green number nobody trusts.
  • If a speculative fix does not help, revert it rather than leaving the complexity behind.

CONNECTED · 1

SEE THE GRAPH

← ALL AGENTS