One Lighthouse run tells you almost nothing

A performance score is one sample from a distribution, not a measurement. Here is how to read it properly, and what the spread between runs is actually telling you.

Somebody sends you a screenshot. A number in a circle, green, usually 100. Sometimes it is an agency pitching you, sometimes it is your own developer, occasionally it is a competitor’s site being used to make a point.

That number is one sample. Run the same test on the same page five minutes later and you can get a different one. Not slightly different. I have watched the same page, same tool, same hour, no changes between runs, score anywhere from 79 to 100.

Why it moves

The test is not measuring your site in isolation. It is loading your site on an emulated mid range phone, over an emulated slow connection, on shared infrastructure somewhere, and timing what happens.

  • The emulated connection is throttled, so anything competing for bandwidth changes the order things arrive in.
  • The machine running the test is shared, so it is sometimes busier than others.
  • Caches at the edge are warm or cold depending on what happened before.
  • If anything on your page is a race between two resources, you win some runs and lose others.

That last one is the interesting case, and it is why the spread is worth reading rather than averaging away.

Run it five times, not once

This is the whole practical instruction and it takes four extra minutes. Run the test five times, write down all five scores, and look at the shape.

  • Five scores clustered within a few points. That is your real number. Trust it.
  • A wide, even spread. Usually network or infrastructure noise. Take the median and move on.
  • Two clusters with a gap in between. Something on the page is a race, and it is losing some of the time. This is a finding, not noise, and it is fixable.

The third shape is the one people throw away by averaging. A page that scores 85 six times and 100 five times does not have an average problem. It has one specific thing that sometimes arrives late.

The trick that actually finds it

Put a fast run and a slow run side by side and compare the individual metrics, not the score. The metrics that stay identical tell you where the problem is not, which narrows things far faster than looking at the slow run on its own.

  • First paint identical, largest paint moves. The page starts drawing at the same moment every time, and one element lands late. Look at that element, not at the whole page.
  • Blocking time moves but paint does not. Script execution. Something is occupying the main thread, usually a third party tag.
  • First paint itself moves. Server response or something blocking the render, which is a different and usually easier problem.
  • Layout shift appears only sometimes. Something is loading without reserved space, and it only pushes things around when it arrives late.

I found a two and a half second problem on a homepage this way. First paint, blocking time and layout shift were byte for byte identical between the fast and slow runs. Only the largest paint moved, and it moved a lot. That ruled out the server, the scripts and the layout in one comparison, and left exactly one thing to look at: whatever was drawing the biggest element on the page.

Lab numbers and real numbers are different things

Worth knowing, because it changes how much any of this should matter to you.

The score you are being shown is a lab test. A simulated phone on a simulated connection, run once, on demand. Separately, browsers report what actually happened to real people on your site, and that field data is what search engines use when performance feeds into ranking.

A site with enough traffic has both. A newer or quieter site often has only the lab number, because there is not enough real world data yet to report. In that situation the lab score is being used to judge you by humans, not by Google, which is a reason to care about it but a different reason than the one usually given.

What the score is genuinely good for

  • Catching a regression. If it was consistently 95 and is now consistently 70, something you did caused that, and you know roughly when.
  • Finding the specific audit that names your problem. The list underneath the score is worth more than the score.
  • Comparing a page to itself over time.

What it is not good for is comparing two different sites, because a brochure page and a store with a thousand products are not attempting the same thing. It is also not worth chasing from 92 to 100 on most sites. The difference between 70 and 92 is a visitor noticing. The difference between 92 and 100 is a number changing colour on a chart.

What to ask when somebody shows you a 100

Two questions, and they are not hostile ones.

How many runs was that, and can I see the metric values. Anybody who measures properly will have the answer immediately, because they ran it more than once and they looked at the numbers underneath. Anybody who ran it until it went green will not.

It is the same instinct as reading any other number somebody hands you, and the same one behind what a caching plugin actually hides, where the measurement looks excellent precisely because it is measuring the wrong thing. If you want the plain English version of which of these metrics matter and which do not, that is here.

Questions

If you have been handed a performance report and cannot tell whether it means anything, that is usually a fifteen minute answer rather than a project. It is also the first thing I do on any build or rebuild, because you cannot tell whether you improved something you never measured properly. Send me the report and I will tell you what it actually says.

Written by Sean Lee, Palm Projects

I build and rank websites for small businesses. If something here applies to your site and you want a second opinion on it, send it over.

Start here

Tell me what you are trying to fix

Send over the site you have now, or the one you wish you had. I will tell you honestly whether I am the right person for it.

info@palmprojects.com