Building RankMetrics: daily rank tracking across 50k+ keywords

Rank tracking is easy for ten keywords and a different problem at fifty thousand. What we built so RankMetrics lands every night without falling behind.

Tracking one keyword is a script. Tracking fifty thousand every day, for hundreds of sites, and having yesterday's numbers ready before the client logs in, is a scheduling and storage problem that has very little to do with SEO.

RankMetrics is the SEO analytics platform we built and run. It does daily tracking across 50,000+ keywords and more than 100 technical audit checks. Most of the engineering is in the word "daily".

Why the naive version falls over

The obvious build is a cron job that loops through keywords and writes results. It works until one of these:

  • The window is fixed. Everything has to finish overnight. A loop that takes eleven hours at forty thousand keywords does not take twelve at fifty thousand — it misses the window and yesterday never lands.
  • Failures are normal, not exceptional. At this volume something always fails. If one bad response stops the run, you lose the whole night's data for everyone.
  • Storage grows faster than you plan for. Every keyword times every day times every site. A schema that stores the obvious thing becomes unqueryable within months.

What the build actually needed

Work as queued units, not a loop

Each check is its own job with its own retry and its own failure state. A failure loses one keyword for one day, visible in the interface as a gap rather than silently missing. The run is resumable, so an interrupted night picks up rather than restarting.

Backpressure, deliberately

The constraint is not how fast we can ask, it is how fast we may. Rate is a scheduling parameter rather than an accident of how many workers happen to be running, so the system stays inside its limits by design.

A schema built for the question

Rank data is written once and read as a trend. That is a very different access pattern from a normal application table, and designing for "show me this keyword's last ninety days" from the start is the difference between a chart that renders instantly and one that times out at year two.

Diagnostics next to the numbers

A ranking drop is not useful on its own. The crawl diagnostics — over a hundred technical checks — run alongside tracking so a drop can be read against what changed on the page. That pairing is the actual product; the number alone is a commodity.

Reporting people outside SEO can read

The hardest design problem was not technical. Most SEO tools report in language only an SEO understands, and their output is usually shown to someone who is not one.

So the reporting answers what moved, what caused it, and what to do — with the raw data underneath for anyone who wants it. That decision shaped the interface more than any performance work.

It runs our own work too

RankMetrics is live at rankmetrics.io. We use it on our own sites and on our clients', which is the only reason we notice the things that would otherwise stay in a backlog.

Running a data pipeline daily for years is a different discipline from building one that works once. If you have something that needs to run unattended and be right in the morning, that is automation work, and we do it because we have had to.

More write-ups