Blog/Job data

Why Collecting Jobs Once Is Not Enough

Getting jobs into a database is the easy part. Keeping the set of available jobs current, as employers publish new roles and remove old ones, is the work that never stops.

Job dataRecuityOctober 20266 min read
In this article
  1. Job data is not static
  2. Stale jobs are more than a data problem
  3. Refreshing jobs means knowing what is new and what is gone
  4. Hundreds of companies means ongoing maintenance
  5. Better data makes the rest of the product better
  6. Recuity maintains the job index for you

Getting jobs into a database is only the beginning.

A one-time crawl gives you a snapshot of the jobs available at that moment. As employers publish new roles and remove old ones, that snapshot quickly becomes outdated. Career sites also move or get redesigned, which can quietly stop new jobs from being found.

A job product that people can rely on needs maintained job data, not a one-time scrape. This article explains why those are two different problems, and why the second one is usually the harder of the two.

Job data is not static
New jobs
Career site checked
New jobs discovered
Added to index
Existing jobs
Previously indexed jobs checked
Still foundKeep in index
No longer foundRemove from index
Every check does two things: adds jobs that are new, and removes jobs that are no longer found on the source.

01Job data is not static

Most web content has a long shelf life. A product page or a news article can sit unchanged for years. Jobs are different. Many roles are open for a few weeks. Some are filled in days.

That means a job board can collect thousands of roles today and have a noticeable share of them out of date by the end of the week. Nothing went wrong with the crawl. The world just moved on.

A few ordinary things that happen all the time:

  • A company fills a role and takes the listing down on Tuesday
  • A graduate scheme closes on its deadline and every listing for it disappears
  • A company opens a new office and adds a batch of local roles
  • A team reopens a role it closed last month, under a new URL
  • A startup posts ten new jobs on the day it announces a funding round

None of this shows up in a snapshot. It only shows up if you keep looking.

02Stale jobs are more than a data problem

It is easy to think of stale jobs as a cleanliness issue in the database. For the people using the product, it feels very different.

A candidate who finds an interesting role, reads the description, gets halfway through a cover letter and then clicks through to a "this job is no longer available" page has wasted real time. If it happens twice, they start to doubt every listing on the site.

Stale data shows up for users as:

  • Clicking jobs that are no longer available
  • Missing new roles that were posted after the last crawl
  • Seeing the same old jobs every time they come back
  • Getting alerts for roles that closed weeks ago

The result is a loss of trust, and trust is hard to win back. Candidates rarely complain. They simply stop returning.

Users do not see your database. They see whether the job they clicked is still real.

03Refreshing jobs means knowing what is new and what is gone

The obvious fix is to crawl every site again. That is part of the answer, but only part.

A second crawl gives you a second snapshot. The useful work is comparing it with what you already have. For each company, the system needs to work out:

  • Which jobs are new?
  • Which previously indexed jobs are no longer present?
  • Has the careers site changed structure?
  • Are there new job pages to discover?
  • Is the crawler still finding the employer's current vacancies?

The last few questions matter more than they look. If a site is redesigned and extraction quietly fails, a naive refresh might remove every job from that company, or simply stop finding new ones. Telling a real drop in vacancies apart from a broken crawl is a large part of keeping the index accurate.

A note on closed jobs

Detecting closed roles has limits, and it is worth being honest about them.

If a job disappears from the employer's career site, that can be detected, and the job can be removed. But if an employer leaves an old job page online after the role is filled, the system may still see it as available. The source is the signal. When the source is out of date, the data can be too.

04Hundreds of companies means ongoing maintenance

With a handful of companies, all of this is manageable by hand. The problem changes shape as the list grows.

Imagine you are tracking 300 companies. The initial crawl works. Everything looks good. Then, over the next few months:

  • One company changes its careers URL
  • One switches to a different ATS
  • One rebuilds its careers page with JavaScript rendering
  • One changes how pagination works
  • One moves its jobs to a separate domain

Each change is small. Each one can silently break collection for that company. And at 300 companies, there is almost always something like this happening somewhere.

300 companies × repeated checks
  • Careers URL changed
  • Switched ATS
  • Added JavaScript rendering
  • Changed pagination
  • Moved jobs to another domain
→
One maintained job indexStill accurate because every source keeps being checked, even when it changes.
A few changed career sites in any given month is normal. Each one needs to be noticed and handled.

This is why the initial crawl is not the real infrastructure challenge. Keeping everything working, week after week, across hundreds of sites that change without warning, is.

05Better data makes the rest of the product better

Freshness does not only affect the job list. Everything built on top of the data inherits its quality:

  • Job search
  • Recommendations
  • Alerts
  • Matching
  • AI workflows

A good matching model that recommends a closed role is still a bad recommendation. An AI assistant that suggests a role which closed weeks ago gives a confident, wrong answer. If the underlying job data is poor, the intelligence built on top is less useful, however well it is designed.

06Recuity maintains the job index for you

This is the part Recuity is built to handle. Once a company is being tracked, Recuity regularly revisits its careers site. Newly discovered jobs are added to the index, and previously indexed jobs are removed when they are no longer found on the source site.

That means:

  • Discovering newly published jobs
  • Adding them to the index
  • Removing jobs that are no longer found on the source
  • Keeping the index searchable throughout

The same limitation described above applies. If an employer leaves an old job page online, Recuity may continue to see that job as available.

The maintained jobs are available through the Recuity workspace, the Job Search API, or as a JSONL export if you prefer to load the data into your own systems.

Fresh job data is a process, not a download

A one-time scrape gives you a snapshot. It is accurate for a moment and then starts to age.

Job freshness is not just collecting jobs once. It means repeatedly checking employers so new jobs enter the index and jobs that disappear from the source can be removed.

That maintenance layer is what Recuity is designed to handle.

Keep your job data current

Choose the companies you want to track. Recuity keeps revisiting them and maintains the index.

Explore Job Tracking & Data →