In this article
Crawling a careers website sounds simple. Find the links, open the job pages and extract the information. A few hundred lines of code and you have jobs.
For one website, that is often true. You look at the page, work out where the jobs live and write something that pulls them out. It works, and it feels like the hard part is done.
The trouble starts when you need it to work across hundreds or thousands of different employer websites, every day, without someone checking each one by hand. That is where job scraping stops being a script and starts being infrastructure.
This article walks through the problems that show up along the way, and why a general-purpose web crawler usually only covers the first part of them.
01Finding pages is not the same as finding jobs
A general-purpose crawler is designed to discover pages. You give it a starting URL, it follows links and it returns what it finds. That is exactly what it should do.
A job crawler has a narrower and harder question to answer: which of these pages are actually jobs?
A typical company website contains far more than job listings. It might have:
- News and press releases
- Blog posts
- Employee and leadership profiles
- Office and location pages
- Policies and legal pages
- Product documentation
Larger companies can have thousands of unrelated links. If the crawler treats them all equally, it spends most of its time and budget on pages that will never contain a job. The real work is reaching the careers area quickly and staying focused once you are there.
02Career websites behave differently
Even once you find the careers section, there is no standard for how it works. Some common patterns we see:
- JavaScript rendering. The initial HTML is almost empty and the jobs only appear after scripts run. A crawler that reads raw HTML sees nothing.
- Jobs behind interaction. Listings only appear after clicking a department, choosing a country or pressing a "View roles" button.
- Pagination and load more. The first page shows ten roles and the other ninety sit behind a button or infinite scroll.
- Custom careers sites. Many employers build their own pages with unique layouts, filters and URL structures.
- Split domains. The main site is on one domain, careers on a subdomain and the actual applications on a third party domain.
- URL explosions. Filters and search parameters can generate thousands of URLs that all show the same handful of jobs.
- Navigation loops. Links that point back to each other in ways that keep a naive crawler busy forever.
None of these are unusual on their own. The difficulty is that every employer combines them differently, and your crawler has to cope with all of them without a human adjusting it for each site.
03Not every employer uses a standard job feed
Some employers use established applicant tracking systems. These platforms tend to have predictable layouts and sometimes offer a feed or public endpoint, which makes life easier.
Many others publish jobs directly on their own websites. Their pages may not include structured job metadata, and there may be no feed at all. The job exists only as text and layout on a page designed for a human to read.
So even if you support the common ATS platforms well, a large share of employers still need a different approach. A job data system has to understand the page itself and turn it into a consistent job record, whatever technology sits behind it.
04Finding the page is only half the problem
Once you have a job page, you still need the useful information out of it. For most hiring products that means:
- Job title
- Location
- Description
- Salary, when available
- Employment type
- Contract details
Every company presents these differently. One puts the location in a heading, another buries it in the third paragraph. Salary might be a range, a single number, a daily rate or missing entirely. "Remote" can mean fully remote, remote within one country or two days a week in an office.
A general-purpose crawler will typically give you the HTML, markdown or plain text of the page. That is useful, but it is not job data. You still have to build the extraction and normalisation layer that turns thousands of different page formats into one consistent schema.
A page is not a job record. Something has to turn one into the other.
05Job data changes
Collecting a job once is not enough. New roles are posted, existing ones are filled and removed, and descriptions get edited. A job index that is not maintained goes stale quickly.
That means revisiting employer websites on a schedule, spotting new jobs, updating changed ones and noticing when a job is no longer listed. Each of those needs its own logic, and all of them need to run reliably across every company you track.
It is worth being honest about the limits here. No system can always know for certain that a job is closed. If an employer removes a role from their listings but leaves the old page online, it can still look available. The best you can do is check regularly and use the signals the site gives you.
06Where a general-purpose crawler still makes sense
None of this means general-purpose crawlers are bad tools. They are very good at what they are built for. They make sense when:
- You want arbitrary web content, not one specific type of data
- The websites you target vary widely in purpose
- You want raw page content to process in your own way
- You are prepared to build and maintain the job-specific logic yourself
If that describes your project, a general crawler may be exactly the right choice. The point is simply that job data needs additional specialised infrastructure on top of crawling: discovery, extraction, normalisation and ongoing maintenance.
07What Recuity does differently
Recuity is built around the job data problem rather than general web crawling.
You choose the companies you care about. Recuity handles the infrastructure needed to turn their career websites into maintained, structured job data. That includes:
- Discovering where each company publishes its jobs
- Reading custom and dynamic career sites, not only standard ATS pages
- Structuring each listing into a consistent job record
- Revisiting tracked companies so the data stays current
- Making the resulting jobs searchable
You can query the data through the Job Search API or take it as a JSONL export into your own systems.
| General-purpose crawler | Recuity |
|---|---|
| Finds web pages | Finds job listings |
| Returns page content | Returns structured job records |
| Customer builds job extraction | Job extraction is part of the system |
| Customer manages recrawling logic | Tracked companies are revisited |
| Customer builds job search | Job Search is available |
Build the hiring product, not the crawling infrastructure
If your goal is simply to crawl websites, a general-purpose crawler may be exactly what you need.
If your goal is to build a job search, job board or hiring product, there is considerably more infrastructure between finding a webpage and having reliable job data. Most teams only discover how much once they are already maintaining it.
Recuity handles that job data layer so teams can focus on the product they are building.