In this article
- Intelligence and job data are different problems
- Your AI can only work with the jobs you give it
- Sometimes you need control over the company universe
- Then you discover you are also building a crawler
- Good inputs make the AI layer easier to build
- The agent does not need to read every job
- Let the AI product own the intelligence
Building the AI is often the interesting part. Getting reliable jobs into the product is the part people underestimate.
A modern hiring AI product can do a lot. It can:
- Understand a CV
- Interpret skills and experience
- Rank jobs against a candidate
- Explain why a role is a good match
- Help with applications
But before it can do any of that, it needs jobs to work with: real, current jobs from the companies its users care about.
The model is only one layer of the product. Job-data infrastructure is another.
01Intelligence and job data are different problems
Large language models are very good at reading, reasoning and explaining. That makes them a natural fit for hiring products. But connecting an LLM does not automatically give you:
- The companies you want to cover
- Their current jobs
- Structured job records
- Ongoing discovery of new roles
- Removal of jobs that disappear
This is not a weakness of ChatGPT, Claude or any other model. They solve a different problem. A model provides the intelligence. Your product still needs a dependable source of jobs for that intelligence to work on.
02Your AI can only work with the jobs you give it
The jobs underneath a product quietly decide what it can be good at. A few examples:
- A product focused on UK startup jobs needs jobs from relevant UK startups.
- A product focused on visa sponsorship needs jobs from suitable employers, plus sponsorship information where a posting mentions it.
- A product for cybersecurity careers needs good coverage of cybersecurity employers.
If those companies are missing, the best ranking model in the world cannot recommend their roles. The quality of the AI experience depends partly on the underlying company and job universe.
03Sometimes you need control over the company universe
For some products, a generic job feed is a perfectly good fit. If you want broad coverage across many industries, it may be exactly what you need.
Other AI products need more control. A founder might want to build:
- An agent for Indian startup jobs
- An assistant for UK fintech roles
- A career tool for biotechnology graduates
- A product focused on sponsorship-friendly opportunities
- An internal tool covering a specific group of employers
In cases like these, the ability to choose the companies can be as important as the AI itself. It is not that generic feeds are bad. It is a different requirement.
04Then you discover you are also building a crawler
Many teams follow a similar path:
- Build the matching or AI experience.
- Realise it needs jobs.
- Start collecting jobs from employer websites.
- Discover different ATS platforms and custom career sites.
- Add extraction and normalisation.
- Add repeated checking for new and removed jobs.
- Spend more and more time maintaining job infrastructure.
We have written before about why job crawling is harder than general crawling and why keeping job data fresh takes ongoing work. The short version is that each step brings its own edge cases, and none of them stay solved for long.
At some point, a team building an AI career product can find itself maintaining two products: the AI product users see and the job-data infrastructure underneath it.
05Good inputs make the AI layer easier to build
An LLM can read a messy job page. That does not mean it should have to, every time. Structured fields give the rest of your system something solid to work with. A useful job record might include:
titleBackend EngineercompanyExample Fintech LtdlocationManchester, hybridsalary£55,000 to £65,000, where availableemployment_typeFull-timeexperience3+ years, where statedsponsorshipMentioned in postingdescriptionFull job description textNot every job will include every field. Many postings do not list a salary, and most say nothing about sponsorship. But when the information is there, structure makes it usable for:
- Filtering before an LLM is used
- Matching
- Ranking
- Recommendations
- Agent reasoning
- Reducing unnecessary model calls
Good inputs do not replace the model. They let the model spend its effort on the parts that actually need intelligence.
06The agent does not need to read every job
It is tempting to picture an AI agent looking through every job and picking the best ones. In practice, it does not need to load thousands of jobs into its context. A search layer can retrieve a smaller, relevant set first.
The search handles what search is good at: location, job type, keywords and other structured filters. The AI handles what it is good at: understanding the candidate, weighing the trade-offs and explaining its reasoning.
This can create a cleaner architecture than asking an LLM to work directly over a huge raw dataset. Each layer does one job, and each is easier to test and improve on its own.
07Let the AI product own the intelligence
Recuity is the job-data layer underneath your product. You choose the companies or market you want to cover. Recuity handles:
- Finding jobs from company career sites
- Structuring those jobs into consistent records
- Repeatedly checking for new and removed jobs
- Maintaining the searchable job index
You can then use the Job Search API, JSONL data exports, and your own AI models and workflows on top.
- AI agent
- Matching
- Recommendations
- User experience
- Career-site crawling
- Job extraction
- Job index
- Ongoing job discovery
The distinction matters. Recuity supplies and maintains the job layer. It does not provide your AI agent, your candidate matching or your recommendation logic. That is the part you build, and the part that makes your product yours.
Build the intelligence. Let someone else handle the job layer.
AI makes it possible to build much richer hiring products than before: agents that understand candidates, explain their suggestions and help people act on them. But those products still need dependable information about real jobs.
A strong hiring AI product needs both intelligence and reliable job data. You can spend your time building both, or focus on the part that makes your product different.
Recuity is the infrastructure for the job-data side. You choose the companies, and Recuity keeps their jobs collected, structured and current.