Firecrawl · Research · Unspecified · Posted 2026-07-12
Research Engineer - Evals
Firecrawl · San Francisco HQ; Toronto Hub · $250k–290k base
This range's midpoint is above 52% of posted research ranges at AI companies right now. See the salary index.
Apply on Firecrawl's site Watch Firecrawl for new roles
RESEARCH ENGINEER - EVALS
You'll build the evaluation systems that tell us whether Firecrawl actually works. That sounds simple. It isn't. Our core promise, convert any URL into clean, structured, LLM-ready data reliably, is hard to measure rigorously across millions of different websites, formats, and edge cases. As the systems we're measuring get more complex, the question "did that work?" gets harder, not easier.
This isn't an eval role where you inherit a framework and run benchmarks. You'll design the metrics, build the pipelines, generate the datasets, and own the feedback loop from output quality back to model and product decisions. If you care about what "good" actually means and have the engineering depth to measure it, this is the role.
Salary Range: $250,000–$290,000 USD/year (SF) / $210,000–$224,000 CAD/year (Toronto)
Equity Range: Competitive equity. Details shared during the process.
Location: San Francisco, CA (SF HQ) or Toronto, ON (Toronto Hub). On-site, five days a week.
Job Type: Full-Time
Experience: 4+ years in ML, research engineering, or data-heavy backend, with real evaluation work
Work Authorization: Must be authorized to work in the United States or Canada. We're not able to sponsor US visas right now. For Canada, we'll consider sponsorship on a case-by-case basis through our Toronto Hub.
ABOUT FIRECRAWL
Firecrawl is the easiest way to turn the web into data AI agents can use. One API call converts any URL into clean, LLM-ready markdown or structured data. It's the boring-hard problem everyone building with LLMs eventually hits, solved.
In September 2026 we raised a $75M Series B led by Smash Capital, and we're spending it building the largest repository of knowledge in the world. We hit 8 figures in ARR in year one and more than doubled it in year two. We have 183k+ GitHub stars, putting us in the top 50 repositories of all time, and developers, agents, and category-defining AI companies build on us every day. Growth like this is rare, and we're just getting started.
We're a small team punching far above our weight, working out of SF HQ and our new Toronto Hub. Everyone here owns a real piece of the product and company, end to end, and runs it themselves. No hiding behind process or headcount.
This is a place for people who want to work at the frontier: an AI company building the infrastructure other AI companies run on, not one bolting AI onto an existing product. We move fast, go deep, and are building the tools superintelligence will rely on to gather data from the web. That library is called Alexandria, and it starts now.
WHAT YOU'LL DO
- Design the metrics that define what "good output" actually means across millions of sites, formats, and edge cases
- Build the pipelines and harnesses that measure quality rigorously and at scale
- Generate and curate the datasets that make evaluation trustworthy
- Own the feedback loop from output quality back to model and product decisions
- Turn "did that work?" into an answer the whole team can act on
WHAT WE'RE LOOKING FOR
- You have the engineering depth to build real evaluation systems, not just run existing ones
- You care deeply about what "good" means and how to measure it rigorously
- You're comfortable owning ambiguous problems where the metric itself has to be invented
- You move fast and close the loop. You'd rather ship, measure, and iterate than perfect on paper
WHAT WE'RE NOT LOOKING FOR
- Someone who only wants to run benchmarks someone else designed
- A pure researcher who won't build the systems, or a pure engineer who won't think about methodology
- Someone who needs a fully-specced ticket to start
A NOTE ON PACE
We operate at an absurd level of urgency because the window for what we're building won't stay open forever. If that excites you, keep reading. If it doesn't, no hard feelings, but this role probably isn't for you.
BENEFITS & PERKS
AVAILABLE TO ALL EMPLOYEES
- Salary that makes sense: $2 …
See also: Research Engineer jobs · AI jobs in San Francisco Bay Area · Firecrawl salaries · LLMs jobs · Evals jobs · Agents jobs.
This listing is reproduced from Firecrawl's public careers feed and links to the original. AI Hiring Index is not the employer and does not accept applications. All Firecrawl roles · AI salaries.