Member of Technical Staff - Research
Why Crucibl
Management consulting is a $400B industry built on selling intelligence-for-hire. The model is breaking - not slowly, but now. The biggest frontier in enterprise AI isn't better models. It's judgment at scale: taking the messy, ambiguous decisions that run 85% of the global economy and making them faster, crisper, and more defensible.
Crucibl is building that. We're profitable, growing fast, and delivering for Fortune 500 clients. We raised a $10M seed from Tier 1 VCs - then kept growing on revenue. The client pipeline is full. We're scaling to meet it.
What You'll Do
Push the Frontier of Judgment
Research how frontier models reason through ambiguous, high-stakes business decisions - and where they fail
Design novel methods for reasoning, evaluation, and calibration that go beyond standard benchmarks
Translate open problems in reasoning, uncertainty, and multi-step decision-making into approaches we can test and ship
Build Evaluation That Matters
Define what "good judgment" looks like for a model, and build the evaluation frameworks to measure it
Design experiments that reveal real failure modes, not just leaderboard scores
Turn findings into concrete recommendations for the product and applied teams
Partner with Founders
Shape technical vision and roadmap alongside the founding team
Bring outside research thinking into a company solving a problem few labs are focused on
Set the Bar
Define what rigorous, applied research looks like in an AI-first organization
Raise the bar for the team as it grows
Who You Are Must-Haves
At least one undeniable signal of excellence - published research at a top venue, research role at a frontier lab, or a track record of novel technical contributions
Deep understanding of how LLMs reason, fail, and can be evaluated - this isn't a theoretical interest, it's core to the job
Strong fundamentals in ML, statistics, or NLP - we're too small for hand-holding on the technical side
Comfortable moving from open-ended research question to a testable hypothesis quickly
Strong Signal
Experience designing evaluation frameworks or benchmarks for reasoning, decision-making, or agentic systems
Published or shipped work on uncertainty, calibration, multi-step reasoning, or LLM evaluation
Has operated in a high-growth, high-ambiguity environment
Thinks like an owner - big picture, not just your part
We're Not Looking For
Researchers who need a clean, self-contained problem before starting - ambiguity is the job
Pure benchmark-chasers - we care about judgment that holds up with real clients, not just leaderboard gains
Researchers without product instincts - your work needs to change what we build and ship, not sit in a paper -- but we are open to publishing work!
Who You'll Work With
Our founders bring deep technology experience (Google-scale systems serving hundreds of millions of users) and domain expertise in high-stakes business decisions from top consulting and private equity firms. You'll be the connective tissue between these worlds, and between the company we are today and the company we're becoming.
Benefits & Perks
Zero to Outcome: Seed stage means you shape the product, the culture, and the trajectory - not inherit them.
See Your Work Matter: No abstractions, no layers - you'll see exactly what you built and what it changed.
The People Around You: A small, elite team that will raise your game. We hire for exceptional, not just experienced.
Meaningful Equity: Every offer includes a comprehensive salary and equity package.
Hybrid Schedule: 3 days in-office with real flexibility around the rest. We care about output, not optics.
Full Health Coverage: Medical, dental, and vision.
Daily Lunch & Snacks: Fueled and focused, on us.
Work Authorization
Crucibl welcomes applications from candidates requiring visa sponsorship. Sponsorship eligibility is determined during the interview process.
Crucibl is an equal-opportunity employer. We hire on merit and outcomes, and welcome applicants of every background.