Clinical Product Specialist — AI Safety & Validation
The Opportunity
Nursing is where patient care happens. Nurses know what patients need, what works in real clinical environments, and how AI can actually help. At Hippocratic AI, we're building AI systems that earn the trust of nurses and clinicians because they're designed with clinical reality in mind from the start. You'll be the clinical voice shaping that design—ensuring our systems are accurate, safe, and useful to the people delivering care.
At Hippocratic AI, we're different. We're building the only generative AI platform designed for safe, autonomous clinical conversations with patients. Our 99.9%+ accuracy isn't accidental—it's earned through rigorous clinical validation by experienced nurses and clinicians who understand real-world care.
We need a Clinical Product Specialist (RN) who can bridge the gap between clinical reality and AI capability. You'll be the voice of frontline nursing in our product and ML teams, working alongside leaders to validate, test, and continuously improve our medical large language models (LLMs). You'll execute clinical validation protocols, author test cases and gold standards, and provide clinical insights that shape product decisions. You'll travel 25-40% to health systems and user research—staying connected to real clinical workflows and learning how our platform performs in the field.
This is hands-on, high-impact work where you'll grow your product muscles while making clinical safety your specialty. You'll own execution. You'll contribute strategic thinking. And you'll be part of a team that treats clinical integrity as our most important feature.
What Success Looks Like in Year One
Validation Protocol Execution: You've successfully executed 2-3 full clinical validation cycles—from protocol design (with guidance from your lead) through systematic testing to documentation and recommendations for improvement.
Gold Standards Library: You've authored or co-authored 30-50 test cases, rubrics, and reference answers for clinical scenarios. You understand the difference between a good test case and a great one. Your work is used by engineering teams to benchmark model performance.
Clinical Gap Identification: Through hands-on testing and analysis, you've identified 5-10 clinical inaccuracies, edge cases, or safety concerns in the model. You've documented them clearly and worked with senior team members to prioritize fixes.
Product Contribution: You've participated in 3-4 product discussions where your clinical input changed a decision—whether that's safety rail design, dialog flow changes, or escalation criteria. You're comfortable speaking up about clinical concerns in technical rooms.
Health System Connection: You've visited 2-3 health systems to observe real patient interactions with our platform. You've gathered direct feedback from nurses and patients. You've synthesized observations into clear clinical insights for the team.
Learning in Public: You've grown your understanding of LLMs, clinical validation frameworks, and product development. You can explain model capabilities and limitations to other clinicians. You're asking better questions and connecting clinical problems to technical solutions.
Quality Mindset: You've developed a rigorous approach to clinical testing. You know what "good" looks like. You've caught issues before release. You take ownership of clinical accuracy and safety.
Your Core Responsibilities
Clinical Validation & Testing
Execute clinical evaluation protocols designed by senior team members. Conduct systematic testing of model behavior against high-stakes scenarios (medication questions, symptom assessment, escalation triggers, patient communication).
Design and run test cases; document findings clearly and identify patterns in model behavior—where it excels and where it struggles.
Identify safety risks, clinical inaccuracies, and edge cases through hands-on interaction with the model. Flag concerns early and escalate appropriately.
Maintain detailed testing documentation that helps the team understand model performance and improvement priorities.
Gold Standards Development
Author and refine test cases, rubrics, and reference answers for clinical scenarios. Work with senior team members to establish what "accurate" and "safe" means in specific contexts.
Build out the gold standards library systematically—documenting clinical best practices, nursing workflows, and patient communication standards that the model should match.
Review existing test cases and suggest refinements based on clinical expertise and testing learnings.
Clinical Product Input
Translate frontline nursing expertise into product and technical feedback. When you see a dialog flow that doesn't match how nurses actually communicate, speak up.
Contribute clinical perspective to product discussions: safety rail design, escalation criteria, patient communication style, feature prioritization.
Review product enhancement opportunities and suggest improvements based on clinical best practices and nurse workflows.
Work with senior leads to understand the reasoning behind product decisions—learning how clinical requirements translate into technical implementation.
Frontline Research & Insights
Conduct user research at health systems (with guidance from senior team members or product leads). Observe real interactions, gather feedback from nurses and patients.
Document clinical observations and user feedback. Synthesize insights and present them to the team.
Maintain connections with pilot health systems; serve as a feedback loop for model validation and product improvement.
Cross-Functional Collaboration
Partner with ML researchers, product managers, and QA teams to understand model capabilities and testing priorities.
Communicate clinical findings clearly to technical audiences—translating clinical language into technical language and vice versa.
Ask smart questions and learn how product decisions get made. Bring a clinical lens to technical discussions.
Contribute to clinical documentation, safety validation reports, and regulatory submissions (with guidance).
What You Bring
Must-Have
Bachelor's degree from an accredited university (BSN, BA, BS, or equivalent).
Active, unrestricted Nursing License (RN) in your practicing state.
3+ years of clinical nursing experience in acute, ambulatory, telehealth, or community settings. You understand patient care, clinical workflows, and what nurses need.
Strong clinical judgment and patient communication skills. You know how to talk to patients, when to escalate, and how to build trust in high-stakes scenarios.
Experience with language models or conversational AI. You've used LLMs (ChatGPT, Claude, etc.), interacted with clinical tools, or worked with voice/conversational systems. You're comfortable learning how they work.
Excellent written and verbal communication. You can explain medical concepts clearly to both clinical and technical audiences. You can document findings precisely.
Commitment to patient safety and clinical ethics. You take privacy, accuracy, and regulatory compliance seriously. You're familiar with HIPAA/PHI and understand the stakes.
Growth mindset. You're curious about product, validation, and AI. You're excited to learn how these intersect. You're not defensive about feedback—you want to get better.
Intellectual rigor. You care about precision. You can think systematically about problems. You notice patterns and edge cases. You ask "why?" and dig deeper.
Nice-to-Have
Background in nursing informatics, quality/safety, or clinical education.
Familiarity with clinical guidelines, evidence-based practice, or patient education.
Experience with clinical validation, quality assurance, or testing frameworks.
Some exposure to healthcare tech, digital health, or clinical AI environments.
Interest in healthcare technology and how AI can improve patient care.
Comfort with ambiguity and learning in fast-moving startup environments.
The Details
Travel: 25-40% travel to health systems, user research, and strategic meetings. This includes multi-day site visits for observing real patient interactions and gathering clinical feedback.
Location: Remote-first. You can be based anywhere in the US. Regular all-hands in Menlo Park (2-3x per year).
Please be aware of recruitment scams impersonating Hippocratic AI. All recruiting communication will come from @hippocraticai.com email addresses. We will never request payment or sensitive personal information during the hiring process.