LLM Red Team Specialist — Failure Modes & Edge Cases (Train AI Models Part Time!)
IndiaPosted Aug 5, 2026
LLM Red Team Specialist — Failure Modes & Edge Cases (Train AI Models Part Time!)
hackajob United StatesLLM Red Team Specialist — Failure Modes & Edge Cases (Train AI Models Part Time!)
hackajob
United States
1 week ago
26 applicants
See who hackajob has hired for this role
hackajob is collaborating with Mercor to connect them with exceptional professionals for this role.- Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced AI models.** ## 1\. Overview A leading AI lab is building the next generation of agentic evaluation benchmarks for frontier models and needs specialists who can find where those models break. Working in a red-teaming setup, you will design and probe complex, multi-step tasks to expose vulnerabilities, edge cases, and failure modes in frontier AI systems — the places where a model looks competent but is quietly wrong. Each task represents one to two days of continuous, focused effort and spans multiple technical skills: coding, experimentation, and careful analysis. You will work in a tight feedback loop with the lab's researchers, turning the failure modes you find into stronger benchmark tasks. This is a full-time W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI lab as part of their extended workforce. This role is fully remote within the United States, at approximately 35 hours per week. ## 2\. Key Responsibilities - **Probe models:** Explore how frontier AI models behave on coding, ML, and analysis tasks, and find the spots where they quietly get things wrong. - **Design challenges:** Turn the weaknesses you find into well-crafted tasks that are hard for models but fair to grade. - **Document findings:** Write up what you discover clearly, with evidence and steps others can reproduce. - **Strengthen tasks:** Team up with task authors to close loopholes, shortcuts, and grading gaps. - **Work as a team:** Share insights with researchers and fellow experts so the benchmark keeps getting better. ## 3\. Core Qualifications - MSc or PhD in a STEM field, or equivalent practical experience in a research-heavy domain requiring data analysis and coding. - 1+ years of experience in a research, research-engineering, security, or AI-evaluation role. - Demonstrated ability to identify vulnerabilities, edge cases, or failure modes in LLMs or ML systems — through red teaming, adversarial testing, security research, or rigorous model evaluation. - Working proficiency in Python and Git, with the ability to script your own probes and analyses. - Strong familiarity with LLM capabilities, limitations, and evaluation techniques. - Past experience in AI training, model evaluation, or benchmark/task authoring is preferred. - A perfectionist mindset: high attention to detail, creativity in finding what others missed, strong written communication, and the ability to work independently through ambiguous, open-ended problems. - Ability to engage reliably for approximately 35 hours per week. ## About Cincinnatus LLC Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. Roles hired through Cincinnatus are not project-based or freelance engagements. They are structured, role-based positions that typically involve part-time or full-time commitments, close collaboration with a client's internal teams, and integration into standard enterprise workflows. Cincinnatus is a legal entity separate from Mercor. While opportunities may be discovered through Mercor's platform, employment, onboarding, payroll, and benefits for these roles are administered by Cincinnatus LLC. ## Equal Employment Opportunity Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic. Cincinnatus is committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans throughout the job application process.
-
Seniority level
Entry level -
Employment type
Full-time -
Job function
Human Resources -
Industries
Software Development
Referrals increase your chances of interviewing at hackajob by 2x
See who you knowGet notified about new Team Specialist jobs in United States.
Sign in to create job alertSimilar jobs
-
AI Red Teamer, LLM Generalist
AI Red Teamer, LLM Generalist
Handshake
Seattle, WA $32.00 - $95.00 6 days ago -
AI Red Team Engineer
AI Red Team Engineer
White Circle
United States $60,000.00 - $90,000.00 1 month ago -
AI Red Teamer
AI Red Teamer
10a Labs
Washington, DC 3 months ago -
GenAI Safety Analyst
GenAI Safety Analyst
Alice (Formerly ActiveFence)
New York, NY 1 week ago -
AI Content Red Team Analyst - Trust and Safety
AI Content Red Team Analyst - Trust and Safety
TikTok
San Jose, CA 1 week ago -
Senior GenAI Safety Researcher
Senior GenAI Safety Researcher
Alice (Formerly ActiveFence)
New York, NY 1 week ago -
AI Trainer Jobs in the United States
AI Trainer Jobs in the United States
Rex.zone
United States $30.00 - $50.00 1 week ago -
AI Research Scientist (United States, Remote)
AI Research Scientist (United States, Remote)
Rex.zone
United States $30.00 - $50.00 1 week ago -
AI/LLM Safety Engineer
AI/LLM Safety Engineer
Propio Language Services
Overland Park, KS 1 month ago -
Oversight Operations Investigator
Oversight Operations Investigator
Meta
United States $100,000.00 - $143,000.00 2 days ago -
Red Team Engineer, Safeguards
Red Team Engineer, Safeguards
Anthropic
San Francisco Bay Area 4 days ago -
AI Evaluation Lead
AI Evaluation Lead
Elly
United States $120,000.00 - $200,000.00 5 days ago -
AI Agent Abuse Prevention Engineer
AI Agent Abuse Prevention Engineer
Zendesk
Washington, United States 6 days ago -
AI Agent Abuse Prevention Engineer
AI Agent Abuse Prevention Engineer
Zendesk
California, United States 6 days ago -
AI Implementation Quality Analyst
AI Implementation Quality Analyst
Granicus
United States 1 month ago -
AI Red Teamer, Cyber
AI Red Teamer, Cyber
10a Labs
Los Angeles, CA 2 months ago -
Safeguards Enforcement Analyst, Safety Evaluations
Safeguards Enforcement Analyst, Safety Evaluations
Anthropic
San Francisco Bay Area 2 weeks ago -
Red Team Engineer
Red Team Engineer
Gray Swan
United States $110,000.00 - $290,000.00 5 days ago -
Member of Technical Staff - Safety
Member of Technical Staff - Safety
Reflection
San Francisco, CA 1 week ago -
AI Penetration Tester
AI Penetration Tester
BMO U.S.
Texas, United States 2 weeks ago -
AI Penetration Tester
AI Penetration Tester
BMO U.S.
Arizona, United States 2 weeks ago -
LLM Evaluation Engineer
LLM Evaluation Engineer
ThirdLaw | Runtime AI Safety
United States 9 months ago -
Research Engineer, AI Safety & Alignment
Research Engineer, AI Safety & Alignment
Character.AI
Redwood City, CA 1 day ago -
Model Policy
Model Policy
OpenAI
San Francisco, CA $207,000.00 - $295,000.00 1 week ago -
Staff Research Engineer, Post-training & Evaluation
Staff Research Engineer, Post-training & Evaluation
Reddit, Inc.
Seattle, WA 2 weeks ago -
Member of Technical Staff - Safety
Member of Technical Staff - Safety
Reflection
New York, NY 1 week ago -
Senior AI Security Researcher
Senior AI Security Researcher
NVIDIA
Washington, United States 6 days ago
People also viewed
-
Senior AI Security Researcher
Senior AI Security Researcher
Texas, United States 6 days ago -
Agent Post-Training, Personality
Agent Post-Training, Personality
San Francisco, CA $295,000.00 - $445,000.00 2 weeks ago -
GenAI Safety Team Lead
GenAI Safety Team Lead
New York, NY 1 week ago -
GenAI Analyst
GenAI Analyst
United States $80.00 - $87.00 4 days ago -
AI Agent Abuse Prevention Engineer
AI Agent Abuse Prevention Engineer
Texas, United States 6 days ago -
AI Red Teamer, Cyber
AI Red Teamer, Cyber
Washington, DC 2 months ago -
AI Red Teamer, Cyber
AI Red Teamer, Cyber
New York, NY 2 months ago -
AI Agent Abuse Prevention Engineer
AI Agent Abuse Prevention Engineer
Arizona, United States 6 days ago -
STEM Careers in the United States
STEM Careers in the United States
United States $30.00 - $50.00 1 week ago -
Safeguards Enforcement Analyst, User Well-being
Safeguards Enforcement Analyst, User Well-being
San Francisco Bay Area 22 hours ago
Similar Searches
-
Dean of Student Affairs jobs
1,002 open jobs -
Associate Human Resources Generalist jobs
100 open jobs -
Associate Project Director jobs
71,818 open jobs -
Regional jobs
322,544 open jobs -
Adoption Specialist jobs
34,902 open jobs -
Choir Director jobs
208 open jobs -
Small Business Advisor jobs
73,239 open jobs -
Special Education Teacher jobs
37,804 open jobs -
Online Marketing Executive jobs
939 open jobs -
Document Reviewer jobs
8,250 open jobs -
Prospect Researcher jobs
1,735 open jobs -
Credentialing Coordinator jobs
33,355 open jobs -
Landscape Manager jobs
6,608 open jobs -
Catalog Manager jobs
17,202 open jobs -
Architectural Manager jobs
62,327 open jobs -
Chief Communications Officer jobs
1,274 open jobs -
Resource Coordinator jobs
49,077 open jobs -
Mental Health Specialist jobs
26,738 open jobs -
Health Information Specialist jobs
39,874 open jobs -
Senior Coordinator jobs
29,514 open jobs -
Banking Advisor jobs
71,154 open jobs -
Team Assistant jobs
181,196 open jobs -
Chaplain jobs
3,177 open jobs -
Benefits Specialist jobs
51,315 open jobs -
Admissions Specialist jobs
35,378 open jobs