AI Safety Practitioner
Mercor connects highly skilled creative and technical professionals with leading artificial intelligence research labs working on advanced AI systems.
Headquartered in San Francisco, the company provides opportunities for experienced professionals to contribute their expertise to cutting-edge AI research and development projects. Mercor is backed by prominent investors including Benchmark, General Catalyst, Peter Thiel, Adam D’Angelo, Larry Summers, and Jack Dorsey.
About the Role
Mercor is looking for an experienced AI Safety Practitioner to support projects focused on evaluating and improving the safety, accuracy, and reliability of advanced AI models.
This is a remote contract opportunity where you will analyze AI-generated responses, identify potential safety and reasoning problems, and provide structured feedback that helps researchers develop safer and more reliable AI systems.
The work involves evaluating sophisticated and sometimes sensitive scenarios, requiring strong analytical judgment, attention to detail, and the ability to consistently apply established safety standards.
Key Responsibilities
- Assess AI-generated responses for safety, factual accuracy, policy compliance, and overall quality.
- Review model outputs involving sensitive areas such as misinformation, political persuasion, self-harm, violence, cybersecurity, biosecurity, and other high-risk topics.
- Apply established evaluation criteria across RLHF, SFT, AI safety benchmarking, and related model-training workflows.
- Help refine evaluation rubrics to improve the consistency and effectiveness of AI assessments.
- Detect hallucinations, unsafe responses, reasoning failures, and potential policy violations.
- Evaluate complex or ambiguous AI responses that require careful contextual judgment.
- Produce clear, structured feedback that can be used to improve model alignment and safety.
- Work alongside AI researchers and safety teams on ongoing model-evaluation projects.
- Maintain consistent quality standards when reviewing large volumes of AI-generated content.
Required Qualifications
Candidates should have:
- A Bachelor’s degree or higher in Journalism, Communications, Psychology, Sociology, Public Policy, Law, Biology, Chemistry, Computer Science, or another relevant discipline.
- At least 5 years of professional experience in areas such as AI Safety, Trust & Safety, journalism, public policy, scientific research, security, or a related field.
- Excellent written English and the ability to communicate evaluations clearly.
- Strong critical-thinking and analytical-reasoning abilities.
- Excellent attention to detail.
- Ability to consistently assess nuanced, complicated, and policy-sensitive situations.
- Sound professional judgment when reviewing potentially high-risk material.
Preferred Qualifications
Candidates may have an advantage if they have:
- Previous experience with AI Safety or AI model evaluation.
- Experience working with RLHF or SFT processes.
- A background in Trust & Safety.
- Familiarity with AI safety policies and content-moderation frameworks.
- Experience creating, refining, or applying evaluation rubrics.
- Previous responsibility for reviewing complicated, ambiguous, or high-risk content.
- Familiarity with common AI issues such as hallucinations, reasoning failures, and unsafe model outputs.
Application Process
The initial application process is expected to take approximately 20–30 minutes.
Applicants will typically be required to:
- Upload their résumé.
- Complete an AI-based interview tailored to their résumé and professional background.
- Complete and submit the required application form.
Applications are reviewed regularly. Candidates should complete all required application and interview stages to be fully considered for the opportunity.
Requirements
Candidates should meet the following requirements:
- A Bachelor’s degree or higher in Journalism, Communications, Psychology, Sociology, Public Policy, Law, Biology, Chemistry, Computer Science, or another relevant discipline.
- At least 5 years of professional experience in AI Safety, Trust & Safety, journalism, public policy, scientific research, security, or a closely related field.
- Excellent written English with the ability to communicate findings clearly and precisely.
- Strong critical-thinking, analytical, and reasoning skills.
- Ability to consistently evaluate complex, nuanced, and policy-sensitive scenarios.
- Strong attention to detail and sound judgment when reviewing potentially sensitive AI-generated content.
Preferred Qualifications
Additional experience in the following areas would be advantageous:
- Previous work involving AI Safety, RLHF, SFT, Trust & Safety, or AI model evaluation.
- Familiarity with AI safety policies and content-moderation frameworks.
- Experience developing, refining, or applying evaluation rubrics.
- Experience assessing complicated, ambiguous, or high-risk content.
- Understanding of common AI model issues, including unsafe outputs, hallucinations, and reasoning failures.