Content Specialist III | AI Evaluation & Prompting
At a glance
Mid-level Content Creation role at Linda Werner & Associates. United States.
Pay not stated Content Creation salary
growthroles summary, based on the employer's posting.
What you'll do
- Try current AI model releases on varied subjects and conversations, noting strong responses and specific failures
- Score model answers using the supplied guidance, then check whether that guidance captures quality effectively
- Shape model tone and conduct by drafting and revising system prompts
- Audit human and automated evaluations, investigate breakdowns, and record findings clearly
- Follow detailed criteria across substantial work volumes as priorities and model behavior shift
What you bring
- At least five years in a relevant discipline such as writing, journalism, linguistics or another listed field
- A bachelor’s credential or equivalent experience, with expert knowledge in at least one subject
- Careful judgment for applying scoring rules, checking claims and spotting quality problems
Who this fits
This role suits a seasoned writer or subject expert who likes examining subtle differences in AI responses and working independently while balancing speed, volume and quality. It is a remote U.S. contract lasting six months, with forty hours expected weekly on a Monday-to-Friday, eight-hour shift. English is required, and the application asks about current or future U.S. work sponsorship.
From the employer
We are seeking an experienced Content Specialist III to help evaluate, refine, and improve advanced AI models and products.
In this role, you will work directly with evolving AI systems, testing how they respond across a wide range of topics and conversation types, identifying where they succeed or fall short, and helping shape the prompts, quality standards, and evaluation frameworks that influence how they communicate and behave.
This is a highly hands on opportunity for someone who brings together strong writing and content expertise, curiosity about AI, thoughtful judgment, and a sharp eye for quality. Your work will directly contribute to improving real world AI product experiences and how these systems perform for users.
What You Will Do
• Test new AI model versions across a variety of topics, use cases, and conversation types
• Evaluate model responses against established rubrics, guidelines, and quality standards
• Identify and document specific examples of successful and unsuccessful model behavior
• Write and refine system prompts that help shape model personality, tone, and behavior
• Assess whether evaluation rubrics and quality standards effectively measure model performance
• Perform quality reviews of both human and agent based evaluations to ensure accuracy and compliance
• Conduct hands on experiments with AI models and products
• Investigate model failures and document findings clearly
• Apply detailed instructions and evaluation criteria consistently across a high volume of work
• Adapt quickly as product priorities, model behavior, and evaluation needs evolve
Top Skills
• AI Model Evaluation
• Prompt Writing and Refinement
• Writing and Editing
• Rubric Based Quality Assessment
• Fact Checking and Analytical Judgment
Qualifications
• 5+ years of experience in writing, editing, journalism, production, linguistics, STEM, coding, policy, or another relevant subject matter field
• 1+ year of hands on AI experience preferred, including prompting, annotation, evaluation, or red teaming
• Bachelor’s degree or equivalent experience
• Deep expertise in at least one subject area with the ability to evaluate content as a subject matter expert
• Strong judgment and attention to detail with the ability to consistently apply detailed rubrics, verify information, and identify quality issues
Preferred Experience
• Experience evaluating generative AI or large language model outputs
• Strong fact checking and research skills
• Ability to verify claims against original sources
• Experience conducting detailed failure investigations
• Clear written communication and documentation skills
• Ability to flag questions and blockers early
• Comfortable shifting priorities as product needs change
This team operates in a fast moving environment where priorities can change from day to day. Successful candidates will be comfortable balancing quality, speed, volume, and independent judgment while following detailed instructions.
You should enjoy getting into the details, identifying subtle differences in AI responses, determining why something works or fails, and translating those observations into clear and actionable feedback.
Location: United States (Remote)
Role type: Contract 6 Month Position
Expected hours: 40 per week
Benefits:
- Dental insurance
- Health insurance
- Health savings account
- Life insurance
- Paid time off
- Retirement plan
- Vision insurance
Schedule:
- 8 hour shift
- Monday to Friday
Application Question(s):
- Do you or will you in the future require any sponsorship to work in the US?
Language:
- English (Required)
Not this one either?
Claude or ChatGPT reads the other 90,612 for you.
Get better matches →