Loading this job…
Loading this job…
MailerMen is hiring a QA Engineer – AI Output Evaluation to review and grade the technical answers produced by AI systems, so that the models behind them can be trained on accurate, well-judged human feedback. This is not an ordinary software-testing position. The day-to-day work is evaluation rather than product release testing, and the project is open only to QA specialists who have already been paid to work on human data for AI training, such as annotation, labelling, RLHF, AI response evaluation, model evaluation or rubric-based grading. RLHF, or reinforcement learning from human feedback, is the training method in which people rate or rank a model's answers and those ratings are then used to teach the model what a good answer looks like. Software QA experience on its own, without hands-on human evaluation of AI output, does not meet this requirement.
The core task is to evaluate and rate AI-generated technical outputs against defined quality criteria and rubrics, with particular attention to answers that look correct on the surface but are in fact inaccurate or incomplete. Alongside that, you will design and assess comprehensive test cases for functional, regression and edge-case scenarios, including negative and boundary tests, and review bug reports and test documentation to confirm that issues are reproducible, that the documentation is complete and that severity has been judged correctly. You will identify, isolate and document defects with precise reproduction steps and record them through structured tracking mechanisms.
Much of the value of the work lies in the written record you leave behind. Feedback and annotations have to be detailed and specific enough that a developer can act on them without coming back for clarification, which is why the project asks for written English at B2 level or above, meaning you can set out a technical argument precisely and be understood first time. You will also work with project teams to refine the evaluation guidelines themselves and to improve testing standards and methods as the project runs. This is a high-volume project, so applying a rubric consistently matters as much as individual judgement. No background in building or training AI systems is expected; your QA domain knowledge, together with the paid human-data experience described above, is what the project needs.
This is a contract engagement, worked remotely from anywhere in India. Compensation is ₹1,500 per hour. The role calls for 7+ years of experience. A LinkedIn profile is required and must be provided with your application. No formal degree is required, because practical, demonstrable testing experience takes precedence, and you will need a reliable internet connection and the readiness to begin promptly. Apply through MailerMen with your updated resume and a summary of your relevant experience.
MailerMen runs a verified job board covering startup and product roles across twelve markets, and takes on interns across engineering, data, design and marketing to build it.
Vistula Software House🇵🇱 Remote (Poland)