MailerMen logo

LLM Applications Intern - RAG and Evaluation

MailerMen

Actively hiring🌍 Remote (Worldwide)RemoteInternship₹17.5k – ₹22.5k /mo0 yrs expData Scientist
Posted 2d agoBe an early applicant

Start here

MailerMen answers roughly 300 support questions a week, and most of them already have an answer buried somewhere in our help documentation. Your three months go into building the retrieval-augmented assistant that finds it.

Scope

You chunk and embed the docs, stand up vector search over them, then spend most of your time on the unglamorous half: an evaluation set built from real questions, a rubric for what counts as a grounded answer, and a regression run that catches the day a prompt tweak quietly makes things worse. You will work in Python with Hugging Face embeddings and whichever hosted model wins on your own benchmark, not on anyone's intuition.

Practicalities

This is the one role on our AI team open worldwide, so we run it asynchronously: written updates in a shared doc, one live call a week scheduled around your timezone, and review through pull requests. You should be a final-year student or recent graduate who has already built something with an LLM API and got frustrated at how hard it is to tell whether a change improved anything. No professional experience needed, only evidence you can measure your own work.

Responsibilities

  • Chunk and embed the help documentation and stand up vector search over the corpus
  • Build a retrieval-augmented answering flow in Python and iterate on prompts using evidence rather than intuition
  • Assemble an evaluation set from real support questions and write the rubric for a grounded answer
  • Run a regression suite on every prompt or retrieval change so quality drops surface the same day
  • Compare hosted model options on your own benchmark and recommend one with costs attached
  • Post a written weekly update covering what changed, what improved and what regressed

Requirements

  • Final-year student or recent graduate who has already built something on top of an LLM API
  • Python fluency and comfort working entirely through pull requests and written updates
  • Understands embeddings, chunking, and why retrieval quality caps answer quality
  • Can define an evaluation criterion for open-ended text output and defend it in review
  • At least two working hours of overlap with India Standard Time

Skills

Benefits

  • Monthly stipend paid wherever in the world you are based
  • Fully asynchronous team with a single scheduled call each week in your timezone
  • The evaluation harness you build becomes the standard the team keeps using
  • Certificate and a recommendation letter from the engineering lead