Loading this job…
MailerMen matches candidates to startup roles, and the matching only works when the underlying data is clean. Job titles reach us as free text from a thousand different employers, which is where you come in.
This two month internship runs from our Pune office, three days a week on site. Your main project is a title normalisation pipeline in Python. You will profile roughly 40,000 raw job titles with pandas, cluster the obvious duplicates, build a mapping table from the mess to a canonical taxonomy, and measure how much the candidate match rate improves once it is applied.
Mornings are usually notebook work. Afternoons often involve walking over to a recruiter and asking whether Growth Hacker II and Growth Marketing Associate should really collapse into the same bucket. Cleaning data at this scale is half code and half asking people questions, and you will do plenty of both.
You like the puzzle of messy strings more than the polish of a finished chart. You are in your final year or recently graduated, you know pandas well enough to be dangerous, and you are happy to have your assumptions checked against reality every few days.
MailerMen runs a verified job board covering startup and product roles across twelve markets, and takes on interns across engineering, data, design and marketing to build it.