
About this role
Two Sigma is a leading quantitative investment management and trading firm. The company applies a scientific approach to investing, combining cutting-edge technology, artificial intelligence, data science, and quantitative research with rigorous human inquiry to capitalize on market opportunities and deliver alpha for investors.
Our team of engineers, quantitative researchers and data scientists looks beyond the traditional to test hypotheses and develop creative solutions to some of the world’s most complex economic problems.
We are applying large language models and transformer-based architectures to problems where ground truth is delayed, noisy, and non-stationary. Our systems generate code, run experiments, and iterate autonomously, and we are looking to go beyond supervised fine-tuning.
We are hiring a Post-Training Research Scientist to build RLHF, DPO, and reward modeling capabilities from the ground up. This is a greenfield role: you will define the infrastructure, research agenda, and evaluation frameworks for aligning LLMs to sophisticated, multi-step workflows in a domain where the reward signal is fundamentally different from existing research on human preference or deterministic task completion.
This hire will help own methodology across training, fine-tuning, context management, and model evaluation. You will shape not only the post-training capability but the broader research direction of the team.