Soumajyoti Sarkar

Background

I work as an applied scientist at Amazon AGI, on the core LLM pre-training and optimization science team. I work across the training stack at the intersection of neural scaling-law science and algorithm-system co-design of language models. My current work span the training dynamics of language models, model architectures designed for hardware efficiency, adaptive compute that mitigates inference inefficiencies, and the science of scaling laws. These days I also occasionally enjoy working on automated discovery of improved training algorithms, mainly toward model-architecture and training-recipe co-design.

Prior to this, I was part of the AI Research and Education (AIRE) group at AWS working on foundational models for structured knowledge grounding and training text embedding models that scale in distributed training environments. Even before, I contributed both as an ML engineer and researcher in the areas of search and recommendations at Twitter in San Francisco. I was part of the Tweet Search Ranking team where I worked on prototyping and deploying Twitter's first content based search relevance model utilizing explicit user survey feedback in Twitter's Search service.

I obtained my PhD in computer science at the Arizona State University, Tempe. My thesis focused on measuring the impact of social network interactions using observational and experimental studies. During my graduate studies, I had the opportunity to spend summers at A9 in Palo Alto and Nokia Bell Labs in New Jersey. I live in and work remotely from San Francisco, California. I finished my undergraduate studies at the Indian Institute of Engineering Science and Technology (IIEST), Shibpur.

Research

My research interests include topics in large-scale machine learning, including data and model efficient pretraining, scaling in the data-limited infinite compute regime and distributed ML optimization. I also enjoy reading up papers related to on-policy distillation. In ancient times, I used to work in computational social science, and their applications in search and recommendation systems. I also enjoy doing independent research in the fields at the intersection of economics and machine learning, mainly with decision making in peer lending platforms and reinforcement learning.

You can find an updated list of published papers and preprints in my Google Scholar page. Please feel free to reach out to me using my email on anything related to my research, paper reviews or any collaborations. I have occasionally kept a few running notes on my ramblings with machine learning over the years which can be found in the Unpublished section in the navigation bar.

News

Check out our recent preprints on scaling law science and systems co-design - Compute-Optimal Is Not Cluster-Optimal: Systems-Aware Scaling for Sparse Mixture-of-Experts and Tokens-per-Parameter Coverage Is Critical for Robust LLM Scaling Law Extrapolation.

Other

I spend quite a large part of my time outside of work on books, movies and swimming. I also occasionally enjoy playing outdoor soccer. I watch an unhealthy amount of tennis matches on youtube (I don't play the sport), and I used to follow Wawrinka for a long time, being my favorite player.