Skip to content

Mononito Goswami

Senior Applied Scientist, AWS

I study how AI systems learn to reason, and how that process can be measured and shaped. Recently, I have been thinking about this through the lens of metacognition: how models learn how to reason.

My recent research at Amazon asks three questions: (1) How can teacher models learn to teach better by adapting to their students (SCOUT)? (2) How can small models learn to reason across contexts and benefit from more test-time compute (Hermes)? (3) How can agents reason about the progress of evolutionary search and intervene when it stalls (MILO)? More broadly, I am interested in post-training, continual learning, and self-improving AI systems.

I also lead science for the AWS DevOps Agent, where these ideas meet real-world constraints. I helped take the agent from inception to general availability, developing agentic swarms, continual learning, and scalable evaluation capabilities. The resulting science now powers a production system used by thousands of customers.

Previously, I studied how advances in language modeling could be extended beyond language to structured data. This work helped usher in foundation modeling for time series through MOMENT and Chronos-2, alongside work on how we evaluate and select models.

I completed my Ph.D. in Robotics at Carnegie Mellon University, advised by Artur Dubrawski, and received the Robotics Institute Distinguished Dissertation Award. I have also worked at Google Research and AWS AI Labs.

I care deeply about mentoring and working with students on ambitious research problems. To talk research or explore working together, reach out.

Mononito Goswami

News

Selected Publications

Post-Training & Test-Time Scaling

SCOUT: On the Off-Policy Teacher in On-Policy Distillation
Langlin Huang, Hao Liu, Mononito Goswami, Xinyu Li, Prithwish Jana, Nikos Kanakaris, Patrick Blöbaum, Purak Jain
ArXiv Preprint, 2026
Identifies teacher-side off-policy degradation in on-policy distillation and mitigates it by adapting the teacher to student-generated prefixes with RL.
Hermes: Learning Contextual Reasoning Unlocks Test-Time Scaling
Xinyu Li*, Mononito Goswami*, Hao Liu, Nikos Kanakaris, Langlin Huang, Prithwish Jana, Patrick Blöbaum, Purak Jain
ArXiv Preprint, 2026
Flexible context management for long-horizon test-time scaling: a parameterized delegation/digestion harness, and an SFT+RL recipe (Hermes-Learn) that teaches a 4B model adaptive context reuse.

Agents

SpIDER: Spatially Informed Dense Embedding Retrieval for Software Issue Localization
Shravan Chaudhari, Rahul Thomas Jacob, Jiajun Cao, Shihab Rashid, Mononito Goswami, Christian Bock
Conference on Empirical Methods in Natural Language Processing (EMNLP), 2026
Retrieval and reasoning framework for software issue localization, combining structural graph exploration with LLM-based reasoning to improve localization accuracy.
MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution
Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc
ArXiv Preprint, 2026
Automated harness discovery through orchestrated multi-agent evolution: the search co-evolves the harness and its own strategy, reaching state-of-the-art results on Terminal-Bench 2.1, PaperBench, and DeepSWE.
TimeSeriesGym: A Scalable Benchmark for (Time Series) Machine Learning Engineering Agents
Yifu Cai, Xinyu Li, Mononito Goswami, Michał Wiliński, Gus Welter, Artur Dubrawski
ArXiv Preprint, 2025
A benchmarking environment for evaluating ML engineering agents on the full time-series modeling stack, from data exploration and modeling to testing and deployment.

Foundation Models

Chronos-2: From Univariate to Universal Forecasting
Abdul Fatir Ansari, Oleksandr Shchur, Jaris Küken, Andreas Auer, Boran Han, Pedro Mercado, Syama Sundar Rangapuram, Huibin Shen, Lorenzo Stella, Xiyuan Zhang, Mononito Goswami, Shubham Kapoor, Danielle C. Maddix, Pablo Guerron, Tony Hu, Junming Yin, Nick Erickson, Prateek Mutalik Desai, Hao Wang, Huzefa Rangwala, George Karypis, Yuyang Wang, Michael Bohlke-Schneider
ArXiv Preprint, 2025
Extends the Chronos forecasting model to handle multivariate, covariate-conditioned, and probabilistic forecasting in a single unified architecture.
Exploring Representations and Interventions in Time Series Foundation Models
Michał Wiliński, Mononito Goswami, Willa Potosnak, Nina Żukowska, Artur Dubrawski
International Conference on Machine Learning (ICML), 2025
First to show that time series foundation models learn interpretable concepts (trends, seasonality) despite self-supervised training, and that these representations can be targeted for intervention.
MOMENT: A Family of Open Time-series Foundation Models
Mononito Goswami, Konrad Szafer, Arjun Choudhry, Yifu Cai, Shuo Li, Artur Dubrawski
International Conference on Machine Learning (ICML), 2024
One of the first open-source time series foundation models. 2.5M+ downloads on HuggingFace, 700+ GitHub stars.

Evaluation Science

TimeSeriesExamAgent: Creating Time Series Reasoning Benchmarks at Scale
Małgorzata Gwiazda, Yifu Cai, Mononito Goswami, Arjun Choudhry, Artur Dubrawski
International Conference on Learning Representations (ICLR), 2026
Benchmark for temporal reasoning in LLMs, with scalable task generation via LLM agents and item response theory.
AQuA: A Benchmarking Tool for Label Quality Assessment
Mononito Goswami, Vedant Sanil, Arjun Choudhry, Arvind Srinivasan, Chalisa Udompanyawit, Artur Dubrawski
Neural Information Processing Systems (NeurIPS), 2023 Datasets and Benchmarks Track
A comprehensive benchmarking tool for evaluating label error detection methods across diverse datasets and annotation types.
Unsupervised Model Selection for Time-series Anomaly Detection Spotlight
Mononito Goswami, Cristian Challu, Laurent Callot, Lenon Minorics, Andrey Kan
International Conference on Learning Representations (ICLR), 2023
Shows that ensembles of unsupervised heuristics and weak supervision can accurately select anomaly detection models without any ground-truth labels.
View all publications →

Awards & Honors

View all awards & honors →