ASSET Seminar: “Distributional robustness under the microscope: Empirical and theoretical case studies in old-school neural network learning”
September 9 at 12:00 PM - 1:00 PM
Organizer
asset-info@seas.upenn.edu
View Website
Venue
Distribution shift, i.e. a change in the distribution of input/output data between training and test time, remains an enduring challenge in modern ML. In practice, methods to improve robustness of ML models to distribution shift are frequently heuristic and require considerable side information about the shifted (target) test distribution to be effective. In theory, fundamental questions like why ML models are initially brittle to distribution shifts, how to best mitigate this brittleness, and how much target data is required for this mitigation remain open.
In this talk, we put distributional robustness under the microscope by studying the “group robustness” setting – where data is comprised of majority and minority groups whose relative proportions shift between training and test time. We present two case studies:
- On the empirical side, we show that simply retraining the last layer of state-of-the-art neural networks on a class-balanced held-out set is surprisingly effective – even if the held-out set remains sizably group-imbalanced – on well-established group robustness benchmarks. We also present a data-driven, active learning-inspired methodology to select and upweight minority group examples which achieves comparable group robustness on these same benchmarks to state-of-the-art methods that do require group labels.
- On the theoretical side, we provide the first end-to-end analysis showing that shallow neural networks learn a simpler spurious correlation and suppress the true (core) predictor. We consider the core predictor to be the nonlinear XOR function while the spurious correlation is linear. Our analysis and techniques reveal a striking phase transition between the situations of “extreme” correlation and “moderate” correlation, and could explain the success of existing robustness-advancing methods as well as inspire new ones.
This is joint work with Tyler LaBonte, John C. Hill, Xinchen Zhang and Abhishek Kumar.
https://upenn.zoom.us/j/97645937545
Meeting ID: 976 4593 7545
Speaker

Vidya Muthukumar
Assistant Professor, Georgia Institute of Technology
Vidya Muthukumar is the Harold R. and Mary Anne Nash Early Career Professor and Assistant Professor in the School of Electrical and Computer Engineering and H. Milton Stewart School of Industrial and Systems Engineering at Georgia Institute of Technology. Her broad interests are in game theory, online and statistical learning. She is particularly interested in designing learning algorithms that provably adapt in strategic environments, fundamental properties of overparameterized ML models, and the foundations of multi-agent decision-making. Before joining Georgia Tech, she spent a semester at the Simons Institute for the Theory of Computing as a research fellow for the program “Theory of Reinforcement Learning.” She is the recipient of an NSF CAREER Award, Amazon Research Award, Adobe Data Science Research Award and Simons-Berkeley Research Fellowship.
Read More
- CBE Doctoral Dissertation Defense: “Machine Learning-Augmented Mechanistic Modeling of Cancer Heterogeneity: Linking Mechanobiology, Microenvironment, and Therapeutic Response” (Sharvari Sudhir Kemkar)
- CBE Seminar Series: “Biology by Design: From Directed Evolution to Autonomous Experimentation” (Huimin Zhao, UIUC)

