FOLDS seminar: Coherence Mechanisms for Provable Self-Improvement
March 19 at 12:00 PM - 1:00 PM
Venue
Zoom link: https://upenn.zoom.us/j/98220304722
Large language models are increasingly trained to improve themselves, yet the mechanisms driving this, such as self-reflection or RLAIF, rely almost entirely on empirical heuristics. Is it possible to mathematically guarantee self-improvement without human supervision?
In this talk, I will introduce a geometric framework that proves self-improvement is not only possible but monotonic, grounded in the principle of coherence. By formalizing self-improvement as a Bregman projection onto a space of logically consistent models, we can guarantee enhanced performance. Furthermore, I will present a surprising characterization theorem: any self-improvement mechanism that offers similar theoretical guarantees must, fundamentally, be a coherence projection in disguise.
(Joint work with Jon Schneider and Yifan Wu.)

