ASSET Seminar: “Improving Generative AI at Inference Time: Alignment, Reasoning, & Efficiency”
March 4 at 12:00 PM - 1:15 PM
Organizer
asset-info@seas.upenn.edu
View Website
Venue
Modern generative AI is often improved through costly post-training (e.g., RLHF or preference tuning). This talk highlights a complementary alternative called inference-time methods. With the right inference-time objectives, search procedures, and compute allocation, we can meaningfully improve generative AI model behavior without updating model weights. We will begin with a principled formulation of inference-time methods for AI alignment and introduce transfer decoding, which estimates token-level optimal values for reward-guided decoding by leveraging an already available baseline-aligned model. Building on this decoding-time optimization viewpoint, we then move to multi-criteria alignment at inference time under preferences and safety constraints. In particular, we will discuss a satisficing view of alignment and an inference-time framework that operationalizes this idea. Next, we will show how inference-time techniques can also improve reasoning performance: why simply thinking longer can create a mirage of improvement and hurt accuracy via overthinking, and how parallel thinking provides a more reliable inference-time scaling strategy under the same compute budget. Finally, we extend these ideas beyond autoregressive LLMs to diffusion LLMs, where flexible generation orders implicitly expose hidden semi-autoregressive experts that enable selective computation at inference time. By ensembling across diverse block schedules, we can allocate compute more effectively and reliably boost performance, again without additional training.

