:::
AIIC

Understanding Transformers


  • 講者 : Jerry Yao-Chieh Hu 博士
  • 日期 : 2026/08/26 (Wed.) 10:30~12:00
  • 地點 : 資創中心122 演講廳
  • 邀請人 : 陳駿丞
Abstract
Transformer foundation models (e.g., LLMs and generative AI) show remarkable capabilities yet remain black boxes. This talk will present scientific understanding of fixed-weight softmax Transformers through a brain-inspired lens with two modes: memory and procedure execution. On the memory side, I develop a statistical‑physics principle that connects attention to energy‑based (Hopfield‑style) associative memory via free‑energy minimization. This model-based perspective yields new Transformers and concrete design knobs (entropy and energy) for sparsity, stability, and knowledge control. On the procedure-execution side, I present a constructive theory of prompt programmability: one frozen Transformer can emulate many algorithms by swapping prompts rather than weights. This provides a mechanistic account of Transformer models' “one model for many tasks” ability: prompts as programs and the model as a CPU-like program executor. The talk will conclude by applying these results to scientific applications in genomics (DNA knowledge handling) and astrophysics (universal time-series forecasting).
Bio
Jerry Yao-Chieh Hu is an incoming Assistant Professor of Computer Science at Vanderbilt University, within the newly founded College of Connected Computing (CCC). He’s also a final-year CS PhD at Northwestern, advised by Han Liu. He holds his B.S. degree in Physics from National Taiwan University, advised by Pisin Chen.

His research focuses on theoretical foundations and principled methodologies for large foundation models. These efforts bridge theory and practice through new model designs and learning methods. Early interdisciplinary applications span physics, genomics, cancer immunology, and drug discovery. He’s the recipient of the 2025 Rising Stars in Data Science, 3 outstanding research awards, 2 named fellowships and 1 best theoretical paper award.