Benchmarking, Correcting, and Accelerating AI Reasoning across Diverse Modalities Open Access
Zhang, Yifei (Spring 2026)
Abstract
AI has achieved remarkable progress in recent years. However, the reasoning processes of modern AI systems often remain opaque, error-prone, and computationally demanding. Although current models achieve strong predictive or generative performance, they often lack reliable mechanisms for evaluating reasoning quality, correcting flawed intermediate processes, and scaling reasoning efficiently across tasks. These limitations pose significant challenges for deploying AI in settings where reasoning must remain auditable, robust, and adaptable to heterogeneous sources of evidence.
This dissertation develops a unified research program for benchmarking, correcting, and accelerating AI reasoning, and further extends these advances to settings that require integrating heterogeneous evidence and multiple reasoning signals.
To benchmark reasoning quality, this dissertation introduces Saliency-Bench, a large-scale benchmark for evaluating visual explanations across eight datasets spanning natural and medical imaging domains. By establishing standardized datasets, evaluation protocols, and complementary metrics for faithfulness and alignment, together with inter-metric reliability analysis, Saliency-Bench provides a rigorous foundation for assessing explanation quality in image classification systems.
To correct flawed reasoning, the dissertation develops explanation-guided learning frameworks that leverage heterogeneous supervision signals. MAGI models annotator variability through multi-annotated supervision, while ESSA improves reasoning fidelity through saliency-guided data augmentation and iterative supervision. Together, these approaches show that explanations can serve not only as interpretive outputs but also as effective supervisory signals for improving model reasoning.
To accelerate reasoning, the dissertation proposes efficient mechanisms for transferring explanation-aware behavior across models. ELAD introduces explanation-guided active distillation from large to small language models, while VAPL enables test-time visual prompting to guide reasoning without retraining. Together, these methods improve the efficiency of reasoning transfer across both language and vision systems.
Finally, the dissertation extends this reasoning-centered framework to settings that require integrating heterogeneous evidence and interacting forms of reasoning. In healthcare, it presents an Explainable Cross-Disease Reasoning Framework for jointly reasoning over pulmonary and cardiovascular signals from low-dose chest CT. In human--AI interaction, it introduces EmoLLM, a framework that integrates cognitive context and emotional state to generate responses that are both factually grounded and contextually appropriate.
Table of Contents
Acknowledgments
List of Figures
List of Tables
Chapter 1. Introduction
Chapter 2. Benchmarking Visual Explanations
Chapter 3. Correcting AI Reasoning
Chapter 4. Accelerating AI Reasoning
Chapter 5. Reasoning Across Diverse Modalities
Chapter 6. Conclusion and Future Directions
Bibliography
About this Dissertation
| School | |
|---|---|
| Department | |
| Degree | |
| Submission | |
| Language |
|
| Research Field | |
| Keyword | |
| Committee Chair / Thesis Advisor | |
| Committee Members |
Primary PDF
| Thumbnail | Title | Date Uploaded | Actions |
|---|---|---|---|
|
|
Benchmarking, Correcting, and Accelerating AI Reasoning across Diverse Modalities () | 2026-03-28 13:16:30 -0400 |
|
Supplemental Files
| Thumbnail | Title | Date Uploaded | Actions |
|---|