Benchmarking, Correcting, and Accelerating AI Reasoning across Diverse Modalities Open Access

Zhang, Yifei (Spring 2026)

Permanent URL: https://etd.library.emory.edu/concern/etds/3197xn83d?locale=en
Published

Abstract

AI has achieved remarkable progress in recent years. However, the reasoning processes of modern AI systems often remain opaque, error-prone, and computationally demanding. Although current models achieve strong predictive or generative performance, they often lack reliable mechanisms for evaluating reasoning quality, correcting flawed intermediate processes, and scaling reasoning efficiently across tasks. These limitations pose significant challenges for deploying AI in settings where reasoning must remain auditable, robust, and adaptable to heterogeneous sources of evidence.

This dissertation develops a unified research program for benchmarking, correcting, and accelerating AI reasoning, and further extends these advances to settings that require integrating heterogeneous evidence and multiple reasoning signals.

To benchmark reasoning quality, this dissertation introduces Saliency-Bench, a large-scale benchmark for evaluating visual explanations across eight datasets spanning natural and medical imaging domains. By establishing standardized datasets, evaluation protocols, and complementary metrics for faithfulness and alignment, together with inter-metric reliability analysis, Saliency-Bench provides a rigorous foundation for assessing explanation quality in image classification systems.

To correct flawed reasoning, the dissertation develops explanation-guided learning frameworks that leverage heterogeneous supervision signals. MAGI models annotator variability through multi-annotated supervision, while ESSA improves reasoning fidelity through saliency-guided data augmentation and iterative supervision. Together, these approaches show that explanations can serve not only as interpretive outputs but also as effective supervisory signals for improving model reasoning.

To accelerate reasoning, the dissertation proposes efficient mechanisms for transferring explanation-aware behavior across models. ELAD introduces explanation-guided active distillation from large to small language models, while VAPL enables test-time visual prompting to guide reasoning without retraining. Together, these methods improve the efficiency of reasoning transfer across both language and vision systems.

Finally, the dissertation extends this reasoning-centered framework to settings that require integrating heterogeneous evidence and interacting forms of reasoning. In healthcare, it presents an Explainable Cross-Disease Reasoning Framework for jointly reasoning over pulmonary and cardiovascular signals from low-dose chest CT. In human--AI interaction, it introduces EmoLLM, a framework that integrates cognitive context and emotional state to generate responses that are both factually grounded and contextually appropriate.

Table of Contents

Acknowledgments

List of Figures

List of Tables

Chapter 1. Introduction

Chapter 2. Benchmarking Visual Explanations

Chapter 3. Correcting AI Reasoning

Chapter 4. Accelerating AI Reasoning

Chapter 5. Reasoning Across Diverse Modalities

Chapter 6. Conclusion and Future Directions

Bibliography

About this Dissertation

Rights statement
  • Permission granted by the author to include this thesis or dissertation in this repository. All rights reserved by the author. Please contact the author for information regarding the reproduction and use of this thesis or dissertation.
School
Department
Degree
Submission
Language
  • English
Research Field
Keyword
Committee Chair / Thesis Advisor
Committee Members
Last modified

Primary PDF

Supplemental Files