Trustworthy and Retrieval-Augmented Machine Learning for Digital Health Restricted; Files Only
Xu, Ran (Spring 2026)
Abstract
Machine learning is increasingly deployed in digital health, yet its effectiveness is often limited by challenges in interpretability, fairness, and reliance on incomplete clinical data. This dissertation presents a set of complementary studies that advance trustworthy and retrieval-augmented machine learning for clinical prediction and medical reasoning. First, we investigate hypergraph-based models for electronic health records that capture higher-order clinical relationships and support factual and counterfactual reasoning, improving interpretability and promoting more balanced predictions across patient subgroups. Second, we explore retrieval-augmented clinical prediction methods that incorporate external biomedical knowledge, including a domain-specialized biomedical retriever for large language models and retrieval-enhanced EHR models that improve robustness and generalization in low-resource settings. Finally, we study self-improving retrieval-augmented generation techniques that adapt large language models to specialized medical domains through iterative refinement of retrieval and generation using synthetic supervision. Together, these works contribute practical and principled approaches toward building more reliable, transparent, and knowledge-aware machine learning systems for digital health.
Table of Contents
1 Introduction 1
2 Counterfactual and Factual Reasoning over Hypergraphs for Inter-
pretable Clinical Predictions 5
2.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6
2.2 Method . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 8
2.2.1 Notations and Definitions . . . . . . . . . . . . . . . . . . . . 9
2.2.2 Hypergraph Construction and Learning . . . . . . . . . . . . . 9
2.2.3 Important Subset Extraction via Counterfactual and Factual
Reasoning . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12
2.2.4 Alternate Training of fθ and gϕ . . . . . . . . . . . . . . . . . 14
2.3 Experiments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 15
2.3.1 Experiment Setup . . . . . . . . . . . . . . . . . . . . . . . . . 15
2.3.2 Implementation Details . . . . . . . . . . . . . . . . . . . . . . 16
2.3.3 Baselines . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 16
2.3.4 Experimental Results . . . . . . . . . . . . . . . . . . . . . . . 17
2.3.5 Parameter and Ablation Studies . . . . . . . . . . . . . . . . . 18
2.3.6 Qualitative Analysis . . . . . . . . . . . . . . . . . . . . . . . 19
2.3.7 Quantitative Study . . . . . . . . . . . . . . . . . . . . . . . . 20
2.4 Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 21
3 Hypergraph Transformer Pretrain-then-Finetuning for Balanced Clin-
ical Predictions 22
3.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 23
3.2 Preliminary Studies . . . . . . . . . . . . . . . . . . . . . . . . . . . . 26
3.2.1 Problem Setup . . . . . . . . . . . . . . . . . . . . . . . . . . 26
3.2.2 Limitations of Traditional ML Methods . . . . . . . . . . . . . 27
3.3 Method . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 29
3.3.1 Hypergraph Learning . . . . . . . . . . . . . . . . . . . . . . . 29
3.3.2 Pretrain-then-Finetune Pipeline . . . . . . . . . . . . . . . . . 32
3.4 Experiments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 35
3.4.1 Datasets and Tasks . . . . . . . . . . . . . . . . . . . . . . . . 35
3.4.2 Baselines . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 37
3.4.3 Implementation Details . . . . . . . . . . . . . . . . . . . . . . 37
3.4.4 Experimental Results . . . . . . . . . . . . . . . . . . . . . . . 38
3.4.5 Ablation Study . . . . . . . . . . . . . . . . . . . . . . . . . . 38
3.4.6 Parameter Study . . . . . . . . . . . . . . . . . . . . . . . . . 39
3.4.7 Study on Different Balancing Methods . . . . . . . . . . . . . 40
3.4.8 A Closer Look at the Finetuning Stage . . . . . . . . . . . . . 41
3.5 Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 42
4 Training Large Language Models for Biomedical Text Retrieval 43
4.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 44
4.2 Method . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 47
4.2.1 Background of Dense Text Retrieval . . . . . . . . . . . . . . 48
4.2.2 Unsupervised Contrastive Pre-training . . . . . . . . . . . . . 48
4.2.3 Supervised Instruction Fine-tuning . . . . . . . . . . . . . . . 49
4.3 Experimental Results . . . . . . . . . . . . . . . . . . . . . . . . . . . 51
4.3.1 Experimental Setups . . . . . . . . . . . . . . . . . . . . . . . 51
4.3.2 Results on Text Representation Tasks . . . . . . . . . . . . . . 52
4.3.3 Results on Retrieval-Oriented Biomedical Applications . . . . 54
4.3.4 Unsupervised Retrieval Performance . . . . . . . . . . . . . . 54
4.3.5 Studies on Instruction Fine-tuning . . . . . . . . . . . . . . . 56
4.3.6 Effect of Data Volume . . . . . . . . . . . . . . . . . . . . . . 56
4.3.7 Case Study . . . . . . . . . . . . . . . . . . . . . . . . . . . . 57
4.4 Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 59
5 Retrieval-Augmented Clinical Predictions on Electronic Health Records 60
5.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 61
5.2 Methodology . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 63
5.2.1 Problem Setup . . . . . . . . . . . . . . . . . . . . . . . . . . 64
5.2.2 Retrieval Augmentation w/ Medical Codes . . . . . . . . . . . 64
5.2.3 Augmenting Patient Visits with Summarized Knowledge via
Co-training . . . . . . . . . . . . . . . . . . . . . . . . . . . . 67
5.3 Experiments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 69
5.3.1 Experiment Setups . . . . . . . . . . . . . . . . . . . . . . . . 69
5.3.2 Main Experimental Results . . . . . . . . . . . . . . . . . . . 72
5.3.3 Additional Studies . . . . . . . . . . . . . . . . . . . . . . . . 73
5.3.4 Case Study . . . . . . . . . . . . . . . . . . . . . . . . . . . . 75
5.4 Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 75
6 Self-Improving Retrieval-Augmented Language Models for Biomed-
ical Reasoning 76
6.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 77
6.2 Methodology . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 80
6.2.1 Problem Setup . . . . . . . . . . . . . . . . . . . . . . . . . . 80
6.2.2 Stage-I: Retrieval-oriented fine-tuning . . . . . . . . . . . . . . 80
6.2.3 Stage-II: Domain Adaptive Fine-tuning . . . . . . . . . . . . . 81
6.3 Experimental Setup . . . . . . . . . . . . . . . . . . . . . . . . . . . . 83
6.3.1 Tasks and Datasets . . . . . . . . . . . . . . . . . . . . . . . . 83
6.3.2 Baselines . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 84
6.3.3 Implementation Details . . . . . . . . . . . . . . . . . . . . . . 85
6.4 Experimental Results . . . . . . . . . . . . . . . . . . . . . . . . . . . 85
6.4.1 Main Results . . . . . . . . . . . . . . . . . . . . . . . . . . . 85
6.4.2 Ablation Studies . . . . . . . . . . . . . . . . . . . . . . . . . 88
6.4.3 Study on Pseudo-labeled Tuples . . . . . . . . . . . . . . . . . 89
6.4.4 Case Studies . . . . . . . . . . . . . . . . . . . . . . . . . . . . 90
6.5 Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 92
7 Conclusions 93
7.1 Summary of Research Contributions . . . . . . . . . . . . . . . . . . 93
7.1.1 Interpretability for Clinical Predictions on EHR . . . . . . . . 93
7.1.2 Fairness and Robustness under Heterogeneous Data . . . . . . 94
7.1.3 Reliable Clinical Predictions via Retrieval-based Knowledge Aug-
mentation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 94
7.1.4 Retrieval-augmented Language Models for Medical Reasoning 95
7.1.5 Overall Impact . . . . . . . . . . . . . . . . . . . . . . . . . . 95
7.2 Future Works . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 96
7.2.1 Agentic Search and Reasoning for Challenging Medical Questions 96
7.2.2 Scaling Medical Reasoning with Verifiable and Non-Verifiable
Learning Signals . . . . . . . . . . . . . . . . . . . . . . . . . 97
Appendix A Dataset and Task Description in CACHE 99
Appendix B Datasets and Tasks Details for HTP-Star 101
Appendix C Task and Dataset Information in BMRetriever 102
C.1 Pre-training . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 102
C.2 Fine-tuning . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 103
C.3 Baseline Information . . . . . . . . . . . . . . . . . . . . . . . . . . . 105
C.3.1 Baselines for Retrieval Tasks in Main Experiments . . . . . . . 105
C.3.2 Baselines for Retrieval-Oriented Downstream Applications . . 108
Appendix D Prompt Template in Ram-EHR 109
D.1 Knowledge Translating Format in Ram-EHR . . . . . . . . . . . . . 109
D.2 Details for Prompt Design in Ram-EHR . . . . . . . . . . . . . . . . 110
Appendix E Details of Data Used in SimRAG 111
E.1 Training Data Details . . . . . . . . . . . . . . . . . . . . . . . . . . . 111
E.2 Test Data Details . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 112
Appendix F Prompt Details in SimRAG 115
F.1 Prompt for Answer Generation . . . . . . . . . . . . . . . . . . . . . 115
F.2 Prompt for Query Generation . . . . . . . . . . . . . . . . . . . . . . 115
F.3 Prompt for Inference . . . . . . . . . . . . . . . . . . . . . . . . . . . 116
Bibliography 117
About this Dissertation
| School | |
|---|---|
| Department | |
| Degree | |
| Submission | |
| Language |
|
| Research Field | |
| Keyword | |
| Committee Chair / Thesis Advisor | |
| Committee Members |
Primary PDF
| Thumbnail | Title | Date Uploaded | Actions |
|---|---|---|---|
|
File download under embargo until 27 May 2027 | 2026-03-10 21:01:19 -0400 | File download under embargo until 27 May 2027 |
Supplemental Files
| Thumbnail | Title | Date Uploaded | Actions |
|---|