Stress Testing Instruction Following in Large Language Models Open Access
Reicin, Noah (Spring 2026)
Abstract
Large Language Models (LLMs) are increasingly deployed in complex, multi-step
workflows, yet their ability to maintain ordered execution across many steps remains
underexplored. This thesis develops and extends RIFT (Reordered Instruction
Following Testbed), an evaluation framework that disentangles prompt structure from
task content using rephrased Jeopardy! question–answer pairs, assessing instruction
following under linear (sequential) and jumping (non-sequential) prompt structures.
Phase 1 established across 10,000 evaluations spanning six open-source LLMs that
accuracy drops by up to 72% under jumping conditions relative to baseline, with
median accuracy near zero on many runs. Error analysis showed that approximately
33 to 60% of failures stem from instruction-order violations rather than knowledge
errors.
Phase 2 extends this work by testing whether prompt-tuning interventions can restore
execution control under the jumping condition at scales up to 300 questions per
prompt. Four prompt families are evaluated: current (baseline), plan-first, explicit-
path, and generate-path. A key finding is the divergence between coverage, whether
the model attempts all required steps, and accuracy, whether it executes them cor-
rectly. Structured prompts raise coverage from ∼21% to ∼90%, largely solving the
completion problem, while accuracy improves only from ∼2% to ∼22%, revealing that
order failure and completion failure are largely independent. A further finding, the
path representation effect, shows that model-generated traversal paths outperform
algorithmically supplied paths of identical format by ∼17%, suggesting execution
fidelity depends on the process by which structural representations are produced
Table of Contents
Contents
1 Introduction 1
1.1 Research Questions . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5
1.2 Thesis Statement . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5
1.3 Overview of Contributions . . . . . . . . . . . . . . . . . . . . . . . . 6
2 Related Work 7
2.1 ComplexBench . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7
2.2 Reasoning-Oriented Prompting . . . . . . . . . . . . . . . . . . . . . 8
2.3 Long-Context Benchmarks and Positional Bias . . . . . . . . . . . . . 8
2.4 Instruction Tuning and Execution Control . . . . . . . . . . . . . . . 9
2.5 Benchmark Limitations and Structural Gaps . . . . . . . . . . . . . . 10
3 Approach 11
3.1 Overview . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11
3.2 Simplicity at the Item Level . . . . . . . . . . . . . . . . . . . . . . . 11
3.3 Conceptual Framework . . . . . . . . . . . . . . . . . . . . . . . . . . 12
3.4 Prompt Construction . . . . . . . . . . . . . . . . . . . . . . . . . . . 13
3.4.1 Phase 2 Prompt Family Development . . . . . . . . . . . . . . 13
3.5 Evaluation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 15
3.5.1 Phase 1 Evaluation . . . . . . . . . . . . . . . . . . . . . . . . 15
3.5.2 Phase 2 Evaluation . . . . . . . . . . . . . . . . . . . . . . . . 15
3.6 Compute and Infrastructure . . . . . . . . . . . . . . . . . . . . . . . 16
4 Experiments 18
i
4.1 Source: Jeopardy! Dataset . . . . . . . . . . . . . . . . . . . . . . . . 18
4.2 Rephrasing Pipeline . . . . . . . . . . . . . . . . . . . . . . . . . . . . 18
4.3 Preprocessing . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 19
4.4 Content Difficulty Analysis . . . . . . . . . . . . . . . . . . . . . . . . 20
4.5 Phase 1: RIFT Baseline Experiments . . . . . . . . . . . . . . . . . . 21
4.5.1 Experimental Settings . . . . . . . . . . . . . . . . . . . . . . 21
4.5.2 RIFT Results . . . . . . . . . . . . . . . . . . . . . . . . . . . 23
4.5.3 Effect of Reasoning Supervision . . . . . . . . . . . . . . . . . 24
4.5.4 Prompt-Length Effects and Effective Context Limits . . . . . 25
4.5.5 Accuracy vs. Prompt Depth . . . . . . . . . . . . . . . . . . . 26
4.5.6 Error Analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . 27
4.5.7 Effect of Jump Distance . . . . . . . . . . . . . . . . . . . . . 28
4.6 Phase 2: Prompt-Tuning Interventions . . . . . . . . . . . . . . . . . 29
4.6.1 Motivation and Extension . . . . . . . . . . . . . . . . . . . . 29
4.6.2 Prompt Families . . . . . . . . . . . . . . . . . . . . . . . . . 30
4.6.3 Metrics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 31
4.6.4 Overall Performance . . . . . . . . . . . . . . . . . . . . . . . 31
4.6.5 Scaling Behavior . . . . . . . . . . . . . . . . . . . . . . . . . 32
4.6.6 Failure Mode Transformation . . . . . . . . . . . . . . . . . . 33
5 Analysis 35
5.1 Taxonomy of Execution Failure . . . . . . . . . . . . . . . . . . . . . 35
5.2 Per-Position Accuracy and Error Progression . . . . . . . . . . . . . . 37
5.3 Sequence-Level Analysis . . . . . . . . . . . . . . . . . . . . . . . . . 39
5.4 Path Generation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 40
5.5 Phase 2 Takeaway . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 42
5.6 Limitations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 42
6 Conclusion 44
A Prompt Templates 46
A.1 Current . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 46
A.2 Plan-First . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 47
A.3 Explicit-Path . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 48
A.4 Generate-Path . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 50
Bibliography 53
About this Honors Thesis
| School | |
|---|---|
| Department | |
| Degree | |
| Submission | |
| Language |
|
| Research Field | |
| Keyword | |
| Committee Chair / Thesis Advisor | |
| Committee Members |
Primary PDF
| Thumbnail | Title | Date Uploaded | Actions |
|---|---|---|---|
|
|
Stress Testing Instruction Following in Large Language Models () | 2026-04-05 15:19:31 -0400 |
|
Supplemental Files
| Thumbnail | Title | Date Uploaded | Actions |
|---|