Stress Testing Instruction Following in Large Language Models Open Access

Reicin, Noah (Spring 2026)

Permanent URL: https://etd.library.emory.edu/concern/etds/08612q018?locale=en
Published

Abstract

Large Language Models (LLMs) are increasingly deployed in complex, multi-step

workflows, yet their ability to maintain ordered execution across many steps remains

underexplored. This thesis develops and extends RIFT (Reordered Instruction

Following Testbed), an evaluation framework that disentangles prompt structure from

task content using rephrased Jeopardy! question–answer pairs, assessing instruction

following under linear (sequential) and jumping (non-sequential) prompt structures.

Phase 1 established across 10,000 evaluations spanning six open-source LLMs that

accuracy drops by up to 72% under jumping conditions relative to baseline, with

median accuracy near zero on many runs. Error analysis showed that approximately

33 to 60% of failures stem from instruction-order violations rather than knowledge

errors.

Phase 2 extends this work by testing whether prompt-tuning interventions can restore

execution control under the jumping condition at scales up to 300 questions per

prompt. Four prompt families are evaluated: current (baseline), plan-first, explicit-

path, and generate-path. A key finding is the divergence between coverage, whether

the model attempts all required steps, and accuracy, whether it executes them cor-

rectly. Structured prompts raise coverage from ∼21% to ∼90%, largely solving the

completion problem, while accuracy improves only from ∼2% to ∼22%, revealing that

order failure and completion failure are largely independent. A further finding, the

path representation effect, shows that model-generated traversal paths outperform

algorithmically supplied paths of identical format by ∼17%, suggesting execution

fidelity depends on the process by which structural representations are produced

Table of Contents

Contents

1 Introduction 1

1.1 Research Questions . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5

1.2 Thesis Statement . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5

1.3 Overview of Contributions . . . . . . . . . . . . . . . . . . . . . . . . 6

2 Related Work 7

2.1 ComplexBench . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7

2.2 Reasoning-Oriented Prompting . . . . . . . . . . . . . . . . . . . . . 8

2.3 Long-Context Benchmarks and Positional Bias . . . . . . . . . . . . . 8

2.4 Instruction Tuning and Execution Control . . . . . . . . . . . . . . . 9

2.5 Benchmark Limitations and Structural Gaps . . . . . . . . . . . . . . 10

3 Approach 11

3.1 Overview . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11

3.2 Simplicity at the Item Level . . . . . . . . . . . . . . . . . . . . . . . 11

3.3 Conceptual Framework . . . . . . . . . . . . . . . . . . . . . . . . . . 12

3.4 Prompt Construction . . . . . . . . . . . . . . . . . . . . . . . . . . . 13

3.4.1 Phase 2 Prompt Family Development . . . . . . . . . . . . . . 13

3.5 Evaluation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 15

3.5.1 Phase 1 Evaluation . . . . . . . . . . . . . . . . . . . . . . . . 15

3.5.2 Phase 2 Evaluation . . . . . . . . . . . . . . . . . . . . . . . . 15

3.6 Compute and Infrastructure . . . . . . . . . . . . . . . . . . . . . . . 16

4 Experiments 18

i

4.1 Source: Jeopardy! Dataset . . . . . . . . . . . . . . . . . . . . . . . . 18

4.2 Rephrasing Pipeline . . . . . . . . . . . . . . . . . . . . . . . . . . . . 18

4.3 Preprocessing . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 19

4.4 Content Difficulty Analysis . . . . . . . . . . . . . . . . . . . . . . . . 20

4.5 Phase 1: RIFT Baseline Experiments . . . . . . . . . . . . . . . . . . 21

4.5.1 Experimental Settings . . . . . . . . . . . . . . . . . . . . . . 21

4.5.2 RIFT Results . . . . . . . . . . . . . . . . . . . . . . . . . . . 23

4.5.3 Effect of Reasoning Supervision . . . . . . . . . . . . . . . . . 24

4.5.4 Prompt-Length Effects and Effective Context Limits . . . . . 25

4.5.5 Accuracy vs. Prompt Depth . . . . . . . . . . . . . . . . . . . 26

4.5.6 Error Analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . 27

4.5.7 Effect of Jump Distance . . . . . . . . . . . . . . . . . . . . . 28

4.6 Phase 2: Prompt-Tuning Interventions . . . . . . . . . . . . . . . . . 29

4.6.1 Motivation and Extension . . . . . . . . . . . . . . . . . . . . 29

4.6.2 Prompt Families . . . . . . . . . . . . . . . . . . . . . . . . . 30

4.6.3 Metrics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 31

4.6.4 Overall Performance . . . . . . . . . . . . . . . . . . . . . . . 31

4.6.5 Scaling Behavior . . . . . . . . . . . . . . . . . . . . . . . . . 32

4.6.6 Failure Mode Transformation . . . . . . . . . . . . . . . . . . 33

5 Analysis 35

5.1 Taxonomy of Execution Failure . . . . . . . . . . . . . . . . . . . . . 35

5.2 Per-Position Accuracy and Error Progression . . . . . . . . . . . . . . 37

5.3 Sequence-Level Analysis . . . . . . . . . . . . . . . . . . . . . . . . . 39

5.4 Path Generation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 40

5.5 Phase 2 Takeaway . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 42

5.6 Limitations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 42

6 Conclusion 44

A Prompt Templates 46

A.1 Current . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 46

A.2 Plan-First . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 47

A.3 Explicit-Path . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 48

A.4 Generate-Path . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 50

Bibliography 53

About this Honors Thesis

Rights statement
  • Permission granted by the author to include this thesis or dissertation in this repository. All rights reserved by the author. Please contact the author for information regarding the reproduction and use of this thesis or dissertation.
School
Department
Degree
Submission
Language
  • English
Research Field
Keyword
Committee Chair / Thesis Advisor
Committee Members
Last modified

Primary PDF

Supplemental Files