StringNLP Lab University of Alberta

LongHarness Bench: documents routed through a language-model harness.

LongHarnessBench

Stress-testing language model harnesses for long-context reasoning

Quang Hieu Pham* · Thuy Duong Nguyen* · Jocelyn Qiaochu Chen† · Xi Ye†

University of Alberta · Alberta Machine Intelligence Institute (Amii) · *Equal contribution · †Equal advising

Harness choice substantially affects accuracy and cost.

Across 25 model–harness combinations, macro-average accuracy spans 0%–68%. No harness leads on every task, and higher cost does not guarantee higher accuracy.

All 25 evaluated combinations.

Macro-average exact accuracy (%) and estimated cost per instance (USD) by model and harness
ModelDirectRLMOpenCodemini-sweReAct
GPT-5.6-sol60.0%$0.8253.5%$4.3040.5%$2.9068.0%$0.8243.0%$1.93
GLM-5.39.5%$0.9527.0%$5.4838.0%$3.3342.5%$3.0013.0%$3.19
Qwen3.8-27B0.0%$0.347.5%$5.6638.5%$7.0633.0%$5.320.0%$3.87
Gemini 3.8 Flash10.0%$0.1610.5%$0.7111.5%$0.4725.0%$0.5521.0%$0.54
Kimi-K2.60.0%$0.640.5%$6.651.5%$5.650.5%$4.030.0%$2.86
Accuracy ↑0%70%
Cost ↓$8$0
Darker is better: higher accuracy (blue), lower cost (teal). “Both” displays them in separate halves of each cell. Values are task averages; cost is USD per instance, rounded to cents. Direct = direct reading; mini-swe = mini-swe-agent. Explore accuracy and cost ↗
~128K
tokens per context
200
reasoning challenges
25
model–harness combinations
68%
best macro-average accuracy

Evaluating adaptive reasoning and efficiency.

Harnesses control how models search, use tools, and make additional model calls over long contexts.

Retrieval must adapt as reasoning unfolds.

Tasks solved through simple retrieval or independent context chunks reveal little about this ability. LongHarness uses confusable evidence and multi-step problems where intermediate results determine what to inspect next.

Accuracy alone does not capture efficiency.

Multiple strategies can solve a task while using very different amounts of computation. LongHarness measures accuracy alongside token use and estimated cost to show what a harness gains from the computation it spends.

A wide spread in accuracy and cost.

LongHarness reveals bigger differences between harnesses and shows that similar accuracy can come at a dramatically different cost.

All 25 model–harness combinations plotted by macro-average accuracy and estimated cost. Color identifies each of five models; marker shape identifies direct reading or one of four agent harnesses. Accuracy ranges from 0 to 68 percent, and mean cost ranges from $0.162 to $7.06 per instance.
All 25 evaluated combinations. Each point averages accuracy and estimated cost across the four tasks. Higher and further left means more accurate and less expensive. Cost uses a logarithmic scale. Download the data.
Four dataset plots: Program Execution Tracing, Outlier Memo Detection, Constraint Solving Search, and Equivalent Program Pair Search. Each shows all 25 configurations on shared accuracy and logarithmic cost axes, with model colors matching the paper and distinct harness shapes.
Each dataset contains 50 instances. All four plots share the same scales; higher and further left is better. Download the data.

Macro-average accuracy

Percent across all LongHarness tasks. Best overall result in bold.

Model Direct Reading RLM OpenCode mini-swe-agent ReAct
GPT-5.6-sol60.053.540.568.043.0
GLM-5.39.527.038.042.513.0
Qwen3.8-27B0.07.538.533.00.0
Gemini 3.8 Flash10.010.511.525.021.0
Kimi-K2.60.00.51.50.50.0

Built to test how reasoning unfolds.

Comparison of long-context benchmarks along properties relevant to evaluating language-model harnesses.
Evaluation property LongHarness LongBench-v2 HELMET LongProc OOLONG
Diverse, adaptive retrieval demand ✓ ~ ~ × ×
Semantically confusable evidence ✓ × ✓ ~ ~
Extensive reasoning steps ✓ × × ✓ ✓
Step-dependent retrieval ✓ × × ✓ ×
Multiple strategies with different costs ✓ × ~ × ~
Wide range of SoTA system accuracy ✓ × × × ×
Wide range of SoTA system cost ✓ × × × ~
✓ Supported ~ Partially supported × Not supported

Adaptive retrieval means choosing among lexical search, semantic retrieval, and direct reading as information needs change. Step-dependent retrieval means intermediate findings determine what must be retrieved next. The final two rows report whether model–harness configurations are widely separated in accuracy and cost.

Find. Verify. Connect. Reason.

Four complementary tasks, each with 50 instances, deterministic exact-match evaluation, and contexts near 128K tokens.

01

Constraint Solving Search

Find every person satisfying three to five conditions using facts scattered across varied documents. Near-matches satisfy all but one condition, and the most selective condition is not identified.

Context
~2,400 documents about 160 people
Answer
5 person IDs
Strategy
Narrow, then verify
Constraint Solving Search example with supporting facts, distractors, a multi-condition query, and correct and mistaken traces.
Constraint Solving Search: combine supporting facts while rejecting plausible near-matches.

The scoring script is available on both GitHub and Hugging Face. GitHub includes two examples per task; Hugging Face adds the complete 200-instance release, answer keys, and manifests.

Hugging Face CLI hf download StringNLP/longharness --repo-type dataset --local-dir longharness

Private access is required during release preparation.

Cite LongHarness

If you use LongHarness in your research, please cite our paper.

Read the paper on arXiv
@misc{pham2026longharness,
  title  = {LongHarness Bench: Stress-Testing Language Model
            Harnesses for Long-Context Reasoning},
  author = {Pham, Quang Hieu and Nguyen, Thuy Duong and
            Chen, Jocelyn Qiaochu and Ye, Xi},
  year   = {2026},
  eprint = {2609.38137},
  archivePrefix = {arXiv},
  primaryClass = {cs.CL},
  url    = {https://arxiv.org/abs/2609.38137}
}