Dataset Viewer
Auto-converted to Parquet Duplicate
model
stringclasses
8 values
prompt
stringclasses
20 values
run
int64
0
4
n_gold
int64
4
14
n_inferred
int64
5
19
n_explicit
int64
4
18
precision_all
float64
0
1
recall
float64
0
1
f1
float64
0
1
precision_explicit
float64
0
1
gemma3:12b
code_review.prompt
0
8
9
8
0.8889
1
0.9412
1
gemma3:12b
content_moderation.prompt
0
8
8
8
1
1
1
1
gemma3:12b
content_moderation.prompt
1
8
8
8
1
1
1
1
gemma3:12b
content_moderation.prompt
2
8
8
8
1
1
1
1
gemma3:12b
customer_support.prompt
0
9
10
9
0.8
0.8889
0.8421
0.8889
gemma3:12b
customer_support.prompt
1
9
10
9
0.8
0.8889
0.8421
0.8889
gemma3:12b
gen_000_triage_support_tickets.prompt
0
6
7
6
0.8571
1
0.9231
1
gemma3:12b
gen_000_triage_support_tickets.prompt
1
6
7
6
0.8571
1
0.9231
1
gemma3:12b
gen_000_triage_support_tickets.prompt
2
6
7
6
0.8571
1
0.9231
1
gemma3:12b
gen_001_review_sql_migrations.prompt
1
8
10
8
0.8
1
0.8889
1
gemma3:12b
gen_001_review_sql_migrations.prompt
2
8
10
8
0.8
1
0.8889
1
gemma3:12b
gen_003_summarize_legal_contracts.prompt
0
7
12
7
0.5833
1
0.7368
1
gemma3:12b
gen_003_summarize_legal_contracts.prompt
1
7
12
7
0.5833
1
0.7368
1
gemma3:12b
gen_003_summarize_legal_contracts.prompt
2
7
12
7
0.5833
1
0.7368
1
gemma3:12b
gen_004_answer_medical_faqs_with_disclaimers.prompt
1
14
15
15
0.9333
1
0.9655
0.9333
gemma3:12b
gen_005_grade_student_essays.prompt
0
4
9
4
0.3333
0.75
0.4615
0.75
gemma3:12b
gen_005_grade_student_essays.prompt
1
4
10
4
0.3
0.75
0.4286
0.75
gemma3:12b
gen_005_grade_student_essays.prompt
2
4
9
4
0.3333
0.75
0.4615
0.75
gemma3:12b
gen_006_route_sales_leads.prompt
0
7
11
7
0.5455
0.8571
0.6667
0.8571
gemma3:12b
gen_006_route_sales_leads.prompt
1
7
11
7
0.5455
0.8571
0.6667
0.8571
gemma3:12b
gen_006_route_sales_leads.prompt
2
7
11
7
0.5455
0.8571
0.6667
0.8571
gemma3:12b
gen_008_screen_job_applications.prompt
0
9
9
9
1
1
1
1
gemma3:12b
gen_008_screen_job_applications.prompt
1
9
9
9
1
1
1
1
gemma3:12b
gen_008_screen_job_applications.prompt
2
9
9
9
1
1
1
1
gemma3:12b
gen_009_translate_with_tone_control.prompt
0
7
12
11
0.5
0.8571
0.6316
0.4545
gemma3:12b
gen_009_translate_with_tone_control.prompt
2
7
12
11
0.5
0.8571
0.6316
0.4545
gemma3:12b
gen_013_recommend_products.prompt
1
7
12
7
0.5
0.8571
0.6316
0.8571
gemma3:12b
gen_013_recommend_products.prompt
2
7
12
7
0.5
0.8571
0.6316
0.8571
gemma3:12b
gen_014_extract_invoice_fields.prompt
0
8
8
8
1
1
1
1
gemma3:12b
gen_014_extract_invoice_fields.prompt
1
8
8
8
1
1
1
1
gemma3:12b
gen_014_extract_invoice_fields.prompt
2
8
8
8
1
1
1
1
gemma3:12b
gen_015_detect_spam_comments.prompt
0
7
8
7
0.875
1
0.9333
1
gemma3:12b
gen_015_detect_spam_comments.prompt
1
7
8
7
0.875
1
0.9333
1
gemma3:12b
gen_015_detect_spam_comments.prompt
2
7
8
7
0.875
1
0.9333
1
glm-4.7
code_review.prompt
1
8
9
8
0.8889
1
0.9412
1
glm-4.7
code_review.prompt
2
8
9
8
0.8889
1
0.9412
1
glm-4.7
content_moderation.prompt
0
8
8
8
1
1
1
1
glm-4.7
content_moderation.prompt
1
8
8
8
1
1
1
1
glm-4.7
content_moderation.prompt
2
8
8
8
1
1
1
1
gpt-oss:120b
code_review.prompt
0
8
8
8
1
1
1
1
gpt-oss:120b
code_review.prompt
1
8
8
8
1
1
1
1
gpt-oss:120b
code_review.prompt
2
8
8
8
1
1
1
1
gpt-oss:120b
content_moderation.prompt
0
8
8
8
1
1
1
1
gpt-oss:120b
content_moderation.prompt
1
8
8
8
1
1
1
1
gpt-oss:120b
content_moderation.prompt
2
8
8
8
1
1
1
1
gpt-oss:120b
customer_support.prompt
0
9
9
9
0.8889
0.8889
0.8889
0.8889
gpt-oss:120b
customer_support.prompt
1
9
9
9
0.8889
0.8889
0.8889
0.8889
gpt-oss:120b
customer_support.prompt
2
9
9
9
0.8889
0.8889
0.8889
0.8889
gpt-oss:120b
gen_000_triage_support_tickets.prompt
0
6
8
6
0.75
1
0.8571
1
gpt-oss:120b
gen_000_triage_support_tickets.prompt
1
6
6
6
0
0
0
0
gpt-oss:120b
gen_000_triage_support_tickets.prompt
2
6
7
7
0.8571
1
0.9231
0.8571
gpt-oss:120b
gen_001_review_sql_migrations.prompt
0
8
11
8
0.7273
1
0.8421
1
gpt-oss:120b
gen_001_review_sql_migrations.prompt
1
8
11
11
0
0
0
0
gpt-oss:120b
gen_001_review_sql_migrations.prompt
2
8
8
8
1
1
1
1
gpt-oss:120b
gen_002_moderate_forum_posts.prompt
0
9
9
9
0
0
0
0
gpt-oss:120b
gen_002_moderate_forum_posts.prompt
1
9
9
9
0
0
0
0
gpt-oss:120b
gen_002_moderate_forum_posts.prompt
2
9
9
9
1
1
1
1
gpt-oss:120b
gen_003_summarize_legal_contracts.prompt
0
7
10
7
0
0
0
0
gpt-oss:120b
gen_003_summarize_legal_contracts.prompt
1
7
11
7
0.6364
1
0.7778
1
gpt-oss:120b
gen_004_answer_medical_faqs_with_disclaimers.prompt
1
14
17
14
0.8235
1
0.9032
1
gpt-oss:120b
gen_004_answer_medical_faqs_with_disclaimers.prompt
2
14
16
14
0
0
0
0
gpt-oss:120b
gen_005_grade_student_essays.prompt
0
4
5
4
0
0
0
0
gpt-oss:120b
gen_005_grade_student_essays.prompt
1
4
6
4
0.6667
1
0.8
1
gpt-oss:120b
gen_005_grade_student_essays.prompt
2
4
5
5
0
0
0
0
gpt-oss:120b
gen_006_route_sales_leads.prompt
0
7
7
7
1
1
1
1
gpt-oss:120b
gen_006_route_sales_leads.prompt
1
7
7
7
1
1
1
1
gpt-oss:120b
gen_006_route_sales_leads.prompt
2
7
9
7
0.7778
1
0.875
1
gpt-oss:120b
gen_007_generate_release_notes.prompt
0
7
7
7
1
1
1
1
gpt-oss:120b
gen_007_generate_release_notes.prompt
1
7
7
7
0
0
0
0
gpt-oss:120b
gen_007_generate_release_notes.prompt
2
7
7
7
0
0
0
0
gpt-oss:120b
gen_008_screen_job_applications.prompt
0
9
9
9
1
1
1
1
gpt-oss:120b
gen_008_screen_job_applications.prompt
2
9
9
9
0
0
0
0
gpt-oss:120b
gen_009_translate_with_tone_control.prompt
0
7
7
7
0
0
0
0
gpt-oss:120b
gen_009_translate_with_tone_control.prompt
1
7
7
7
0
0
0
0
gpt-oss:120b
gen_009_translate_with_tone_control.prompt
2
7
7
7
1
1
1
1
gpt-oss:120b
gen_010_classify_incoming_emails.prompt
0
7
8
7
0
0
0
0
gpt-oss:120b
gen_010_classify_incoming_emails.prompt
1
7
8
8
0
0
0
0
gpt-oss:120b
gen_010_classify_incoming_emails.prompt
2
7
7
7
1
1
1
1
gpt-oss:120b
gen_011_draft_incident_postmortems.prompt
0
7
7
7
0
0
0
0
gpt-oss:120b
gen_012_review_pull_requests.prompt
0
6
10
10
0
0
0
0
gpt-oss:120b
gen_012_review_pull_requests.prompt
1
6
10
10
0.6
1
0.75
0.6
gpt-oss:120b
gen_012_review_pull_requests.prompt
2
6
6
6
1
1
1
1
gpt-oss:120b
gen_013_recommend_products.prompt
0
7
12
7
0
0
0
0
gpt-oss:120b
gen_013_recommend_products.prompt
1
7
9
7
0.7778
1
0.875
1
gpt-oss:120b
gen_013_recommend_products.prompt
2
7
7
7
1
1
1
1
gpt-oss:120b
gen_014_extract_invoice_fields.prompt
0
8
11
8
0.7273
1
0.8421
1
gpt-oss:120b
gen_014_extract_invoice_fields.prompt
1
8
8
8
1
1
1
1
gpt-oss:120b
gen_014_extract_invoice_fields.prompt
2
8
8
8
0
0
0
0
gpt-oss:120b
gen_015_detect_spam_comments.prompt
0
7
7
7
0
0
0
0
gpt-oss:120b
gen_015_detect_spam_comments.prompt
1
7
11
7
0.6364
1
0.7778
1
gpt-oss:120b
gen_015_detect_spam_comments.prompt
2
7
7
7
1
1
1
1
gpt-oss:120b
gen_016_answer_hr_policy_questions.prompt
1
12
12
12
0
0
0
0
minimax-m2.1
code_review.prompt
0
8
8
8
1
1
1
1
minimax-m2.1
code_review.prompt
1
8
8
8
1
1
1
1
minimax-m2.1
code_review.prompt
2
8
8
8
1
1
1
1
minimax-m2.1
code_review.prompt
3
8
8
8
1
1
1
1
minimax-m2.1
code_review.prompt
4
8
8
8
1
1
1
1
minimax-m2.1
content_moderation.prompt
0
8
8
8
1
1
1
1
minimax-m2.1
content_moderation.prompt
1
8
8
8
1
1
1
1
minimax-m2.1
content_moderation.prompt
2
8
8
8
1
1
1
1
End of preview. Expand in Data Studio

LLM Evaluation Self-Audit

Data for How Reproducible Are Evaluation Conclusions? A Self-Audit of LLM-Inferred Prompt Structure, Dipankar Sarkar, Skelf Research.

Code and paper source: github.com/sarkar-dipankar/llm-evaluation-self-audit (this package is built from commit 687ee8a).

The finding

We asked eight open model variants to infer the structure of a prompt, then asked again with the identical call. Caching was off.

They often gave a different answer.

  • Mean node-set Jaccard across repeated identical calls ranges from 0.39 (gpt-oss:120b) to 0.96 (minimax-m2.1).
  • Only 35 of 127 prompt-model cells were node-set-perfect on every run. So 72% never were.
  • We then ranked the eight models by reproducibility and bootstrapped the ranking over prompts (10,000 replicates, seed 20260817). The bottom holds. The two least reproducible models keep their rank in 99% and 86% of replicates. The middle four keep theirs in 27% to 48%. The top two in 68% each. The table finds the worst model. It does not reliably find the best.
  • Reproducible is not the same as correct. Mean F1 against ground-truth annotations ranges from 0.56 (gpt-oss:120b) to 0.99 (minimax-m2.1) over 293 persisted runs.
  • On the audit date, 2026-08-17, 4 of the 8 evaluated variants returned HTTP 410 (retired). The study as specified can no longer be re-run. The raw outputs here are what survives.

Every number above is recomputed from the files in this repo: ir_stability_cells, ir_stability_summary, ir_correctness, model_availability, and raw/results/stats_analysis.json.

If a small evaluation can look this definitive on this little evidence, how many published leaderboards would survive the same audit?

Configs

Nine viewer configs, all JSONL, built by build_hf.py (in this repo, stdlib only, deterministic). Six are the release's results/*.csv files converted row for row, one typed value per column. The original CSVs are unchanged under raw/results/.

Config File Rows What it is
ir_correctness (default) data/ir_correctness.jsonl 293 Each persisted inferred structure (IR) scored against gold annotations: precision, recall, F1, explicit-node precision. One row per model, prompt, run.
ir_stability_cells data/ir_stability_cells.jsonl 127 One row per prompt-model cell: mean Jaccard, ID stability, perfect flag, and the campaign CSV the cell came from. Source: results/ir_stability_merged_cells.csv.
ir_stability_summary data/ir_stability_summary.jsonl 8 Per-model reproducibility table (the ranked table the paper audits). Source: results/ir_stability_merged.csv.
drift data/drift.jsonl 98 What changes when an IR drifts: id-label, count, or metadata disagreement per cell. Source: results/drift_decompose.csv. The column hierarchy_drift is what the paper calls metadata drift, the IR is a flat node list.
downstream_cost data/downstream_cost.jsonl 98 Node-count trajectory per cell and the run at which its rolling mean stabilises. The paper documents two specification errors in this analysis, read its appendix before using stabilises_at.
model_availability data/model_availability.jsonl 15 Live endpoint check on 2026-08-17: HTTP status and provider message for every model the study named, tried, or probed.
ir_runs data/ir_runs.jsonl 850 Every raw IR file. 293 from results/irs (the main study) and 557 from results/irs_crossed (the crossed run on public SKILL.md prompts). The full IR is in ir as a JSON string.
behaviour data/behaviour.jsonl 80 Companion mutation study: original and mutated rendered prompts, the mutation operator, and model responses to each. Response lists are JSON strings.
corpus data/corpus.jsonl 123 The prompts. 83 synthetic templates and 40 SKILL.md files from public GitHub repos, 12 of them with hand annotations. Contexts are a JSON string.

Column notes

  • ir_runs.model is the on-disk directory slug (gemma3_12b), the CSVs use the provider name (gemma3:12b). run is parsed from the _r<N> filename suffix. n_nodes and confidence are lifted from the IR for filtering.
  • corpus.commit is the pinned revision parsed from html_url. corpus.git_blob_sha is the MANIFEST column named commit, which holds the file's git blob sha (it matches git hash-object), not a commit.
  • Nested values are serialised as JSON strings so each column has one type. Parse them with json.loads.

raw/

raw/results/ and raw/corpus/ are exact copies of the release directories, minus four collection *.log files. No config points at them. They hold every CSV the paper uses (including the per-campaign ir_stability_*.csv files), stats_analysis.json, the coverage and mutation outputs, and the annotation notes. Nothing from the analysis is dropped.

Reproduce

Clone the GitHub repo and run from its root:

python3 -m venv .venv
. .venv/bin/activate
pip install -r requirements.txt
(cd tool && cargo build --release -p promptdbg)   # needed by ir_correctness.py
python scripts/merge_ir_stability.py
python scripts/stats_analysis.py
python scripts/drift_decompose.py
python scripts/downstream_cost.py
python scripts/ir_correctness.py
python scripts/stability_vs_correctness.py

ir_correctness.py calls the promptdbg binary in tool/target/release/, so build it first (Rust toolchain, about a minute). These scripts replay the saved outputs offline.

The live checks in model_availability.py and failure_attribution.py need provider access and cannot bring back retired endpoints. The saved responses make the analysis repeatable. They do not make the collection repeatable.

To rebuild this package from a checkout:

python3 build_hf.py /path/to/llm-evaluation-self-audit .

Licensing

  • Author material (measurements, synthetic prompts, annotations, analysis outputs, this card) is released under CC BY 4.0.
  • The 40 files in corpus/github/ (rows with source: github in the corpus config) are third-party SKILL.md files. They keep their upstream licenses: 38 MIT, 1 Apache-2.0, 1 Unlicense. Each row carries license, repo, path, commit, and html_url, from corpus/MANIFEST.csv. CC BY 4.0 does not replace those terms.

Citation

@misc{sarkar2026reproducible,
  title = {How Reproducible Are Evaluation Conclusions? A Self-Audit of LLM-Inferred Prompt Structure},
  author = {Dipankar Sarkar},
  year = {2026},
  eprint = {2609.30074},
  archivePrefix = {arXiv},
  primaryClass = {cs.CL},
  doi = {10.48550/arXiv.2609.30074},
  url = {https://arxiv.org/abs/2609.30074}
}

Links

Downloads last month
293

Paper for dipankarsarkar/llm-evaluation-self-audit