Datasets:
model stringclasses 8
values | prompt stringclasses 20
values | run int64 0 4 | n_gold int64 4 14 | n_inferred int64 5 19 | n_explicit int64 4 18 | precision_all float64 0 1 | recall float64 0 1 | f1 float64 0 1 | precision_explicit float64 0 1 |
|---|---|---|---|---|---|---|---|---|---|
gemma3:12b | code_review.prompt | 0 | 8 | 9 | 8 | 0.8889 | 1 | 0.9412 | 1 |
gemma3:12b | content_moderation.prompt | 0 | 8 | 8 | 8 | 1 | 1 | 1 | 1 |
gemma3:12b | content_moderation.prompt | 1 | 8 | 8 | 8 | 1 | 1 | 1 | 1 |
gemma3:12b | content_moderation.prompt | 2 | 8 | 8 | 8 | 1 | 1 | 1 | 1 |
gemma3:12b | customer_support.prompt | 0 | 9 | 10 | 9 | 0.8 | 0.8889 | 0.8421 | 0.8889 |
gemma3:12b | customer_support.prompt | 1 | 9 | 10 | 9 | 0.8 | 0.8889 | 0.8421 | 0.8889 |
gemma3:12b | gen_000_triage_support_tickets.prompt | 0 | 6 | 7 | 6 | 0.8571 | 1 | 0.9231 | 1 |
gemma3:12b | gen_000_triage_support_tickets.prompt | 1 | 6 | 7 | 6 | 0.8571 | 1 | 0.9231 | 1 |
gemma3:12b | gen_000_triage_support_tickets.prompt | 2 | 6 | 7 | 6 | 0.8571 | 1 | 0.9231 | 1 |
gemma3:12b | gen_001_review_sql_migrations.prompt | 1 | 8 | 10 | 8 | 0.8 | 1 | 0.8889 | 1 |
gemma3:12b | gen_001_review_sql_migrations.prompt | 2 | 8 | 10 | 8 | 0.8 | 1 | 0.8889 | 1 |
gemma3:12b | gen_003_summarize_legal_contracts.prompt | 0 | 7 | 12 | 7 | 0.5833 | 1 | 0.7368 | 1 |
gemma3:12b | gen_003_summarize_legal_contracts.prompt | 1 | 7 | 12 | 7 | 0.5833 | 1 | 0.7368 | 1 |
gemma3:12b | gen_003_summarize_legal_contracts.prompt | 2 | 7 | 12 | 7 | 0.5833 | 1 | 0.7368 | 1 |
gemma3:12b | gen_004_answer_medical_faqs_with_disclaimers.prompt | 1 | 14 | 15 | 15 | 0.9333 | 1 | 0.9655 | 0.9333 |
gemma3:12b | gen_005_grade_student_essays.prompt | 0 | 4 | 9 | 4 | 0.3333 | 0.75 | 0.4615 | 0.75 |
gemma3:12b | gen_005_grade_student_essays.prompt | 1 | 4 | 10 | 4 | 0.3 | 0.75 | 0.4286 | 0.75 |
gemma3:12b | gen_005_grade_student_essays.prompt | 2 | 4 | 9 | 4 | 0.3333 | 0.75 | 0.4615 | 0.75 |
gemma3:12b | gen_006_route_sales_leads.prompt | 0 | 7 | 11 | 7 | 0.5455 | 0.8571 | 0.6667 | 0.8571 |
gemma3:12b | gen_006_route_sales_leads.prompt | 1 | 7 | 11 | 7 | 0.5455 | 0.8571 | 0.6667 | 0.8571 |
gemma3:12b | gen_006_route_sales_leads.prompt | 2 | 7 | 11 | 7 | 0.5455 | 0.8571 | 0.6667 | 0.8571 |
gemma3:12b | gen_008_screen_job_applications.prompt | 0 | 9 | 9 | 9 | 1 | 1 | 1 | 1 |
gemma3:12b | gen_008_screen_job_applications.prompt | 1 | 9 | 9 | 9 | 1 | 1 | 1 | 1 |
gemma3:12b | gen_008_screen_job_applications.prompt | 2 | 9 | 9 | 9 | 1 | 1 | 1 | 1 |
gemma3:12b | gen_009_translate_with_tone_control.prompt | 0 | 7 | 12 | 11 | 0.5 | 0.8571 | 0.6316 | 0.4545 |
gemma3:12b | gen_009_translate_with_tone_control.prompt | 2 | 7 | 12 | 11 | 0.5 | 0.8571 | 0.6316 | 0.4545 |
gemma3:12b | gen_013_recommend_products.prompt | 1 | 7 | 12 | 7 | 0.5 | 0.8571 | 0.6316 | 0.8571 |
gemma3:12b | gen_013_recommend_products.prompt | 2 | 7 | 12 | 7 | 0.5 | 0.8571 | 0.6316 | 0.8571 |
gemma3:12b | gen_014_extract_invoice_fields.prompt | 0 | 8 | 8 | 8 | 1 | 1 | 1 | 1 |
gemma3:12b | gen_014_extract_invoice_fields.prompt | 1 | 8 | 8 | 8 | 1 | 1 | 1 | 1 |
gemma3:12b | gen_014_extract_invoice_fields.prompt | 2 | 8 | 8 | 8 | 1 | 1 | 1 | 1 |
gemma3:12b | gen_015_detect_spam_comments.prompt | 0 | 7 | 8 | 7 | 0.875 | 1 | 0.9333 | 1 |
gemma3:12b | gen_015_detect_spam_comments.prompt | 1 | 7 | 8 | 7 | 0.875 | 1 | 0.9333 | 1 |
gemma3:12b | gen_015_detect_spam_comments.prompt | 2 | 7 | 8 | 7 | 0.875 | 1 | 0.9333 | 1 |
glm-4.7 | code_review.prompt | 1 | 8 | 9 | 8 | 0.8889 | 1 | 0.9412 | 1 |
glm-4.7 | code_review.prompt | 2 | 8 | 9 | 8 | 0.8889 | 1 | 0.9412 | 1 |
glm-4.7 | content_moderation.prompt | 0 | 8 | 8 | 8 | 1 | 1 | 1 | 1 |
glm-4.7 | content_moderation.prompt | 1 | 8 | 8 | 8 | 1 | 1 | 1 | 1 |
glm-4.7 | content_moderation.prompt | 2 | 8 | 8 | 8 | 1 | 1 | 1 | 1 |
gpt-oss:120b | code_review.prompt | 0 | 8 | 8 | 8 | 1 | 1 | 1 | 1 |
gpt-oss:120b | code_review.prompt | 1 | 8 | 8 | 8 | 1 | 1 | 1 | 1 |
gpt-oss:120b | code_review.prompt | 2 | 8 | 8 | 8 | 1 | 1 | 1 | 1 |
gpt-oss:120b | content_moderation.prompt | 0 | 8 | 8 | 8 | 1 | 1 | 1 | 1 |
gpt-oss:120b | content_moderation.prompt | 1 | 8 | 8 | 8 | 1 | 1 | 1 | 1 |
gpt-oss:120b | content_moderation.prompt | 2 | 8 | 8 | 8 | 1 | 1 | 1 | 1 |
gpt-oss:120b | customer_support.prompt | 0 | 9 | 9 | 9 | 0.8889 | 0.8889 | 0.8889 | 0.8889 |
gpt-oss:120b | customer_support.prompt | 1 | 9 | 9 | 9 | 0.8889 | 0.8889 | 0.8889 | 0.8889 |
gpt-oss:120b | customer_support.prompt | 2 | 9 | 9 | 9 | 0.8889 | 0.8889 | 0.8889 | 0.8889 |
gpt-oss:120b | gen_000_triage_support_tickets.prompt | 0 | 6 | 8 | 6 | 0.75 | 1 | 0.8571 | 1 |
gpt-oss:120b | gen_000_triage_support_tickets.prompt | 1 | 6 | 6 | 6 | 0 | 0 | 0 | 0 |
gpt-oss:120b | gen_000_triage_support_tickets.prompt | 2 | 6 | 7 | 7 | 0.8571 | 1 | 0.9231 | 0.8571 |
gpt-oss:120b | gen_001_review_sql_migrations.prompt | 0 | 8 | 11 | 8 | 0.7273 | 1 | 0.8421 | 1 |
gpt-oss:120b | gen_001_review_sql_migrations.prompt | 1 | 8 | 11 | 11 | 0 | 0 | 0 | 0 |
gpt-oss:120b | gen_001_review_sql_migrations.prompt | 2 | 8 | 8 | 8 | 1 | 1 | 1 | 1 |
gpt-oss:120b | gen_002_moderate_forum_posts.prompt | 0 | 9 | 9 | 9 | 0 | 0 | 0 | 0 |
gpt-oss:120b | gen_002_moderate_forum_posts.prompt | 1 | 9 | 9 | 9 | 0 | 0 | 0 | 0 |
gpt-oss:120b | gen_002_moderate_forum_posts.prompt | 2 | 9 | 9 | 9 | 1 | 1 | 1 | 1 |
gpt-oss:120b | gen_003_summarize_legal_contracts.prompt | 0 | 7 | 10 | 7 | 0 | 0 | 0 | 0 |
gpt-oss:120b | gen_003_summarize_legal_contracts.prompt | 1 | 7 | 11 | 7 | 0.6364 | 1 | 0.7778 | 1 |
gpt-oss:120b | gen_004_answer_medical_faqs_with_disclaimers.prompt | 1 | 14 | 17 | 14 | 0.8235 | 1 | 0.9032 | 1 |
gpt-oss:120b | gen_004_answer_medical_faqs_with_disclaimers.prompt | 2 | 14 | 16 | 14 | 0 | 0 | 0 | 0 |
gpt-oss:120b | gen_005_grade_student_essays.prompt | 0 | 4 | 5 | 4 | 0 | 0 | 0 | 0 |
gpt-oss:120b | gen_005_grade_student_essays.prompt | 1 | 4 | 6 | 4 | 0.6667 | 1 | 0.8 | 1 |
gpt-oss:120b | gen_005_grade_student_essays.prompt | 2 | 4 | 5 | 5 | 0 | 0 | 0 | 0 |
gpt-oss:120b | gen_006_route_sales_leads.prompt | 0 | 7 | 7 | 7 | 1 | 1 | 1 | 1 |
gpt-oss:120b | gen_006_route_sales_leads.prompt | 1 | 7 | 7 | 7 | 1 | 1 | 1 | 1 |
gpt-oss:120b | gen_006_route_sales_leads.prompt | 2 | 7 | 9 | 7 | 0.7778 | 1 | 0.875 | 1 |
gpt-oss:120b | gen_007_generate_release_notes.prompt | 0 | 7 | 7 | 7 | 1 | 1 | 1 | 1 |
gpt-oss:120b | gen_007_generate_release_notes.prompt | 1 | 7 | 7 | 7 | 0 | 0 | 0 | 0 |
gpt-oss:120b | gen_007_generate_release_notes.prompt | 2 | 7 | 7 | 7 | 0 | 0 | 0 | 0 |
gpt-oss:120b | gen_008_screen_job_applications.prompt | 0 | 9 | 9 | 9 | 1 | 1 | 1 | 1 |
gpt-oss:120b | gen_008_screen_job_applications.prompt | 2 | 9 | 9 | 9 | 0 | 0 | 0 | 0 |
gpt-oss:120b | gen_009_translate_with_tone_control.prompt | 0 | 7 | 7 | 7 | 0 | 0 | 0 | 0 |
gpt-oss:120b | gen_009_translate_with_tone_control.prompt | 1 | 7 | 7 | 7 | 0 | 0 | 0 | 0 |
gpt-oss:120b | gen_009_translate_with_tone_control.prompt | 2 | 7 | 7 | 7 | 1 | 1 | 1 | 1 |
gpt-oss:120b | gen_010_classify_incoming_emails.prompt | 0 | 7 | 8 | 7 | 0 | 0 | 0 | 0 |
gpt-oss:120b | gen_010_classify_incoming_emails.prompt | 1 | 7 | 8 | 8 | 0 | 0 | 0 | 0 |
gpt-oss:120b | gen_010_classify_incoming_emails.prompt | 2 | 7 | 7 | 7 | 1 | 1 | 1 | 1 |
gpt-oss:120b | gen_011_draft_incident_postmortems.prompt | 0 | 7 | 7 | 7 | 0 | 0 | 0 | 0 |
gpt-oss:120b | gen_012_review_pull_requests.prompt | 0 | 6 | 10 | 10 | 0 | 0 | 0 | 0 |
gpt-oss:120b | gen_012_review_pull_requests.prompt | 1 | 6 | 10 | 10 | 0.6 | 1 | 0.75 | 0.6 |
gpt-oss:120b | gen_012_review_pull_requests.prompt | 2 | 6 | 6 | 6 | 1 | 1 | 1 | 1 |
gpt-oss:120b | gen_013_recommend_products.prompt | 0 | 7 | 12 | 7 | 0 | 0 | 0 | 0 |
gpt-oss:120b | gen_013_recommend_products.prompt | 1 | 7 | 9 | 7 | 0.7778 | 1 | 0.875 | 1 |
gpt-oss:120b | gen_013_recommend_products.prompt | 2 | 7 | 7 | 7 | 1 | 1 | 1 | 1 |
gpt-oss:120b | gen_014_extract_invoice_fields.prompt | 0 | 8 | 11 | 8 | 0.7273 | 1 | 0.8421 | 1 |
gpt-oss:120b | gen_014_extract_invoice_fields.prompt | 1 | 8 | 8 | 8 | 1 | 1 | 1 | 1 |
gpt-oss:120b | gen_014_extract_invoice_fields.prompt | 2 | 8 | 8 | 8 | 0 | 0 | 0 | 0 |
gpt-oss:120b | gen_015_detect_spam_comments.prompt | 0 | 7 | 7 | 7 | 0 | 0 | 0 | 0 |
gpt-oss:120b | gen_015_detect_spam_comments.prompt | 1 | 7 | 11 | 7 | 0.6364 | 1 | 0.7778 | 1 |
gpt-oss:120b | gen_015_detect_spam_comments.prompt | 2 | 7 | 7 | 7 | 1 | 1 | 1 | 1 |
gpt-oss:120b | gen_016_answer_hr_policy_questions.prompt | 1 | 12 | 12 | 12 | 0 | 0 | 0 | 0 |
minimax-m2.1 | code_review.prompt | 0 | 8 | 8 | 8 | 1 | 1 | 1 | 1 |
minimax-m2.1 | code_review.prompt | 1 | 8 | 8 | 8 | 1 | 1 | 1 | 1 |
minimax-m2.1 | code_review.prompt | 2 | 8 | 8 | 8 | 1 | 1 | 1 | 1 |
minimax-m2.1 | code_review.prompt | 3 | 8 | 8 | 8 | 1 | 1 | 1 | 1 |
minimax-m2.1 | code_review.prompt | 4 | 8 | 8 | 8 | 1 | 1 | 1 | 1 |
minimax-m2.1 | content_moderation.prompt | 0 | 8 | 8 | 8 | 1 | 1 | 1 | 1 |
minimax-m2.1 | content_moderation.prompt | 1 | 8 | 8 | 8 | 1 | 1 | 1 | 1 |
minimax-m2.1 | content_moderation.prompt | 2 | 8 | 8 | 8 | 1 | 1 | 1 | 1 |
LLM Evaluation Self-Audit
Data for How Reproducible Are Evaluation Conclusions? A Self-Audit of LLM-Inferred Prompt Structure, Dipankar Sarkar, Skelf Research.
Code and paper source: github.com/sarkar-dipankar/llm-evaluation-self-audit (this package is built from commit 687ee8a).
The finding
We asked eight open model variants to infer the structure of a prompt, then asked again with the identical call. Caching was off.
They often gave a different answer.
- Mean node-set Jaccard across repeated identical calls ranges from 0.39 (gpt-oss:120b) to 0.96 (minimax-m2.1).
- Only 35 of 127 prompt-model cells were node-set-perfect on every run. So 72% never were.
- We then ranked the eight models by reproducibility and bootstrapped the ranking over prompts (10,000 replicates, seed 20260817). The bottom holds. The two least reproducible models keep their rank in 99% and 86% of replicates. The middle four keep theirs in 27% to 48%. The top two in 68% each. The table finds the worst model. It does not reliably find the best.
- Reproducible is not the same as correct. Mean F1 against ground-truth annotations ranges from 0.56 (gpt-oss:120b) to 0.99 (minimax-m2.1) over 293 persisted runs.
- On the audit date, 2026-08-17, 4 of the 8 evaluated variants returned HTTP 410 (retired). The study as specified can no longer be re-run. The raw outputs here are what survives.
Every number above is recomputed from the files in this repo: ir_stability_cells, ir_stability_summary, ir_correctness, model_availability, and raw/results/stats_analysis.json.
If a small evaluation can look this definitive on this little evidence, how many published leaderboards would survive the same audit?
Configs
Nine viewer configs, all JSONL, built by build_hf.py (in this repo, stdlib only, deterministic). Six are the release's results/*.csv files converted row for row, one typed value per column. The original CSVs are unchanged under raw/results/.
| Config | File | Rows | What it is |
|---|---|---|---|
ir_correctness (default) |
data/ir_correctness.jsonl |
293 | Each persisted inferred structure (IR) scored against gold annotations: precision, recall, F1, explicit-node precision. One row per model, prompt, run. |
ir_stability_cells |
data/ir_stability_cells.jsonl |
127 | One row per prompt-model cell: mean Jaccard, ID stability, perfect flag, and the campaign CSV the cell came from. Source: results/ir_stability_merged_cells.csv. |
ir_stability_summary |
data/ir_stability_summary.jsonl |
8 | Per-model reproducibility table (the ranked table the paper audits). Source: results/ir_stability_merged.csv. |
drift |
data/drift.jsonl |
98 | What changes when an IR drifts: id-label, count, or metadata disagreement per cell. Source: results/drift_decompose.csv. The column hierarchy_drift is what the paper calls metadata drift, the IR is a flat node list. |
downstream_cost |
data/downstream_cost.jsonl |
98 | Node-count trajectory per cell and the run at which its rolling mean stabilises. The paper documents two specification errors in this analysis, read its appendix before using stabilises_at. |
model_availability |
data/model_availability.jsonl |
15 | Live endpoint check on 2026-08-17: HTTP status and provider message for every model the study named, tried, or probed. |
ir_runs |
data/ir_runs.jsonl |
850 | Every raw IR file. 293 from results/irs (the main study) and 557 from results/irs_crossed (the crossed run on public SKILL.md prompts). The full IR is in ir as a JSON string. |
behaviour |
data/behaviour.jsonl |
80 | Companion mutation study: original and mutated rendered prompts, the mutation operator, and model responses to each. Response lists are JSON strings. |
corpus |
data/corpus.jsonl |
123 | The prompts. 83 synthetic templates and 40 SKILL.md files from public GitHub repos, 12 of them with hand annotations. Contexts are a JSON string. |
Column notes
ir_runs.modelis the on-disk directory slug (gemma3_12b), the CSVs use the provider name (gemma3:12b).runis parsed from the_r<N>filename suffix.n_nodesandconfidenceare lifted from the IR for filtering.corpus.commitis the pinned revision parsed fromhtml_url.corpus.git_blob_shais the MANIFEST column namedcommit, which holds the file's git blob sha (it matchesgit hash-object), not a commit.- Nested values are serialised as JSON strings so each column has one type. Parse them with
json.loads.
raw/
raw/results/ and raw/corpus/ are exact copies of the release directories, minus four collection *.log files. No config points at them. They hold every CSV the paper uses (including the per-campaign ir_stability_*.csv files), stats_analysis.json, the coverage and mutation outputs, and the annotation notes. Nothing from the analysis is dropped.
Reproduce
Clone the GitHub repo and run from its root:
python3 -m venv .venv
. .venv/bin/activate
pip install -r requirements.txt
(cd tool && cargo build --release -p promptdbg) # needed by ir_correctness.py
python scripts/merge_ir_stability.py
python scripts/stats_analysis.py
python scripts/drift_decompose.py
python scripts/downstream_cost.py
python scripts/ir_correctness.py
python scripts/stability_vs_correctness.py
ir_correctness.py calls the promptdbg binary in tool/target/release/, so build it first (Rust toolchain, about a minute). These scripts replay the saved outputs offline.
The live checks in model_availability.py and failure_attribution.py need provider access and cannot bring back retired endpoints. The saved responses make the analysis repeatable. They do not make the collection repeatable.
To rebuild this package from a checkout:
python3 build_hf.py /path/to/llm-evaluation-self-audit .
Licensing
- Author material (measurements, synthetic prompts, annotations, analysis outputs, this card) is released under CC BY 4.0.
- The 40 files in
corpus/github/(rows withsource: githubin thecorpusconfig) are third-party SKILL.md files. They keep their upstream licenses: 38 MIT, 1 Apache-2.0, 1 Unlicense. Each row carrieslicense,repo,path,commit, andhtml_url, fromcorpus/MANIFEST.csv. CC BY 4.0 does not replace those terms.
Citation
@misc{sarkar2026reproducible,
title = {How Reproducible Are Evaluation Conclusions? A Self-Audit of LLM-Inferred Prompt Structure},
author = {Dipankar Sarkar},
year = {2026},
eprint = {2609.30074},
archivePrefix = {arXiv},
primaryClass = {cs.CL},
doi = {10.48550/arXiv.2609.30074},
url = {https://arxiv.org/abs/2609.30074}
}
Links
- Paper: https://arxiv.org/abs/2609.30074
- Code, paper source, and full artifact: https://github.com/sarkar-dipankar/llm-evaluation-self-audit
- Downloads last month
- 293