How many repeated LLM queries are enough? Testing a pilot-based reliability protocol [R]
Heat trend
Collecting trend data
The percentage is based on available heat signal, not comment count or independent people.
The author of a new preprint and founder of Rankfor.AI tested a pilot-based reliability protocol for repeated LLM queries, specifically auditing brand recommendations. The study involved three independently collected corpora covering political-orientation questionnaires and benchmark stability. Out of 39 prediction cells, 37 met the prespecified replication criterion, with two partial matches, indicating strong reliability. External validation materials are available on GitHub.
I’m the author of a new preprint on repeated-query auditing of LLM brand recommendations, and the founder of Rankfor.AI.
The practical question: how many times should we repeat a prompt before comparing results?
The paper applies generalizability theory: estimate variance components from a pilot, then calculate the repeat count needed for a chosen reliability target.
Tested the reliability predictions on three independently collected corpora covering political-orientation questionnaires and benchmark stability. Across 39 prediction cells, 37 met the prespecified replication criterion and two were partial matches.
The fixed iteration thresholds did not transfer. Other preregistered tests, including parts of the drift diagnostics, also failed. Those results are reported in the paper.
An important limitation i see is that these external corpora do not contain brand recommendations. They test the statistical machinery outside our original application which is independent replication on repeated brand-recommendation data remains outstanding.
I’d particularly welcome criticism of the pilot-based variance estimates and the reliability validation design. Does anyone know an independently collected brand-recommendation dataset with repeated identical prompts?
Preprint: https://arxiv.org/abs/2609.04047
External validation materials: https://github.com/Rankfor/rankfor-open/tree/main/research/dice-roll-method/external-validation