Current events and new studies, read on a calibrated ruler.
The same model, the same benchmark, a different score every run. That instability is not noise in your results. It is your instrument telling you it was never measuring.