A blinded comparison of AI-generated outputs. Same task, two different ways of asking. You decide which response works better.
I'm submitting an essay to the John Locke Essay Competition (UK, judged by Oxford academics) on whether the way we phrase requests to AI changes the quality of what we get back. Your ratings are the empirical evidence in the essay.
The data is anonymous. Results appear in aggregate only ("X of N evaluators preferred…").
Your ratings have been recorded. They go directly into the empirical section of the essay.