Skip to main content
All forms
Loading the form…
Sources & notes

OpenAI’s 2020 study used human comparisons of summaries with reviewer guidance and quality checks. Fillo supplies a form for one pair, adding tie and unable-to-judge options. It does not reproduce the study’s evaluation or training pipeline.

Give reviewers the source text, both summaries, and the same review criteria. One shared form contains one fixed pair.

Keep ties and unable-to-judge answers separate. A model comparison needs many cases, blinded identities, and varied A/B order; this form does not assign those. A preference is not proof of factual accuracy.