Loading the form…
Sources & notes
OpenAI’s 2020 study used human comparisons of summaries with reviewer guidance and quality checks. Fillo supplies a form for one pair, adding tie and unable-to-judge options. It does not reproduce the study’s evaluation or training pipeline.
- Learning to summarize from human feedbackStiennon et al. / OpenAI · 2020-09-02
- Learning to summarize with human feedbackOpenAI · 2020-09-04
Give reviewers the source text, both summaries, and the same review criteria. One shared form contains one fixed pair.
Keep ties and unable-to-judge answers separate. A model comparison needs many cases, blinded identities, and varied A/B order; this form does not assign those. A preference is not proof of factual accuracy.