Experiment analysisConfiguration & exports

Condition performance

Models are analyzed separately. Rankings use regularized Bradley–Terry strengths, normalized to sum to 1; raw wins are descriptive only.

Ties contribute half a win to each condition in the fit. A symmetric 0.5 pseudo-win per observed edge keeps estimates finite. Invalid judgments are excluded. Pairwise win rates below are wins / (wins + losses + ties).

Combined evaluator results

All evaluator response files for this experiment are combined below. 0 selections collected.