Hi, could the organizers please clarify the rule about probability thresholds?
The competition page says that setting a probability threshold is strictly forbidden and that TargetF1 should be based on the default threshold of 0.5.
Does this requirement apply to every submission made during the public leaderboard phase, including experimental submissions, or only to the final two submissions selected for private evaluation?
For example, would it be allowed to submit temporary threshold variants to study model calibration, provided that the final selected and reproducible model uses:
TargetF1 = (predicted_probability >= 0.5)
Could you also clarify whether threshold-adjusted submissions are detected automatically by the platform, checked during code review, or reviewed in some other way?
This matters because, if non-compliant threshold tuning can improve public leaderboard scores without being detected until the end, participants may end up chasing leaderboard results that cannot be reproduced under the official rules. It would be helpful to know whether the current leaderboard should be assumed to reflect only submissions using the required 0.5 threshold.
An explicit clarification would help everyone experiment fairly and avoid unintentionally violating the rules.
i agree with this question, pls. answer, organizers! i have been very hesitant to do any type of class weighting as it appears to be prohibitied by the rules - but there is no way to tell if leaderboard submissions are using it!
Hi @meganomaly / @AJoel,
Follow-up to the earlier "Threshold clarification" thread — I want to check a specific technique against the rule that raw probabilities must reflect the default 0.5 threshold, with no threshold-setting or rounding to improve leaderboard placement.
You've confirmed the test set's positive rate is inherently higher than the ~40% train rate. Given that, is it acceptable to apply a prior-probability-shift correction (Saerens, Latinne & Decaestecker, 2002) to our raw model outputs before submission?Concretely: we estimate the test-set positive rate from our own out-of-fold TPR/FPR bias (Adjusted Classify & Count), then apply the standard closed-form Bayesian correction to each row's probability, given the known train prior and the estimated test prior:corrected_p = (test_prior/train_prior * p) / (test_prior/train_prior * p + (1-test_prior)/(1-train_prior) * (1-p))
This produces a genuinely updated per-row probability estimate reflecting the known prevalence shift you've described — not a uniform quantile remap chosen to hit a target leaderboard score. TargetRAUC would be this corrected probability, and TargetF1 would be (corrected_p >= 0.5) at the standard threshold. Is this within the rules, or does it fall under the prohibited threshold/probability-adjustment behavior? A clear yes/no would help everyone avoid either under- or over-correcting for the prevalence difference you've flagged.
Thanks!