Hello everyone, I hope you're doing well. I have two quick questions about the rules for the ROAD Barbados Historic Handwriting Challenge:
- Are hyperparameter optimization libraries such as Optuna allowed? I understand the rules restrict AutoML tools — I'd like to confirm whether standard hyperparameter search (learning rate, dropout, etc.) on our own models falls under that restriction or not.
- During cross-validation we identified ~20 training rows with corrupted labels (e.g., image–label mismatches, an empty image). Is it allowed to exclude these rows from training? To be clear: no test data is touched in any way, no labels are modified, and the original Train.csv stays unchanged — we would only skip these rows when training.
Thanks a lot!
I agree we have tiny noise in train set, for example:
- 79tMUVyfIdy3GzkG is empty/not readable.
- F8DYDDp2AvW9Dytw text does not match to the label.
- JU7lRwk3jKkus24Z text does not match to the label.
I guess we're allowed to at least drop them as it's regular cleanup of train data.
Optuna is not AutoML and is just a tool to optimize parameters of a model (it is not the model itselt) - it is allowed.
excluding/cleaning train data is also allowed.
Thank you for answer.
Hi @Alpan_Cepik
Thanks for the query.
Thanks for flagging the affected rows - please also share the IDs with the Zindi team so we can review them.
Thank you for the clarification! As requested, here are the 21 training-row IDs we excluded.
79tMUVyfIdy3GzkG, rh8o7bdCGOIBFPwH, F8DYDDp2AvW9Dytw, mfQxfOmeRmwBh0g8, EcxuqKeZl7OQexfB, PO7QQLWIFOT65BTz, yNyf3Tp0zc7DFj5F, JU7lRwk3jKkus24Z, R6iYPb7MHFiHtXH6, VmrEALeZiP1Y6nF9, t0UrASljcgzvBAnO, KH5g92Q3DA5Bo6xi, N78M3v6GKmUzF7sk, 0CCrVKAom8EK53jj, 3ZxOeKcOr5wUYyk0, PJbM7Q1SrblrSWt6, WwlTCykxjP3c4kfo, baTY3OlGskirWgFc, u3b4JNo5bqpfE7Js, MfT9S5oghk9ywNSC, 8H2ITJSWZhAD6eh0
@Alpcan_Cepik
Have you found issues with all these images. Do all the images fall in empty/mislabelled categories or you have other reasons to exclude some of them.