Primary competition visual

R.O.A.D. Barbados Historic Handwriting Challenge

Helping Barbados
$25 000 USD
~2 months left
Optical Character Recognition
Natural Language Processing
1221 joined
273 active
Starti
03 Jul 26
Closei
04 Oct 26
Reveali
04 Oct 26
Rules clarification: hyperparameter search & training-data cleaning
14 Jul 2026, 09:10 · 6

Hello everyone, I hope you're doing well. I have two quick questions about the rules for the ROAD Barbados Historic Handwriting Challenge:

  1. Are hyperparameter optimization libraries such as Optuna allowed? I understand the rules restrict AutoML tools — I'd like to confirm whether standard hyperparameter search (learning rate, dropout, etc.) on our own models falls under that restriction or not.
  2. During cross-validation we identified ~20 training rows with corrupted labels (e.g., image–label mismatches, an empty image). Is it allowed to exclude these rows from training? To be clear: no test data is touched in any way, no labels are modified, and the original Train.csv stays unchanged — we would only skip these rows when training.

Thanks a lot!

Discussion 6 answers

I agree we have tiny noise in train set, for example:

- 79tMUVyfIdy3GzkG is empty/not readable.

- F8DYDDp2AvW9Dytw text does not match to the label.

- JU7lRwk3jKkus24Z text does not match to the label.

I guess we're allowed to at least drop them as it's regular cleanup of train data.

14 Jul 2026, 16:05
Upvotes 1
User avatar
nymfree

Optuna is not AutoML and is just a tool to optimize parameters of a model (it is not the model itselt) - it is allowed.

excluding/cleaning train data is also allowed.

16 Jul 2026, 07:10
Upvotes 1

Thank you for answer.

User avatar
meganomaly
Zindi

Hi @Alpan_Cepik

Thanks for the query.

  1. Optuna is allowed.
  2. Yes, you may exclude clearly corrupted training rows from your training process. Please keep the original training data unchanged and document which rows were excluded and why in your final solution. This should be based only on issues identified in the provided training data, with no use of test data or external labels.

Thanks for flagging the affected rows - please also share the IDs with the Zindi team so we can review them.

16 Jul 2026, 12:03
Upvotes 0

Thank you for the clarification! As requested, here are the 21 training-row IDs we excluded.

79tMUVyfIdy3GzkG, rh8o7bdCGOIBFPwH, F8DYDDp2AvW9Dytw, mfQxfOmeRmwBh0g8, EcxuqKeZl7OQexfB, PO7QQLWIFOT65BTz, yNyf3Tp0zc7DFj5F, JU7lRwk3jKkus24Z, R6iYPb7MHFiHtXH6, VmrEALeZiP1Y6nF9, t0UrASljcgzvBAnO, KH5g92Q3DA5Bo6xi, N78M3v6GKmUzF7sk, 0CCrVKAom8EK53jj, 3ZxOeKcOr5wUYyk0, PJbM7Q1SrblrSWt6, WwlTCykxjP3c4kfo, baTY3OlGskirWgFc, u3b4JNo5bqpfE7Js, MfT9S5oghk9ywNSC, 8H2ITJSWZhAD6eh0

User avatar
Nayal_17

@Alpcan_Cepik

Have you found issues with all these images. Do all the images fall in empty/mislabelled categories or you have other reasons to exclude some of them.