Primary competition visual

R.O.A.D. Barbados Historic Handwriting Challenge

Helping Barbados
$25 000 USD
Closing soon! (13 days left)
Optical Character Recognition
Natural Language Processing
1740 joined
574 active
Starti
03 Jul 26
Closei
04 Oct 26
Reveali
04 Oct 26
Pretrained model licensing: does training-data provenance matter, or only the weights licence?
8 Sep 2026, 11:48 · 0

Following the ruling on stanford-oval/churro-3B (denied because the base model carried a non-commercial Qwen Research License), I would like to ask for a clarification that applies to several candidate models.

Some widely used open-weight OCR/HTR models are released under a permissive licence but were fine-tuned by their own authors on datasets whose terms restrict commercial use. The concrete example is Microsoft's TrOCR.

My questions:

1. For this competition, is the relevant licence the one attached to the model weights, or must the licence of the data the authors used to produce those weights also permit commercial use? 2. Specifically, is microsoft/trocr-handwritten permitted? 3. If not, are the stage1 checkpoints permitted, given that they predate the IAM fine-tuning?

Discussion 0 answers