Hi all,
While going through the training images, I noticed that most multi-line images are labeled with only the text of the centered/main line, but a portion of the training set has labels that include the full multi-line text instead.
Before I invest time building special handling for multi-line cases, I'd like to clarify: for the test set, are multi-line images expected to follow the same mixed convention (i.e., some evaluated as center-line-only, some as full multi-line), or is there a single consistent rule that applies to test specifically?
Concretely:
Any clarification would help before I decide how much engineering effort to put into detecting/handling multi-line cases specifically. Thanks!
@FNF
I'd treat it as full transcription unless the organizers say otherwise. I haven't found a reliable way to tell when only the center line should be used, so I'd avoid adding special-case logic.
yeah I noticed the same thing, and it is unclear what to do actually
Hi @youssefilo Thanks for letting us know. Could you share a few example IDs where you've found this please. We'll then investigate and come back with a more solid response re handling. Thanks very much!