We're sitting around 0.906 and have spent the last stretch testing ideas that sounded
obvious and turned out to be worth nothing. Posting the negative results so nobody else
pays for them twice. Everything below is measured on held-out training lines, not
guessed, and not on the test set.
HOW TO READ THE NUMBERS
All of it is measured on 820 held-out lines from a 5-fold split (train on 4 folds,
predict the 5th). Sampling noise on a set that size is sd 0.0048, so anything under
about 0.005 is indistinguishable from luck. If you are tuning on a validation set of a
few hundred lines and chasing 0.003 differences, you are chasing noise -- that alone
cost us one bad decision.
Worth knowing about the metric: the leaderboard's WER/CER columns are mean edit COUNTS
per line, not rates, and the score is
public = 0.5 * (1 - word_edits_per_line / 12) + 0.5 * (1 - char_edits_per_line / 55)
That reproduces the Benchmark row to 2e-10 and our own submissions to nine decimals.
(Credit to J0NNY, who posted this form first.) The practical consequence: one word edit
costs the same as 55/12 = 4.58 character edits. Getting a word exactly right is worth
far more than shaving a character off a wrong one.
1. RECASING FROM A CORPUS LEXICON -- doesn't work
Casing is about 13% of our word errors, and a perfect casing oracle would be worth
+0.014. Tempting. So: build a table of how each word is capitalised across the training
text, and push each predicted word to its dominant form.
It LOSES at every confidence threshold we tried (0.60 through 1.00), from -0.016 to
-0.0006. Even at 1.00 -- words that take that form in the training data without a single
exception -- it still loses.
Why: the errors are symmetric. We write lower case where the truth is upper about as
often as the reverse. That's genuine ambiguity in a 17th-century hand, not a bias you can
correct from frequency. The model can see the ink; your lexicon can't.
2. FIXING THE MARKUP (^ : &) -- not worth your time
The caret superscripts, the colon suspensions and the ampersands look like the obvious
place a general-purpose VLM would go wrong. We assumed our model was expanding & into
"and" and dropping carets.
It isn't. We emit MORE ampersands than the truth contains (ratio 1.09). Eliminating the
&/and confusion entirely is worth -0.0002 -- i.e. nothing. Perfect caret handling is
+0.0027, perfect colon handling +0.0016, both inside the noise band. All markup slips
together are 3% of our word edits.
Check your own ratio before building anything here. It's a five-line script.
3. SNAPPING RARE WORDS TO THE TRAINING VOCABULARY -- actively harmful
Take any predicted word that never appears in the training text, and replace it with the
most frequent word within edit distance 1. Given the 4.58x premium on word-exactness this
should be the highest-value trick available.
Applied to our held-out predictions it scores -0.0120. It fixes 67 words and breaks 79.
Why it can't work: about 43% of our wrong words are cases where BOTH our word and the
truth are ordinary training-vocabulary words -- an OOV filter cannot even see those. And
the rare words the rule does reach are disproportionately names and places, exactly the
tokens you must not touch. Gating it on model confidence might rescue it; ungated it is
a net loss.
4. GUARDING AGAINST OVER-TRANSCRIPTION -- not a real problem
There was a thread about multi-line crops and models transcribing neighbouring lines.
Under this metric one hallucinated extra line is genuinely expensive, so we checked.
Zero held-out lines are more than twice the length of their label. 96.5% are within 10%
of it. If your model is well fine-tuned this failure mode does not appear, and a guard
against it is dead code. (Genuinely multi-line crops are also rare: only ~0.1% of images
have an aspect ratio below 6.)
5. PER-BATCH RESOLUTION NORMALISATION -- doesn't fix what it looks like it fixes
This one is the most interesting, because the premise is TRUE and the fix still fails.
The images come in two scan batches. Split your validation by image height (>150px is a
clean cut) and you'll likely see what we see: the high-resolution batch scores far worse
-- 0.823 vs 0.900 for us. It's ~27% of lines carrying ~42% of all character edits.
The natural conclusion is that a fixed-height squash mistreats them. It looks well
supported: those crops end up with fewer output pixels per character (35 vs 47) because
they carry more vertical margin, and within the batch the error rises monotonically as
you squash harder.
It's still wrong. Cropping to the ink band lifts them to parity on pixels-per-character
and buys +0.0003. And at MATCHED pixels-per-character the gap is still ~0.05 in every
bucket -- same at matched line length. They're harder originals, not mishandled ones.
More pixels is not the answer for them.
We also lost 0.009 on the leaderboard to a band-crop preprocessing change earlier, so
that's two preprocessing experiments that cost real submissions. Test the matched
comparison before you rebuild your pipeline.
ONE CAUTION, NOT A TECHNIQUE
Roughly 1 line in 70 appears to be mislabelled -- the image paired with a different
line's transcription. On our held-out fold, 12 lines out of 820 (1.5%) carry 18% of all
our character edits, and on those the model outputs perfectly fluent, coherent text that
has nothing to do with the label. That's what a wrong label looks like, not a wrong
reading.
Two implications. Don't tune against your worst validation lines -- a chunk of them are
unfixable. And there is a floor under everyone's score from the same noise in the test
set.
The staff have said clearly that you may exclude clearly corrupted training rows if you
document which and why (thread 33891). Worth a look. But note that the 21 IDs circulated
earlier are NOT all bad -- we checked them, and most are correct centre-line labels on
tall multi-line crops. Only a handful are genuinely junk or mismatched. Don't drop all 21
blindly; you'd be removing exactly the examples that teach the model not to over-read.
WHAT WE'RE NOT SAYING
We haven't cracked this. We're mid-table and most of what's left on our list is
unproven. This post is only the negative half of our notebook -- the stuff we can say
with confidence because it's measured, and the stuff that costs you compute to discover
for yourself.
If you've tested any of the above and got a different result, please say so. Different
model families may behave differently, and 820 lines is 820 lines.
Good luck all.
Thanks for these tips. I have to say, I've been finding this challange a lot harder than I was expecting. The highest I've gotten is 0.83 or so. I've tried many different techniques, models, etc and not been able to get anywhere near the top scores. I've tried large models, small models, OCR-tuned models, generic vision models, general current SOTA models. I've tried ensembles, I've tried using multiple models in sequence for different parts... still not getting very far. Your tip about some bad data in the training corpus is where I'm going to look next to see how I get on.
focus on cleaning the images . ull get +0.9
Thanks Vaultguard. Will work on that
Thanks🫡️
thank you
thank you man