Repository navigation
fix: keep letter-spaced words whole on runs with explicit spaces (#377) - #385
seanphopkins wants to merge 3 commits into
Conversation
…ling-project#377) On a run that carries explicit spaces, every gap past the word-gap threshold started a new word, so a tracked (letter-spaced) word came out one word per letter. The contraction now drops such a boundary when the gap is under half the run's own explicit-space width and either * under half a font space (the reference capped at 1.3 spaces, because justified stretch and table gutters inflate the median), or * no wider than 1.5 times the font's usual letter gap on the run, judged from at least three OTHER gaps in the same font, and no wider than 1.5 times the other letter gaps of its own segment (the glyphs between two explicit spaces). Never across a font change; the rule can only remove a split. The same-font bound stops a tracked span in another font vouching for a gap; the segment bound stops a tracked span elsewhere on the line vouching for a gap between two untracked words. The gaps are grouped by font and by segment and sorted once per run, so each judgement is a lookup. The new unit tests are the control streams from the review of docling-project#377, including the documented limit: a boundary inside uniform tracking leaves no evidence and is joined. The parser groundtruth changes on the pages this repairs; it is refreshed as a separate dataset revision. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Sean Hopkins <sean.p.hopkins@gmail.com>
Merge Protections🟢 Merge protection satisfied — ready to merge. Show 1 satisfied protection🟢 Enforce conventional commitMake sure that we follow https://www.conventionalcommits.org/en/v1.0.0/
|
|
✅ DCO Check Passed Thanks @seanphopkins, all your commits are properly signed off. 🎉 |
|
Built Your tests. All eight in Why. Neither of those lines carries a space glyph (the char cells are Regression set, word cells on the 940 selected pages, each branch against the same
149 joins are shared (mostly the Portuguese table). Your 36 exclusive joins are the floor's: One exclusive join I checked against 7.22.0 rather than a dictionary: Proposal. Keep both PRs and make them disjoint:
If that works for you and @PeterStaar-IBM, I will re-scope #358 that way and post its before/after against this branch there. |
|
@wittjeff thank you for building it and running both sets. Agreed on all three points. The split is clean: everything this PR adds sits behind |
|
A correction to my comment above, after checking the char cells with Only one of #358's two lines is unspaced. The same applies to most of the 81 joins I listed as "#358 exclusive, on unspaced runs". I checked a sample; each of these lines has a space glyph and is a spaced run where the half-space gate blocks the repair, not an unspaced one:
So the proposal I made, dropping #358's spaced-run branch, is not free. I built that variant (this PR plus #358 restricted to The underlying point is that real tracking is routinely 0.6 to 0.8 of a space, so half a space is the right floor for the Until then the plain merge of the two branches is the safer reference point: it is the exact union on the 940 regression pages (never more split than either PR, 56 pages against Also: HYRIMOZ p89 and p96 give |
|
@wittjeff thank you for checking the cells; the table makes it concrete. Agreed on the diagnosis. The I'll tune the Agreed that the plain merge of both branches is the reference until then. On your corpus point in #377: agreed. The movements are small enough that I have kept every claim per page rather than as a rate. A larger and more varied corpus would make this bound a measurement rather than a judgement call. |
…-project#377) The letter-spacing branch shared the half-space gate of the small-gap branch, so tracking at 0.6-0.8 of a space never reached it (found on docling-project#385: Portuguese Calibri text tracked at ~0.7 of a space stayed one word per letter). The small-gap branch keeps its half-space floor; the letter-spacing branch, which carries its own same-font and same-segment evidence, now admits gaps under 0.8 of the run's space. Swept 0.5 to 1.0 on the 940 regression pages: 0.8 makes those words whole with no join beyond what this PR or docling-project#358 already make, apart from one split dot leader joined and a partial join of a letter-split Japanese word; at 1.0 positioned maths symbols join (:,). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Sean Hopkins <sean.p.hopkins@gmail.com>
|
@wittjeff the sweep is done, and I have pushed the result as a separate commit, e418bb4. What changed. A correction first. The 1,050-page set is @rasokiwayami's, not mine. I measured on the 940 regression groundtruth pages plus 483 pages of my own corpus, 1,423 pages in all. The sweep, against
"#358-only" means the 580 words #358 joins on the 28 pages where it changes words and 587ec48 does not. At 0.8, Why widening gets so few of #358's joins. I traced the C++ judgement on the lines from your table. Your reading of the Portuguese document is right: it is tracked at 0.6 to 0.7 of a space, with 12 explicit spaces measuring the run's space, and 0.8 takes it. The others are spaced runs too, as you said, but each has a single space glyph at its edge, so there is no inner space to measure, and
The same-font and same-segment evidence accepts all three, so only the gate stops them. To reach them the gate would have to go past 1.0, and that is where the maths joins start. Display tracking of one to two spaces, on a run with nothing inside it to measure, is beyond what this rule can separate from a real boundary. #358's line-relative threshold handles exactly that case. So I agree that #358's spaced-run branch is still needed, and that the plain merge of the two branches is the reference. #358's before/after can target e418bb4. If you have |
rpcndr.h defines 'small' as 'char', so 'const bool small' fails to compile on win_arm64 (C2632). No behaviour change. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Sean Hopkins <sean.p.hopkins@gmail.com>
|
I pushed 9f94c75 to fix the |
|
A data point from a real document with letter-spaced emphasis, which is common in Norwegian parliamentary and legal typesetting: Innst. 240 S (2013–2014), 5 pages. The word The pdfium text layer has 70 tracked runs (4 or more single letters separated by spaces). I matched their letters against
The text is justified, so the space width changes from line to line and the tracking does not. Page 1:
So this PR joins the word when the gap is under half of the space of the line, but not on the tight lines (0.57 and 0.69). I think the cause is in |
|
I am increasingly of the opinion that this problem will only be fully solved with an AI model trained for the purpose, not an increasingly subtle pile of heuristics. Model idea to follow in Discussions. |
|
For reference: I opened a longer-term proposal in docling Discussions, docling-project/docling#4731. It proposes a small learned seam tagger that uses geometry and the characters on each side. A few fixed rules handle the structural cases, and a written convention covers dot leaders and inline math. @morten-lagabote's This does not change anything for this PR. #385 is a clear improvement over 7.22.2, and I think it should continue as it is. |
Draft, for the review in #377.
On a run that carries explicit spaces, every gap past the word-gap threshold starts a new word, so a tracked (letter-spaced) word comes out one word per letter. This change drops such a boundary when the gap is under half the run's own explicit-space width and either
It never joins across a font change, and it can only remove a split. The same-font bound stops a tracked span in another font from vouching for a gap. The segment bound stops a tracked span elsewhere on the line from vouching for a gap between two untracked words. Gaps are grouped by font and by segment and sorted once per run, so each judgement is a lookup.
Tests.
tests/test_unit_word_gap_letter_spacing.pyholds the eight control streams from #377, built in memory withtests/pdf_builder.py. Three fail onmain(the tracked-text cases, one of them the documented limit, where a boundary inside uniform tracking leaves no evidence and is joined) and five pass on both (the guards).Suite, this branch against an unpatched build of
mainfrom the same source: 353 tests compared, and six change outcome. Three aretest_unit_word_gap_letter_spacingcases going from fail to pass. The other three are the groundtruth tests going from pass to fail, because the references hold the old split words:test_regression_parse::test_reference_documents_from_filenames,test_regression_threaded_parse::test_threaded_reference_documents_from_filenamesandtest_regression_threaded_render::test_render_reference_documents_from_filenames. Nothing else changes. (test_threaded_results_match_sequentialfails on both builds on Windows, unrelated.)Scope. This PR changes only runs that carry explicit spaces: everything it adds sits behind
space_reference > 0.0. Letter-spaced runs with no space glyph (elsevier-00.pdfARTICLE INFO,14289803404128846560-14.pdfp14Tinta) go throughinfer_word_gap()and are #358's, so the two PRs edit different code.References. The groundtruth is in the dataset repo, so it is not part of this PR. As agreed in #377 and below, the maintainers regenerate it, once, from a build carrying both this PR and #358 (their page sets overlap on 16 pages and differ on 13), and bump
HF_DATASET_REVISION. Until then the three groundtruth tests fail as expected. Reviewed against the published references, this PR changes the word and line cells of 27 pages in 16 documents and no character cells; the words, page by page, are in this comment on #377.07f5395c8b3e7d1c_0001.pdfp1;082b97f3d239a9c5_0006.pdfp1;10572911635253446040-13.pdfp8;11794545469969901016-2.pdfp2;14289803404128846560-14.pdfp2, p3, p5, p10, p12;17791c05056ff856_0022.pdfp1;335fea0ba454ba15_0004.pdfp1;4796728975040539044-1.pdfp1;5081873815222802242-2.pdfp2;7816024906253388048-2.pdfp1;79db8838970047b5_0002.pdfp1.10749817875312063576_009.pdfp5, p6, p9;19246029784032167-4.pdfp2, p4, p6.GLYPH<n>text with no Unicode mapping), 6 pages:10976580690960943929_004.pdfp7;11816493302388412838-4.pdfp1, p6, p7, p8;right_to_left_04.pdfp1.Known costs, discussed in #377: seven Greek-math operator joins (
Θ=,α∈,η>×2,ρ-×3) on two documents, and the limit above.Refs #377
🤖 Generated with Claude Code