Skip to content

fix: keep letter-spaced words whole on runs with explicit spaces (#377) - #385

Draft
seanphopkins wants to merge 3 commits into
docling-project:mainfrom
seanphopkins:fix/word-gap-tracked-runs
Draft

seanphopkins wants to merge 3 commits into
docling-project:mainfrom
seanphopkins:fix/word-gap-tracked-runs

Conversation

@seanphopkins

@seanphopkins seanphopkins commented Oct 6, 2026 •

Copy link
Copy Markdown

Draft, for the review in #377.

On a run that carries explicit spaces, every gap past the word-gap threshold starts a new word, so a tracked (letter-spaced) word comes out one word per letter. This change drops such a boundary when the gap is under half the run's own explicit-space width and either

  • under half a font space (the reference capped at 1.3 spaces, because justified stretch and table gutters inflate the median), or
  • no wider than 1.5 times the font's usual letter gap on the run, judged from at least three other gaps in the same font, and no wider than 1.5 times the other letter gaps of its own segment (the glyphs between two explicit spaces).

It never joins across a font change, and it can only remove a split. The same-font bound stops a tracked span in another font from vouching for a gap. The segment bound stops a tracked span elsewhere on the line from vouching for a gap between two untracked words. Gaps are grouped by font and by segment and sorted once per run, so each judgement is a lookup.

Tests. tests/test_unit_word_gap_letter_spacing.py holds the eight control streams from #377, built in memory with tests/pdf_builder.py. Three fail on main (the tracked-text cases, one of them the documented limit, where a boundary inside uniform tracking leaves no evidence and is joined) and five pass on both (the guards).

Suite, this branch against an unpatched build of main from the same source: 353 tests compared, and six change outcome. Three are test_unit_word_gap_letter_spacing cases going from fail to pass. The other three are the groundtruth tests going from pass to fail, because the references hold the old split words: test_regression_parse::test_reference_documents_from_filenames, test_regression_threaded_parse::test_threaded_reference_documents_from_filenames and test_regression_threaded_render::test_render_reference_documents_from_filenames. Nothing else changes. (test_threaded_results_match_sequential fails on both builds on Windows, unrelated.)

Scope. This PR changes only runs that carry explicit spaces: everything it adds sits behind space_reference > 0.0. Letter-spaced runs with no space glyph (elsevier-00.pdf ARTICLE INFO, 14289803404128846560-14.pdf p14 Tinta) go through infer_word_gap() and are #358's, so the two PRs edit different code.

References. The groundtruth is in the dataset repo, so it is not part of this PR. As agreed in #377 and below, the maintainers regenerate it, once, from a build carrying both this PR and #358 (their page sets overlap on 16 pages and differ on 13), and bump HF_DATASET_REVISION. Until then the three groundtruth tests fail as expected. Reviewed against the published references, this PR changes the word and line cells of 27 pages in 16 documents and no character cells; the words, page by page, are in this comment on #377.

  • Word repairs (rejoined fragments of one word, or a word and its punctuation), 15 pages: 07f5395c8b3e7d1c_0001.pdf p1; 082b97f3d239a9c5_0006.pdf p1; 10572911635253446040-13.pdf p8; 11794545469969901016-2.pdf p2; 14289803404128846560-14.pdf p2, p3, p5, p10, p12; 17791c05056ff856_0022.pdf p1; 335fea0ba454ba15_0004.pdf p1; 4796728975040539044-1.pdf p1; 5081873815222802242-2.pdf p2; 7816024906253388048-2.pdf p1; 79db8838970047b5_0002.pdf p1.
  • Mathematics (the seven operator joins below, with these documents' repairs), 6 pages: 10749817875312063576_009.pdf p5, p6, p9; 19246029784032167-4.pdf p2, p4, p6.
  • Cannot be judged by reading (mis-encoded Arabic, a right-to-left test file, GLYPH<n> text with no Unicode mapping), 6 pages: 10976580690960943929_004.pdf p7; 11816493302388412838-4.pdf p1, p6, p7, p8; right_to_left_04.pdf p1.

Known costs, discussed in #377: seven Greek-math operator joins (Θ=, α∈, η> ×2, ρ- ×3) on two documents, and the limit above.

Refs #377

🤖 Generated with Claude Code

…ling-project#377)

On a run that carries explicit spaces, every gap past the word-gap threshold
started a new word, so a tracked (letter-spaced) word came out one word per
letter. The contraction now drops such a boundary when the gap is under half
the run's own explicit-space width and either

* under half a font space (the reference capped at 1.3 spaces, because
  justified stretch and table gutters inflate the median), or
* no wider than 1.5 times the font's usual letter gap on the run, judged
  from at least three OTHER gaps in the same font, and no wider than 1.5
  times the other letter gaps of its own segment (the glyphs between two
  explicit spaces).

Never across a font change; the rule can only remove a split. The
same-font bound stops a tracked span in another font vouching for a gap;
the segment bound stops a tracked span elsewhere on the line vouching for
a gap between two untracked words. The gaps are grouped by font and by
segment and sorted once per run, so each judgement is a lookup.

The new unit tests are the control streams from the review of docling-project#377,
including the documented limit: a boundary inside uniform tracking leaves
no evidence and is joined.

The parser groundtruth changes on the pages this repairs; it is refreshed
as a separate dataset revision.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Sean Hopkins <sean.p.hopkins@gmail.com>
@mergify

mergify Bot commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

Merge Protections

🟢 Merge protection satisfied — ready to merge.

Show 1 satisfied protection

🟢 Enforce conventional commit

Make sure that we follow https://www.conventionalcommits.org/en/v1.0.0/

  • title ~= ^(fix|feat|docs|style|refactor|perf|test|build|ci|chore|revert)(?:\(.+\))?(!)?:

@github-actions

github-actions Bot commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

✅ DCO Check Passed

Thanks @seanphopkins, all your commits are properly signed off. 🎉

@wittjeff

wittjeff commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Built 587ec48 and ran it as promised in #377. Short version: this PR and #358 fix different runs, not the same one, so I will keep #358 open and narrow it rather than close it.

Your tests. All eight in test_unit_word_gap_letter_spacing.py pass on the build. #358's test_regression_cell_contraction.py also passes on it except its own letter-spaced test: elsevier-00.pdf p1 still reads a r t i c l e i n f o and 14289803404128846560-14.pdf p14 still reads T i nt a.

Why. Neither of those lines carries a space glyph (the char cells are a r t i c l e i n f o with nothing between the words), so has_semantic_spaces is false, space_reference stays at -1.0, and this PR's branch never runs. Those lines go through infer_word_gap() and hit the flat 0.35 cap there. That is the path #358 changes. So the split is clean: this PR covers spaced runs, #358 covers unspaced ones.

Regression set, word cells on the 940 selected pages, each branch against the same main (7.22.2):

pages / files changed joins new splits joins only this branch makes of those, the 7.22.0 word
this PR 28 / 16 212 0 36 28
#358 44 / 26 287 0 81 74

149 joins are shared (mostly the Portuguese table). Your 36 exclusive joins are the floor's: Inflation, Ref., the kana, aanvullingen's neighbours, and the known Θ= / α∈ / η> / ρ-. #358's 81 are on unspaced runs: the elsevier heading, Cyrillic поздравляем, the Dutch aanvullingen, and two math-heavy papers (19246029784032167-4.pdf, 8624879949050867564-1.pdf) where 7.22.1 had split tracked symbol runs that 7.22.0 kept whole.

One exclusive join I checked against 7.22.0 rather than a dictionary: 5081873815222802242-2.pdf p2 gives 27日6, 時間36, 7.25度 here, where main has 時間, 36, 7.25. 7.22.0 had 27日6時間36分 and 7.25度, so this PR gets partway back and #358 does not; neither restores the whole phrase.

Proposal. Keep both PRs and make them disjoint:

  • This PR stays as it is: the space_reference > 0.0 gate already limits it to spaced runs.
  • I drop fix: keep letter-spaced text whole under the word-gap cap #358's spaced-run branch (return std::min(1.0, letter_spacing_cap(1.0)) on has_semantic_spaces) and keep only the unspaced path, so it no longer competes with yours on the same line. Then the two change different lines of contract() and infer_word_gap(), and whichever lands second is a trivial rebase.
  • The groundtruth refresh should be one regeneration from a build carrying both changes, not two reference sets merged: the page sets overlap on 16 pages and the outputs differ on 13 of them (each side repairs different lines there, e.g. aanvullingen from fix: keep letter-spaced text whole under the word-gap cap #358 and (op-/af-)levertijd from this PR on 7816024906253388048-2.pdf p1).

If that works for you and @PeterStaar-IBM, I will re-scope #358 that way and post its before/after against this branch there.

@seanphopkins

Copy link
Copy Markdown
Author

@wittjeff thank you for building it and running both sets. Agreed on all three points. The split is clean: everything this PR adds sits behind space_reference > 0.0, so a run with no space glyph never reaches it, and #358's two cases are rightly left to #358. I'll keep #385 as it is. One regeneration from a build carrying both changes makes sense given the 13 overlapping pages that differ; I'll leave that, and the HF_DATASET_REVISION bump, to the maintainers and put the per-page summary in this PR's description. @PeterStaar-IBM, does that arrangement work for you?

@wittjeff

wittjeff commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

A correction to my comment above, after checking the char cells with keep_glyphs=True.

Only one of #358's two lines is unspaced. elsevier-00.pdf p1 (a r t i c l e i n f o) has no space glyph and goes through infer_word_gap(), as I said. 14289803404128846560-14.pdf p14 (Tinta) does not: the cell carries a suppressed PDF space after a (and one before T), so has_semantic_spaces is true and this PR's branch does run there. It declines the join because the gate gap < 0.5 * space_reference rejects every letter gap: with no measurable inner space, space_reference is 1.0, and the Calibri letter gaps are 1.5 pt over a 2.26 pt space, 0.68.

The same applies to most of the 81 joins I listed as "#358 exclusive, on unspaced runs". I checked a sample; each of these lines has a space glyph and is a spaced run where the half-space gate blocks the repair, not an unspaced one:

line file main / this PR #358
п о з д р а в л я е м 10588776208842536523_006.pdf p8 split поздравляем
F I N A N C I A L S 2014 dln_096631dd…pdf p1 F INANCI ALS 2014 FINANCIALS 2014
Α Π Ο Φ Α Σ Η 8313873609013042575-4.pdf p1 split ΑΠΟΦΑΣΗ
文 責 : 事 務 室 c879c7ccf5b9de33_0007.pdf p1 split 文責:事務室
Dez emb ro ;, Fr egu esia . 14289803404128846560-14.pdf p2, p10 split whole

So the proposal I made, dropping #358's spaced-run branch, is not free. I built that variant (this PR plus #358 restricted to infer_word_gap()): your eight tests pass and article info is whole, but Tinta and 16 other lines on 12 pages, the ones above among them, go back to letters, and #358's own test_letter_spaced_text_stays_whole fails on Tinta.

The underlying point is that real tracking is routinely 0.6 to 0.8 of a space, so half a space is the right floor for the small branch but too tight a gate for the lettered branch. I also tried removing the gate from lettered only: everything from both PRs passes, but three pages regress against 7.22.0 (positioned maths symbols join as :,∈,∈⊂ on 19246029784032167-4.pdf, 脳血管疾患(4) absorbs its count, Dutch hoeveelheden), loses its comma). So the gate wants widening for lettered, not removal, and I have not tuned the bound; it should be set against your 1,050 pages.

Until then the plain merge of the two branches is the safer reference point: it is the exact union on the 940 regression pages (never more split than either PR, 56 pages against main), with the one known cost that #358's line-wide cap rejoins ONE / TWO on your two tracked controls. I will hold #358 as it is rather than re-scope it, and post its before/after against this branch once the lettered bound is settled.

Also: HYRIMOZ p89 and p96 give efficacy on this build (7.22.1 ef f icacy), matching the corrected issue body.

@seanphopkins

Copy link
Copy Markdown
Author

@wittjeff thank you for checking the cells; the table makes it concrete.

Agreed on the diagnosis. The gap < 0.5 * space_reference gate sits in front of both branches, so tracking at 0.6–0.8 of a space never reaches letter_spaced_like(). small should keep the half-space floor. lettered already carries its own evidence (at least three other same-font gaps, and the veto against a gap wider than its own segment's letter gaps), so it can take a wider gate.

I'll tune the lettered bound here. I'll sweep it at 0.6 / 0.7 / 0.8 / 0.9 / 1.0 of space_reference on the 1,050 pages, plus your 81 joins and the three pages that regress without the gate (:,∈,∈⊂, 脳血管疾患(4), hoeveelheden),). For each bound I'll report the repairs gained, the regressions against 7.22.0, and the two tracked controls. If one bound clears all of that, I'll push it to this PR as a separate commit, so #358's before/after has a fixed target.

Agreed that the plain merge of both branches is the reference until then.

On your corpus point in #377: agreed. The movements are small enough that I have kept every claim per page rather than as a rate. A larger and more varied corpus would make this bound a measurement rather than a judgement call.

…-project#377)

The letter-spacing branch shared the half-space gate of the small-gap
branch, so tracking at 0.6-0.8 of a space never reached it (found on docling-project#385:
Portuguese Calibri text tracked at ~0.7 of a space stayed one word per
letter). The small-gap branch keeps its half-space floor; the
letter-spacing branch, which carries its own same-font and same-segment
evidence, now admits gaps under 0.8 of the run's space.

Swept 0.5 to 1.0 on the 940 regression pages: 0.8 makes those words whole
with no join beyond what this PR or docling-project#358 already make, apart from one
split dot leader joined and a partial join of a letter-split Japanese
word; at 1.0 positioned maths symbols join (:,).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Sean Hopkins <sean.p.hopkins@gmail.com>
@seanphopkins

Copy link
Copy Markdown
Author

@wittjeff the sweep is done, and I have pushed the result as a separate commit, e418bb4.

What changed. small keeps its half-space floor. lettered now admits gaps under 0.8 of the run's space (it was gated at 0.5 with small). There are two new control tests: a word tracked at 0.7 of a measured space stays whole, and the same 0.7-space gap between two untracked words stays split. The first test fails on main and on 587ec48 and passes on e418bb4; the second passes on all three. All 10 tests in the file pass.

A correction first. The 1,050-page set is @rasokiwayami's, not mine. I measured on the 940 regression groundtruth pages plus 483 pages of my own corpus, 1,423 pages in all. Tinta (p14) is not among the 940 groundtruth pages, so I could not check it.

The sweep, against main (7.22.2):

lettered gate #358-only joins also made joins neither PR makes
0.5 (587ec48) 0 / 580 none
0.6 1 / 580 one split dot leader joined
0.7 1 / 580 partial Portuguese fragments (rada, oriz, Incumpr)
0.8 1 / 580 the dot leader; ー病 (a partial join of a letter-split アルツハイマー病)
0.9 1 / 580 same as 0.8
1.0 2 / 580 adds :, on 19246029784032167-4.pdf p2, p4 and p6

"#358-only" means the 580 words #358 joins on the 28 pages where it changes words and 587ec48 does not. At 0.8, Dezembro; and Freguesia. come out whole. 1.0 is where your maths regression starts, so 0.8 is the widest clean bound. :,∈,∈⊂ and 脳血管疾患(4) do not appear at any of these bounds. hoeveelheden), on 7816024906253388048-2.pdf p1 is a different case: it already loses its comma on main and on 587ec48, and the gate does not change that at any bound. 7.22.0 keeps the comma, and #358 and the plain merge restore it, so that one is #358's to fix, not a cost of the gate.

Why widening gets so few of #358's joins. I traced the C++ judgement on the lines from your table. Your reading of the Portuguese document is right: it is tracked at 0.6 to 0.7 of a space, with 12 explicit spaces measuring the run's space, and 0.8 takes it. The others are spaced runs too, as you said, but each has a single space glyph at its edge, so there is no inner space to measure, and space_reference falls back to 1.0 font space. Their letter gaps are wider than that space:

line file letter gap letter_spaced_like
п о з д р а в л я е м 10588776208842536523_006.pdf p8 2.05 spaces yes
文 責 : 事 務 室 c879c7ccf5b9de33_0007.pdf p1 1.56 spaces yes
Α Π Ο Φ Α Σ Η 8313873609013042575-4.pdf p1 1.01 to 1.07 spaces yes
a r t i c l e i n f o elsevier-00.pdf p1 (unspaced run) (branch not reached)

The same-font and same-segment evidence accepts all three, so only the gate stops them. To reach them the gate would have to go past 1.0, and that is where the maths joins start. Display tracking of one to two spaces, on a run with nothing inside it to measure, is beyond what this rule can separate from a real boundary. #358's line-relative threshold handles exactly that case. So I agree that #358's spaced-run branch is still needed, and that the plain merge of the two branches is the reference.

#358's before/after can target e418bb4. If you have Tinta p14 to hand, I would be glad to know how it comes out there.

rpcndr.h defines 'small' as 'char', so 'const bool small' fails to
compile on win_arm64 (C2632). No behaviour change.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Sean Hopkins <sean.p.hopkins@gmail.com>
@seanphopkins

Copy link
Copy Markdown
Author

I pushed 9f94c75 to fix the win_arm64 build failures in the last run. small is a macro in Windows' rpcndr.h, so the local is now narrow; no behaviour change. I ran the full CI on my fork with the same commit: all 25 wheels (including the 5 win_arm64), code-checks and rhel-build pass. In regression-tests, 91 pass and the only 3 failures are the expected groundtruth tests, which pass once the references are regenerated. Could a maintainer approve the workflow run when convenient?

@morten-lagabote

Copy link
Copy Markdown

A data point from a real document with letter-spaced emphasis, which is common in Norwegian parliamentary and legal typesetting: Innst. 240 S (2013–2014), 5 pages. The word Presidentskapet is tracked by 1.65 pt per letter, with no space characters inside the word.

The pdfium text layer has 70 tracked runs (4 or more single letters separated by spaces). I matched their letters against word_cells:

Build Whole words
7.20.0, 7.22.0 60 of 70
7.22.1, 7.22.2 2 of 70
this PR (9f94c75, the macOS arm64 wheel from CI) 35 of 70

The text is justified, so the space width changes from line to line and the tracking does not. Page 1:

Line Letter gap Space gap on the line Ratio 7.22.0 7.22.2 This PR
Presidentskapet viser til representantforsla- 1.66 pt 2.41 pt 0.69 whole Pr e s i den t skape t Pr e s i den t skape t
Presidentskapet viser til at representantfor- 1.65 pt 2.89 pt 0.57 whole split split
Presidentskapet viser til at spørsmålet om 1.65 pt 3.97 pt 0.42 whole split whole
Presidentskapet arrangerte i forbindelse 1.65 pt 7.68 pt 0.21 whole split whole

So this PR joins the word when the gap is under half of the space of the line, but not on the tight lines (0.57 and 0.69). I think the cause is in letter_spaced_like(): the tracked word is one word among untracked words in the same font, so most of the other gaps of the run are near zero and typical is below 0.3.

@wittjeff

wittjeff commented Oct 9, 2026

Copy link
Copy Markdown
Contributor

I am increasingly of the opinion that this problem will only be fully solved with an AI model trained for the purpose, not an increasingly subtle pile of heuristics. Model idea to follow in Discussions.

@wittjeff

wittjeff commented Oct 9, 2026

Copy link
Copy Markdown
Contributor

For reference: I opened a longer-term proposal in docling Discussions, docling-project/docling#4731. It proposes a small learned seam tagger that uses geometry and the characters on each side. A few fixed rules handle the structural cases, and a written convention covers dot leaders and inline math.

@morten-lagabote's Presidentskapet line is the main example. Its letter gap is 0.69 of the line's space, so no threshold on gap ÷ space can keep it whole without joining real words on other lines.

This does not change anything for this PR. #385 is a clear improvement over 7.22.2, and I think it should continue as it is.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants