fix(docx): import DOCX blocks in document order - #332
Merged
petertzy merged 2 commits intoOct 5, 2026
Merged
Conversation
`_convert_docx_to_markdown` read `document.paragraphs` to exhaustion and then `document.tables`. Those are two independent lists, so every table was emitted after every paragraph no matter where it sat in the document. A report that opens with a paragraph, then a summary table, then a conclusion imported as paragraph / conclusion / table. The table still renders, and the delimiter row still works, so nothing looks broken — the document just no longer says what it said, and the reordering is worst for exactly the reports that read best with the table inline. python-docx 1.1.0 added `Document.iter_inner_content`, which yields the same top-level paragraphs and tables interleaved in document order. It is the declared floor of the `python-docx>=1.1.0` requirement. Its selection matches the two properties being replaced: paragraphs inside revision marks such as `w:ins` and tables nested in a cell are excluded by all three, so the set of exported content is unchanged and only its order moves. The row loop moves out into `_docx_table_to_markdown` so the interleaving loop stays readable and the row logic can be tested on its own. Adds 15 tests: prose after a table, prose before one, two tables around a paragraph, consecutive tables keeping their own delimiter rows, headings before a table, a table at either end of the body, an exact expected string for a five block document, list items around a table, an empty paragraph not shifting a table, a nested table still not emitted, a revision-marked paragraph still skipped, ordering surviving the renderer, and both helper-level cases. 9 of the 15 fail against the previous implementation; 9 of 9 mutants killed.
Repository owner
locked and limited conversation to collaborators
Oct 5, 2026
Repository owner
unlocked this conversation
Oct 5, 2026
petertzy
self-requested a review
October 5, 2026 13:41
Owner
|
Approved. This is a necessary and well-scoped fix for DOCX imports that previously moved all tables to the end of the generated Markdown. Using Document.iter_inner_content() correctly preserves the order of top-level paragraphs and tables while remaining compatible with the project's declared python-docx>=1.1.0 requirement. The refactoring of table serialization into a helper keeps the conversion logic clear, and the added regression coverage thoroughly validates interleaving, headings, lists, consecutive tables, nested tables, revision-marked content, rendering, and delimiter placement. The targeted tests, full test suite, Ruff checks, and formatting checks all pass. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Importing a DOCX moved every table to the end of the document, regardless of where it actually sat.
Root cause
document.paragraphsanddocument.tablesare two independent lists. The converter drained one and then the other:So the output order was always all prose, then all tables.
Effect
A document whose body is Intro → summary table → Conclusion imports as:
The table still renders and the GFM delimiter row is still correct, so nothing looks broken — the import just no longer says what the document said. The reordering bites hardest on exactly the reports that read best with the table inline: reports, comparisons, changelogs, meeting notes with a metrics table.
It is invisible in the simplest case, which is why it survived: a document that ends with a table already looks right, because "tables last" is what the old code did anyway.
Fix
python-docx 1.1.0 added
Document.iter_inner_content(), which yields the same top-level paragraphs and tables interleaved in document order. It is the declared floor of the repo'spython-docx>=1.1.0requirement, so no dependency bump is needed.The set of exported content is unchanged — only its order moves. I checked the selection rules against the two properties being replaced, because they are not identical by accident:
.paragraphs.tables.iter_inner_content()w:ins/w:delAll three agree, so nothing new starts being exported and nothing stops. There is a test pinning the nested-table and revision-mark rows specifically, since those are the two rules that would silently widen the export if they ever diverged.
The row loop moves out into
_docx_table_to_markdownunchanged — including therow_index == 0check that decides where the delimiter row goes — so the interleaving loop stays readable and the row logic is testable on its own.Testing
15 new tests in
TestDocxImportKeepsDocumentOrder.Order is asserted by index comparison rather than by substring presence, because presence checks are exactly what fails to notice a reordering:
-markers and their positionsw:ins) paragraph is still skipped and does not let the table leak into that gaprender_markdown: the prose,<table>and</table>come out in source order_docx_table_to_markdownreturns the rows with no trailing blank, and is empty for a row-less table9 of the 15 fail against the previous implementation — the other 6 are the cases the old code got right by accident (a document ending in a table, a document with no prose at all), which are worth pinning so they stay right.
9 of 9 mutants killed: reverting to paragraphs-then-tables, inverting to tables-then-paragraphs, emitting only each table's header row, dropping the blank line after a table, emitting the delimiter row under every row, emitting it under the second row, exporting only the first body block, dropping the final
.strip(), and no longer skipping empty paragraphs.Verification
pytest tests/test_docx_import.py— 25 passed (10 onmain; this PR adds 15)pytest tests— unchanged failures, +15 passingruff check .— cleanruff format --check— cleaniter_inner_contentwas confirmed present at thev1.1.0tag, not just in the 1.2.0 the lockfile pins, so the declared floor is honest.