Skip to content

fix(docx): import DOCX blocks in document order - #332

Merged
petertzy merged 2 commits into
petertzy:mainfrom
harsh-thakkar7:fix/docx-import-document-order
Oct 5, 2026
Merged

petertzy merged 2 commits into
petertzy:mainfrom
harsh-thakkar7:fix/docx-import-document-order

Conversation

@harsh-thakkar7

Copy link
Copy Markdown
Contributor

Summary

Importing a DOCX moved every table to the end of the document, regardless of where it actually sat.

Root cause

document.paragraphs and document.tables are two independent lists. The converter drained one and then the other:

for paragraph in document.paragraphs:      # every paragraph…
    ...
for table in document.tables:              # …then every table

So the output order was always all prose, then all tables.

Effect

A document whose body is Intro → summary table → Conclusion imports as:

source imported
1 Introduction paragraph. Introduction paragraph.
2 summary table Conclusion paragraph that comes AFTER the table.
3 Conclusion paragraph that comes AFTER the table. | Metric | Value |
4 | Latency | 42ms |

The table still renders and the GFM delimiter row is still correct, so nothing looks broken — the import just no longer says what the document said. The reordering bites hardest on exactly the reports that read best with the table inline: reports, comparisons, changelogs, meeting notes with a metrics table.

It is invisible in the simplest case, which is why it survived: a document that ends with a table already looks right, because "tables last" is what the old code did anyway.

Fix

python-docx 1.1.0 added Document.iter_inner_content(), which yields the same top-level paragraphs and tables interleaved in document order. It is the declared floor of the repo's python-docx>=1.1.0 requirement, so no dependency bump is needed.

for block in document.iter_inner_content():
    if isinstance(block, Table):
        lines.extend(_docx_table_to_markdown(block))
        lines.append("")
        continue

    line = _docx_paragraph_to_markdown(block)
    if line:
        lines.append(line)
        lines.append("")

The set of exported content is unchanged — only its order moves. I checked the selection rules against the two properties being replaced, because they are not identical by accident:

.paragraphs .tables .iter_inner_content()
top-level blocks yes yes yes
table nested in a cell — excluded excluded
paragraph inside w:ins / w:del excluded — excluded

All three agree, so nothing new starts being exported and nothing stops. There is a test pinning the nested-table and revision-mark rows specifically, since those are the two rules that would silently widen the export if they ever diverged.

The row loop moves out into _docx_table_to_markdown unchanged — including the row_index == 0 check that decides where the delimiter row goes — so the interleaving loop stays readable and the row logic is testable on its own.

Testing

15 new tests in TestDocxImportKeepsDocumentOrder.

Order is asserted by index comparison rather than by substring presence, because presence checks are exactly what fails to notice a reordering:

  • prose after a table, prose before one, a heading before one
  • two tables separated by a paragraph are not swapped
  • consecutive tables each keep their own delimiter row
  • a table at the very start and at the very end of the body
  • an exact expected string for a five-block document (paragraph → table → paragraph → heading → paragraph), so any drift in block sequence, blank-line separation or delimiter placement fails
  • list items around a table keep both their - markers and their positions
  • an empty paragraph between blocks does not shift the table
  • a table nested in a cell is still not emitted
  • a revision-tracked (w:ins) paragraph is still skipped and does not let the table leak into that gap
  • order survives render_markdown: the prose, <table> and </table> come out in source order
  • _docx_table_to_markdown returns the rows with no trailing blank, and is empty for a row-less table

9 of the 15 fail against the previous implementation — the other 6 are the cases the old code got right by accident (a document ending in a table, a document with no prose at all), which are worth pinning so they stay right.

9 of 9 mutants killed: reverting to paragraphs-then-tables, inverting to tables-then-paragraphs, emitting only each table's header row, dropping the blank line after a table, emitting the delimiter row under every row, emitting it under the second row, exporting only the first body block, dropping the final .strip(), and no longer skipping empty paragraphs.

Verification

  • pytest tests/test_docx_import.py — 25 passed (10 on main; this PR adds 15)
  • pytest tests — unchanged failures, +15 passing
  • ruff check . — clean
  • ruff format --check — clean

iter_inner_content was confirmed present at the v1.1.0 tag, not just in the 1.2.0 the lockfile pins, so the declared floor is honest.

`_convert_docx_to_markdown` read `document.paragraphs` to exhaustion and then
`document.tables`. Those are two independent lists, so every table was emitted
after every paragraph no matter where it sat in the document.

A report that opens with a paragraph, then a summary table, then a
conclusion imported as paragraph / conclusion / table. The table still renders,
and the delimiter row still works, so nothing looks broken — the document just
no longer says what it said, and the reordering is worst for exactly the
reports that read best with the table inline.

python-docx 1.1.0 added `Document.iter_inner_content`, which yields the same
top-level paragraphs and tables interleaved in document order. It is the
declared floor of the `python-docx>=1.1.0` requirement. Its selection matches
the two properties being replaced: paragraphs inside revision marks such as
`w:ins` and tables nested in a cell are excluded by all three, so the set of
exported content is unchanged and only its order moves.

The row loop moves out into `_docx_table_to_markdown` so the interleaving loop
stays readable and the row logic can be tested on its own.

Adds 15 tests: prose after a table, prose before one, two tables around a
paragraph, consecutive tables keeping their own delimiter rows, headings before
a table, a table at either end of the body, an exact expected string for a five
block document, list items around a table, an empty paragraph not shifting a
table, a nested table still not emitted, a revision-marked paragraph still
skipped, ordering surviving the renderer, and both helper-level cases. 9 of the
15 fail against the previous implementation; 9 of 9 mutants killed.
@petertzy petertzy closed this Oct 5, 2026
Repository owner locked and limited conversation to collaborators Oct 5, 2026
Repository owner unlocked this conversation Oct 5, 2026
@petertzy petertzy reopened this Oct 5, 2026
@petertzy
petertzy self-requested a review October 5, 2026 13:41
@petertzy

petertzy commented Oct 5, 2026

Copy link
Copy Markdown
Owner

Approved. This is a necessary and well-scoped fix for DOCX imports that previously moved all tables to the end of the generated Markdown. Using Document.iter_inner_content() correctly preserves the order of top-level paragraphs and tables while remaining compatible with the project's declared python-docx>=1.1.0 requirement. The refactoring of table serialization into a helper keeps the conversion logic clear, and the added regression coverage thoroughly validates interleaving, headings, lists, consecutive tables, nested tables, revision-marked content, rendering, and delimiter placement. The targeted tests, full test suite, Ruff checks, and formatting checks all pass.

@petertzy
petertzy merged commit d7a0084 into petertzy:main Oct 5, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants