Skip to content

Stage cloud PDF sources before synchronous parsing - #11

Merged
KSemenenko merged 1 commit into
mainfrom
codex/pdf-cloud-staging
Oct 1, 2026
Merged

KSemenenko merged 1 commit into
mainfrom
codex/pdf-cloud-staging

Conversation

@KSemenenko

Copy link
Copy Markdown
Member

Seekable cloud PDF streams were passed directly to PdfPig, whose synchronous byte reads and random seeks can repeatedly fetch remote ranges during page batches. Stage those streams asynchronously to bounded temporary files before parsing; continue to reuse local file and memory sources by default.

Expose staging mode, copy-buffer size and temporary directory through FileContextOptions. Return pooled buffers, keep temporary sources private on Unix, and clean up files and inputs after completion, cancellation, size rejection and read failures. Release 1.0.14 through the existing main-branch Release workflow.

Validation: 33 focused PDF tests and all 183 tests passed; line coverage 95.65%; Release build, format verification and pack passed; no direct or transitive NuGet vulnerabilities. Regression inputs use real files behind an async-only seekable stream, including a large source and parser/rendering calls after staging.

@KSemenenko
KSemenenko merged commit 069cb39 into main Oct 1, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant