Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 6 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,12 @@

All notable changes to ManagedCode.FileContext are documented here.

## 1.0.14

- Stage seekable cloud PDF streams asynchronously to bounded temporary files before synchronous parser reads and seeks.
- Preserve local file/memory sources by default and expose typed staging mode, buffer size and temporary-directory options.
- Return pooled staging buffers and delete temporary sources on completion, cancellation and failures.

## 1.0.13

- Parse PDF text, pages and embedded images from bounded seekable streams instead of whole-document arrays.
Expand Down
2 changes: 1 addition & 1 deletion Directory.Build.props
Original file line number Diff line number Diff line change
Expand Up @@ -12,7 +12,7 @@
<AnalysisMode>Recommended</AnalysisMode>
<TreatWarningsAsErrors>true</TreatWarningsAsErrors>
<NoWarn>$(NoWarn);CS1591;MAAI001</NoWarn>
<Version>1.0.13</Version>
<Version>1.0.14</Version>
<PackageVersion>$(Version)</PackageVersion>
</PropertyGroup>

Expand Down
6 changes: 6 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -185,6 +185,12 @@ Results include `StartLine`, `EndLine`, `HasMore`, and `TotalLines` when the end

## Read PDFs and send pages to vision models

Storage-backed PDF tools stage seekable cloud streams asynchronously before parsing. PdfPig's
synchronous reads and seeks then use a temporary file rather than repeated network ranges. Local
seekable files and memory streams remain reusable. `PdfSourceStagingMode` can force `TemporaryFile`
staging, `PdfSourceBufferBytes` bounds the copy buffer, and `PdfTemporaryDirectory` optionally selects
an existing host directory. Disposal, cancellation and staging failures remove temporary sources.

`IFileContextPdf.ReadPdfTextAsync(path)` returns bounded text, `PageCount`, and one-based `PagesWithoutText`. It does not perform OCR. A scanned page can instead be rendered with `RenderPdfPageAsync(path, pageNumber)`, which returns PNG `DataContent`. Use `CountPdfPageImagesAsync` and `ExtractPdfImageAsync` when the original embedded pictures are needed rather than the complete page. The four read-only `file_context_pdf_*` tools expose the same operations from scoped storage.

For an authenticated PDF already held as bytes, `FileContextPdfTextExtractor.Extract`, `FileContextPdfImages.RenderPagePng`, and `FileContextPdfImages.ExtractPageImagesPng` work without storing it. PDF source reads default to 100 MiB and accept `FileContextOptions` for a different limit; page rasterization also uses configured pixel and PNG limits. `FileContextImageContent` creates model-visible `DataContent` from PNG bytes or base64 and `UriContent` from an HTTPS URL. A URL reference is not fetched by FileContext, so the model provider must be able to access it. A host must pass image content to its model as image content. A generic OpenAI Chat function result serializes it as text, so hosts must explicitly bridge image tool results into a multimodal model message.
Expand Down
2 changes: 1 addition & 1 deletion docs/Architecture.md
Original file line number Diff line number Diff line change
Expand Up @@ -82,7 +82,7 @@ source and retaining it through parsing/rendering. `MaximumConcurrentPdfOperatio
`FileContextPdfRenderDocument` exposes page count and sequential page rendering from one parsed PDF;
dispose it after the batch. The low-level synchronous image helpers remain caller-scheduled APIs.

PDF text, page rendering and embedded-image APIs accept bounded seekable streams. Storage-backed PDF tools keep seekable provider streams directly and stage non-seekable sources to an automatically deleted temporary file, never a whole-document managed array. Native raster decoding has a separate per-page source-image pixel budget; lowering output scale does not reduce source bitmap allocation.
PDF text, page rendering and embedded-image APIs accept bounded seekable streams. Storage-backed PDF tools retain seekable local `FileStream`/`MemoryStream` inputs but asynchronously stage other streams, including seekable cloud streams, to an automatically deleted temporary file. PdfPig's synchronous byte reads and seeks then remain local, without blocking on repeated network ranges. `PdfSourceStagingMode.TemporaryFile` also stages local inputs. `PdfSourceBufferBytes` bounds every copy read and `PdfTemporaryDirectory` optionally selects an existing host directory. No whole-document managed array is created. Native raster decoding has a separate per-page source-image pixel budget; lowering output scale does not reduce source bitmap allocation.

All potentially large operations are controlled by `IOptions<FileContextOptions>`: PDF source/page/image budgets, full-read bytes, range bytes, files scanned, bytes per searched file, matches per file, total search results, graph documents, graph source bytes, and exported graph characters. Non-seekable cloud streams are supported by sequential streaming.

Expand Down
7 changes: 7 additions & 0 deletions docs/Features/file-context.md
Original file line number Diff line number Diff line change
Expand Up @@ -115,6 +115,13 @@ Verification: DocumentCreationTests and DocumentValidationTests reopen real form

## PDF reads and vision images

PDF source staging keeps synchronous parser I/O local. The default `Automatic` mode reuses
seekable `FileStream` and `MemoryStream` sources and asynchronously stages all other inputs,
including seekable cloud streams. `TemporaryFile` stages every source. The configured
`PdfSourceBufferBytes` bounds each asynchronous copy read; `PdfTemporaryDirectory` may name an
existing host directory. Size rejection, cancellation and copy failures dispose the input and
delete any staged file. Source ownership continues through document rendering and caller disposal.

DOCX reading uses the native `file_context_docx_text` tool. It reads ordinary paragraph and table
text from the scoped `.docx` package in bounded windows. Each result includes the next paragraph
and character offset when more text remains, so an agent can continue without loading a long
Expand Down
1 change: 1 addition & 0 deletions docs/Testing/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -10,6 +10,7 @@ The suite is integration-first:
- timeout tests cover configured operation expiry, cancellation of every public operation, disabled deadlines, duration validation, and timeout tool results through restored sessions;
- concurrent storage tests write and range-read eight independent files through one shared adapter/service;
- a sparse 1 GiB filesystem test reads bounded line windows repeatedly, rejects full-file loading, caps allocations, and proves that an oversized line fails before it can be buffered in memory.
- PDF cloud-source tests use real files behind an async-only seekable stream, proving parsing uses the staged local file; they cover large inputs, configured buffers/directories, forced staging, local-file reuse, limits, mid-copy cancellation, failure cleanup and private Unix permissions.

Every filesystem test owns a unique temporary root and removes it on disposal. Test execution is serialized so process-wide allocation assertions cannot be distorted by another test. No `IStorage`, Agent Framework, Markdown-LD, or LlmTck mocks are used.

Expand Down
1 change: 1 addition & 0 deletions src/ManagedCode.FileContext/FileContextDefaults.cs
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,7 @@ public static class FileContextDefaults
public const int FirstLineNumber = 1;
public const int MaximumPdfReadBytes = 100 * 1024 * 1024;
public const int MaximumConcurrentPdfOperations = 1;
public const int PdfSourceBufferBytes = 81920;
public const int MaximumImageBytes = 8 * 1024 * 1024;
public const int MaximumDecodedPdfImagePixels = 32_000_000;
public const int MaximumRenderedPagePixels = 4_000_000;
Expand Down
21 changes: 21 additions & 0 deletions src/ManagedCode.FileContext/FileContextOptions.cs
Original file line number Diff line number Diff line change
@@ -1,3 +1,5 @@
using ManagedCode.FileContext.Pdf;

namespace ManagedCode.FileContext;

/// <summary>Controls file access, approval, search, and graph limits for one context provider.</summary>
Expand All @@ -23,6 +25,15 @@ public sealed class FileContextOptions

public int MaximumPdfReadBytes { get; set; } = FileContextDefaults.MaximumPdfReadBytes;

/// <summary>Stages cloud streams before synchronous parsing; Automatic reuses local files and memory.</summary>
public FileContextPdfSourceStagingMode PdfSourceStagingMode { get; set; } = FileContextPdfSourceStagingMode.Automatic;

/// <summary>Maximum bytes requested by one asynchronous PDF source staging read.</summary>
public int PdfSourceBufferBytes { get; set; } = FileContextDefaults.PdfSourceBufferBytes;

/// <summary>Existing directory for temporary PDF sources. Null uses the operating system's temp directory.</summary>
public string? PdfTemporaryDirectory { get; set; }

/// <summary>Maximum simultaneous PDF reads/renders per shared processor, including source buffering.</summary>
public int MaximumConcurrentPdfOperations { get; set; } = FileContextDefaults.MaximumConcurrentPdfOperations;

Expand Down Expand Up @@ -81,6 +92,16 @@ internal void Validate()
{
ValidatePositive(MaximumGeneratedFileBytes, nameof(MaximumGeneratedFileBytes));
ValidatePositive(MaximumPdfReadBytes, nameof(MaximumPdfReadBytes));
ValidatePositive(PdfSourceBufferBytes, nameof(PdfSourceBufferBytes));
if (PdfSourceStagingMode is not FileContextPdfSourceStagingMode.Automatic
and not FileContextPdfSourceStagingMode.TemporaryFile)
{
throw new InvalidOperationException("The PDF source staging mode is invalid.");
}
if (PdfTemporaryDirectory is not null && string.IsNullOrWhiteSpace(PdfTemporaryDirectory))
{
throw new InvalidOperationException("The PDF temporary directory must be a nonempty path or null.");
}
ValidatePositive(MaximumConcurrentPdfOperations, nameof(MaximumConcurrentPdfOperations));
ValidatePositive(MaximumImageBytes, nameof(MaximumImageBytes));
ValidatePositive(MaximumDecodedPdfImagePixels, nameof(MaximumDecodedPdfImagePixels));
Expand Down
70 changes: 51 additions & 19 deletions src/ManagedCode.FileContext/Pdf/FileContextPdfSource.cs
Original file line number Diff line number Diff line change
@@ -1,9 +1,10 @@
using System.Buffers;

namespace ManagedCode.FileContext.Pdf;

/// <summary>Owns a bounded seekable PDF source; non-seekable inputs are staged to disk.</summary>
/// <summary>Owns a bounded local PDF source; cloud inputs are staged before synchronous random reads.</summary>
public sealed class FileContextPdfSource : IAsyncDisposable
{
private const int CopyBufferBytes = 81920;
private const string TemporaryFilePrefix = "filecontext-pdf-";

private FileContextPdfSource(Stream stream) => Stream = stream;
Expand All @@ -24,28 +25,26 @@ public static async Task<FileContextPdfSource> OpenAsync(Stream source, FileCont
if (source.CanSeek)
{
Validate(source, options);
source.Position = 0;
return new FileContextPdfSource(source);
}
staged = CreateTemporaryFile();
var buffer = new byte[CopyBufferBytes];
int read;
while ((read = await source.ReadAsync(buffer, cancellationToken).ConfigureAwait(false)) > 0)
else if (!source.CanRead)
{
if (staged.Length + read > options.MaximumPdfReadBytes)
{
throw new IOException("The PDF exceeds the read limit.");
}
await staged.WriteAsync(buffer.AsMemory(0, read), cancellationToken).ConfigureAwait(false);
throw new ArgumentException("The PDF source must be readable.", nameof(source));
}
if (options.PdfSourceStagingMode == FileContextPdfSourceStagingMode.Automatic
&& source.CanSeek && source is FileStream or MemoryStream)
{
return new FileContextPdfSource(source);
}
staged = CreateTemporaryFile(options);
await CopyAsync(source, staged, options, cancellationToken).ConfigureAwait(false);
staged.Position = 0;
await source.DisposeAsync().ConfigureAwait(false);
return new FileContextPdfSource(staged);
}
catch
{
if (staged is not null) { await staged.DisposeAsync().ConfigureAwait(false); }
await source.DisposeAsync().ConfigureAwait(false);
try { if (staged is not null) { await staged.DisposeAsync().ConfigureAwait(false); } }
finally { await source.DisposeAsync().ConfigureAwait(false); }
throw;
}
}
Expand All @@ -64,10 +63,43 @@ internal static void Validate(Stream source, FileContextOptions options)
source.Position = 0;
}

private static FileStream CreateTemporaryFile() => new(
Path.Combine(Path.GetTempPath(), TemporaryFilePrefix + Guid.NewGuid().ToString("N")),
FileMode.CreateNew, FileAccess.ReadWrite, FileShare.None, CopyBufferBytes,
FileOptions.Asynchronous | FileOptions.DeleteOnClose);
private static async Task CopyAsync(Stream source, Stream staged, FileContextOptions options,
CancellationToken cancellationToken)
{
var buffer = ArrayPool<byte>.Shared.Rent(options.PdfSourceBufferBytes);
try
{
int read;
while ((read = await source.ReadAsync(buffer.AsMemory(0, options.PdfSourceBufferBytes), cancellationToken)
.ConfigureAwait(false)) > 0)
{
if (staged.Length + read > options.MaximumPdfReadBytes)
{
throw new IOException("The PDF exceeds the read limit.");
}
await staged.WriteAsync(buffer.AsMemory(0, read), cancellationToken).ConfigureAwait(false);
}
}
finally { ArrayPool<byte>.Shared.Return(buffer, clearArray: true); }
}

private static FileStream CreateTemporaryFile(FileContextOptions options)
{
var fileOptions = new FileStreamOptions
{
Mode = FileMode.CreateNew,
Access = FileAccess.ReadWrite,
Share = FileShare.None,
BufferSize = options.PdfSourceBufferBytes,
Options = FileOptions.Asynchronous | FileOptions.DeleteOnClose
};
if (!OperatingSystem.IsWindows())
{
fileOptions.UnixCreateMode = UnixFileMode.UserRead | UnixFileMode.UserWrite;
}
return new FileStream(Path.Combine(options.PdfTemporaryDirectory ?? Path.GetTempPath(),
TemporaryFilePrefix + Guid.NewGuid().ToString("N")), fileOptions);
}

public ValueTask DisposeAsync() => Stream.DisposeAsync();
}
11 changes: 11 additions & 0 deletions src/ManagedCode.FileContext/Pdf/FileContextPdfSourceStagingMode.cs
Original file line number Diff line number Diff line change
@@ -0,0 +1,11 @@
namespace ManagedCode.FileContext.Pdf;

/// <summary>Controls where synchronous PDF parsers perform random reads.</summary>
public enum FileContextPdfSourceStagingMode
{
/// <summary>Reuse seekable local files or memory; stage other streams asynchronously to disk.</summary>
Automatic,

/// <summary>Stage every input to a temporary file, including seekable local sources.</summary>
TemporaryFile
}
39 changes: 39 additions & 0 deletions tests/ManagedCode.FileContext.Tests/AsyncOnlySeekablePdfStream.cs
Original file line number Diff line number Diff line change
@@ -0,0 +1,39 @@
namespace ManagedCode.FileContext.Tests;

// Real file contents, with the async-read/synchronous-seek separation of a cloud source.
internal sealed class AsyncOnlySeekablePdfStream(FileStream source) : Stream
{
public bool WasDisposed { get; private set; }
public int MaximumReadRequestBytes { get; private set; }
public CancellationTokenSource? CancelAfterRead { get; set; }
public bool FailRead { get; set; }
public bool HideSeekability { get; set; }
public override bool CanRead => source.CanRead;
public override bool CanSeek => !HideSeekability;
public override bool CanWrite => false;
public override long Length => source.Length;
public override long Position { get => source.Position; set => source.Position = value; }

public override async ValueTask<int> ReadAsync(Memory<byte> buffer, CancellationToken cancellationToken = default)
{
MaximumReadRequestBytes = Math.Max(MaximumReadRequestBytes, buffer.Length);
if (FailRead) { throw new IOException("Source read failed."); }
var count = await source.ReadAsync(buffer, cancellationToken).ConfigureAwait(false);
if (CancelAfterRead is { } cancellation) { await cancellation.CancelAsync().ConfigureAwait(false); }
return count;
}

public override int Read(byte[] buffer, int offset, int count) =>
throw new InvalidOperationException("The parser must never read this remote source synchronously.");
public override long Seek(long offset, SeekOrigin origin) => source.Seek(offset, origin);
public override void Flush() => throw new NotSupportedException();
public override void SetLength(long value) => throw new NotSupportedException();
public override void Write(byte[] buffer, int offset, int count) => throw new NotSupportedException();

protected override void Dispose(bool disposing)
{
WasDisposed = true;
if (disposing) { source.Dispose(); }
base.Dispose(disposing);
}
}
Loading
Loading