> ## Documentation Index
> Fetch the complete documentation index at: https://docs.eigenpal.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Parse Document

> Automatic native-first parsing with OCR/vision recovery, explicit completeness, and professional controls.

`ai.parse-v2` reads supported PDF, image, Office, and text files. Start with just the input:

```yaml theme={null}
- name: parse
  type: ai.parse-v2
  with:
    input: '{{ input.document }}'
```

PDFs use local embedded text first. Empty or suspect pages and pages containing embedded raster images are read with the first available allowed image-reading backend: OCR, then vision. Clean native pages are retained. Text and Office files use local parsers; the optical recovery policy applies to PDFs and raster images. Raster images go directly to image reading; multipage TIFF/image frames keep their original page indices.

The worker conservatively treats embedded images as potential text, including scans below a native header. It reads the entire affected page and replaces that page's text, avoiding duplicated native headers. This can process pages containing photos or logos even when their native text is usable. It is a completeness safeguard, not semantic proof that every extracted character is correct. Unicode diagnostics cannot detect valid-looking wrong font mappings.

A uniformly white rendered page is marked `blank` without a provider call. Vision can explicitly report `no-text` for a visual page without readable text. Empty provider responses, missing pages, and missing required evidence remain unresolved. Ambiguous multi-page vision results are retried as individual pages.

Vision selection prefers the deployment parsing default or vision role, then another registered image-capable client. A vision-capable model with a text role can still recover scans; a text-only model cannot. Configuration readiness uses the same client capability inference and requires a live probe to verify custom endpoints.

## Policy and provider selection

```yaml theme={null}
policy:
  native: prefer
  imageReading:
    order: [ocr, vision]
providers:
  ocr: onprem/paddle
  vision: openai/gpt-5.4
```

| Setting | Meaning |
| - | - |
| `policy.native: prefer` | Try native PDF text, then recover unread content. Default. |
| `policy.native: require` | Require local extraction. Never send document content for OCR or vision transcription. Explicit figure enrichment is independent. Scanned/unread pages fail. Raster images are rejected. |
| `policy.native: skip` | Read every nonblank PDF/image page visually. Text and Office files still parse locally. |
| `policy.imageReading.order` | Ordered **allowed** backends: `[ocr, vision]`, `[vision, ocr]`, `[ocr]`, or `[vision]`. A single backend prohibits the other. |
| `providers.ocr` / `providers.vision` | Exact model/provider IDs. Missing or non-vision pins fail; they never silently use a different provider. |

Omitted providers use deployment parsing defaults. Vision otherwise uses the configured vision role, then an available vision-capable registered client. V2 does not inherit the workflow's text-generation `defaultModel` or the tenant's general text default.

On-prem administrators can configure defaults and enforce provider boundaries once:

```yaml theme={null}
# eigenpal.config.yaml; models must also be configured as usual.
deployment:
  parsing:
    ocrProvider: onprem/paddle
    visionProvider: onprem/vision
    allowedProviders: [onprem/paddle, onprem/vision]
```

`allowedProviders` applies to provider overrides and figure enrichment, too. With no OCR, a configured allowed vision provider reads scans automatically. With neither capability, scans cannot be read; configure an approved OCR/vision endpoint. Eigenpal does not bundle a new local OCR service in this release.

Recoverable provider outages use client retry budgets and may proceed to the next allowed backend. Authentication, authorization, cancellation, quota, and validation failures propagate. A sparse or empty OCR result can recover only unresolved pages through vision.

## Output and evidence

```yaml theme={null}
output:
  textFormat: markdown
  nativeWhitespace: reading-order
  require: []
```

`textFormat` supports `markdown`, `plain`, `html`, and `djot`. OCR/vision Markdown is converted locally to the requested format. Use `nativeWhitespace: spatial` with `textFormat: plain` to preserve native PDF column spacing. Spatial whitespace is PDF-only and does not invent coordinates for OCR/vision output.

`output.require` accepts `wordCoordinates` and `tables`. Every nonblank text page must provide the requested evidence. Vision transcription cannot satisfy these requirements. A native page lacking required evidence is eligible for OCR recovery. Table capability does not mean a table exists on every page.

The familiar `text`, `pages`, `document`, and `processingStrategy` fields remain available for downstream extraction, splitting, and review. V2 adds:

* `parserVersion: "2"`.
* `completeness.status`: `complete` or `partial`.
* `completeness.unresolvedPageIndexes`: original 0-based page indices.
* Per-page `status`: `complete`, `blank`, `no-text`, or `unresolved`.
* Per-page `provenance`: source, provider/model when applicable, reason, actual text format, and available evidence capabilities.
* Per-page warnings and native-text diagnostics.
* `usage.ocrPagesProcessed` and `usage.visionPagesProcessed`.

Original page order and indices are retained. Document-wide structured output is retained for fully native output or a complete result from one OCR call. There is no fabricated merged document graph for mixed backends; page-level evidence remains available.

OCR selection, egress, and billing depend on provider capabilities: a provider may receive or process the whole PDF even when only some pages need recovery. Vision renders and sends selected pages. Provider calls and figure descriptions can incur credits.

## Spreadsheet fidelity

XLS and XLSX parse locally into one page per nonempty sheet. Cells retain addresses and their displayed text, formatted using recorded number formats and the workbook's 1900/1904 date system. Formulas use cached values and are never calculated. Empty declared trailing extent does not generate text.

Set `output.includeCellMetadata: true` for ambiguous or mixed date/identifier columns. Each page adds `spreadsheet: {dateSystem, declaredRange, cells}` with raw/displayed values, storage types, formats, addresses, and formulas. The same evidence is annotated in page text, so downstream extraction using only `text` can still see it. A valid date-formatted serial also exposes `dateValue` without a timezone; Excel's fictitious 1900-02-29 is explicitly marked. This describes Excel formatting, not the value's business meaning.

The parser preserves identifiers and strings even under misleading headers. It does not infer dates from unformatted numbers, normalize names, or classify birth dates versus identifiers. Let the extraction model interpret source evidence and domain rules; flag conflicting or ambiguous evidence. For row objects from one selected sheet, use [`transform.xlsx-to-json`](/steps/transform/xlsx-to-json) with `valueMode: displayed` and `includeCellMetadata: true`, passing the full output to extraction.

## Enrichment and advanced controls

```yaml theme={null}
enrichment:
  figures:
    enabled: false
    # provider: onprem/vision
    # instructions: Describe charts and signatures.
    # reasoningEffort: low
advanced:
  cache: false
  allowPartial: false
  # languages: [en, de]
  vision:
    maxConcurrency: 3
    pagesPerBatch: 5
    renderScale: 2
    imageQuality: 85
    # instructions: Faithfully transcribe all visible text; do not summarize.
    # reasoningEffort: low
```

Figure descriptions are a separate opt-in vision pass and cannot run under native-only policy. Enrichment errors fail v2 instead of silently omitting a requested pass.

`advanced.allowPartial: true` returns unresolved pages with explicit diagnostics. The default fails whenever a page remains unresolved. An empty provider response is never accepted as completed parsing.

`advanced.cache: true` reuses only complete results. Cache identity includes v2 policy/version, file identity, settings, and resolved provider IDs. Cache hits do not incur new parsing usage. Cache JSON remains in tenant storage without automatic expiry.

## Readiness and migration

```bash theme={null}
eigenpal workflow step-type get ai.parse-v2
eigenpal models parser-readiness --json
eigenpal models list --json
```

Readiness reports configured, deployment-approved image-reading capabilities. It does **not** probe worker binaries or live provider health. Before accepting customer traffic, run native, scanned, and mixed fixtures through the deployment. For the bundled headless `parse-document` workflow:

```bash theme={null}
bun scripts/smoke/parse-v2.ts --base-url http://127.0.0.1:8080
```

The smoke test sends only a generated native/scanned fixture and checks that both page values were recovered. Use `HEADLESS_API_TOKEN` for authenticated headless deployments.

Legacy `ai.parse` keeps its existing behavior. Migration writes a new local workflow file and never changes or publishes the source:

```bash theme={null}
# Keep deliberate native/OCR/vision restrictions.
eigenpal workflow migrate-parser workflow.yaml --policy preserve --out workflow-v2.yaml

# Adopt native-first automatic OCR/vision recovery.
eigenpal workflow migrate-parser workflow.yaml --policy auto --out workflow-v2-auto.yaml
```

Preserving `native-or-ocr` preserves its OCR-only restriction: it still cannot use vision. Choose automatic policy to fix the OCR-absent/vision-available case. Migration reports broadened backend use, formatting changes, and document-wide `nativeText` behavior changes. Unknown settings require review rather than being silently dropped. The builder offers the same two reversible migration choices.

Validate and compare representative documents before publishing, including downstream fields, evidence, page coverage, latency, and costs. V2 completeness checks can expose failures that legacy parsing previously accepted. Migration is not a byte-identical output promise.

Figure enrichment is independent of text transcription policy and must be explicitly enabled. It may call vision even with `policy.native: require`; keep it disabled when all document processing must remain local. Requested enrichment failures fail the step.

## Configuration reference

<ParamField path="policy" type="object" required>
  <Expandable title="policy properties">
    <ParamField path="native" type="&#x22;prefer&#x22; | &#x22;require&#x22; | &#x22;skip&#x22;" default="prefer">
      Prefer local text extraction, require local text extraction with no image transcription, or skip native PDF extraction. Figure enrichment is independent.
    </ParamField>

    <ParamField path="imageReading" type="object" required>
      <Expandable title="imageReading properties">
        <ParamField path="order" type="array<&#x22;ocr&#x22; | &#x22;vision&#x22;>" default="[&#x22;ocr&#x22;,&#x22;vision&#x22;]">
          Ordered allowed backends. A single entry prohibits the other backend.
        </ParamField>
      </Expandable>
    </ParamField>
  </Expandable>
</ParamField>

<ParamField path="providers" type="object" required>
  <Expandable title="providers properties">
    <ParamField path="ocr" type="string">
      Exact configured OCR provider ID. Missing pins fail; they never resolve to another provider.
    </ParamField>

    <ParamField path="vision" type="string">
      Exact configured vision provider ID. Otherwise uses the deployment parsing default or vision role.
    </ParamField>
  </Expandable>
</ParamField>

<ParamField path="output" type="object" required>
  <Expandable title="output properties">
    <ParamField path="textFormat" type="&#x22;plain&#x22; | &#x22;markdown&#x22; | &#x22;html&#x22; | &#x22;djot&#x22;" default="markdown" />

    <ParamField path="includeCellMetadata" type="boolean" default="false">
      Include spreadsheet cell evidence in pages\[].spreadsheet and annotate text for downstream extraction. Preserves raw/displayed values and workbook date system.
    </ParamField>

    <ParamField path="nativeWhitespace" type="&#x22;reading-order&#x22; | &#x22;spatial&#x22;" default="reading-order">
      Spatial preserves native PDF column spacing. It does not promise OCR/vision coordinates.
    </ParamField>

    <ParamField path="require" type="array<&#x22;wordCoordinates&#x22; | &#x22;tables&#x22;>" default="[]">
      Required evidence capabilities on nonblank pages. Backends unable to satisfy them are rejected.
    </ParamField>
  </Expandable>
</ParamField>

<ParamField path="enrichment" type="object" required>
  <Expandable title="enrichment properties">
    <ParamField path="figures" type="object" required>
      <Expandable title="figures properties">
        <ParamField path="enabled" type="boolean" default="false" />

        <ParamField path="provider" type="string" />

        <ParamField path="instructions" type="string" />

        <ParamField path="reasoningEffort" type="&#x22;none&#x22; | &#x22;minimal&#x22; | &#x22;low&#x22; | &#x22;medium&#x22; | &#x22;high&#x22; | &#x22;xhigh&#x22; | &#x22;max&#x22;" />
      </Expandable>
    </ParamField>
  </Expandable>
</ParamField>

<ParamField path="advanced" type="object" required>
  <Expandable title="advanced properties">
    <ParamField path="vision" type="object" required>
      <Expandable title="vision properties">
        <ParamField path="maxConcurrency" type="integer" default="3" />

        <ParamField path="pagesPerBatch" type="integer" default="5" />

        <ParamField path="renderScale" type="number" default="2" />

        <ParamField path="imageQuality" type="integer" default="85" />

        <ParamField path="instructions" type="string">
          Transcription instructions, never document-specific extraction/schema instructions.
        </ParamField>

        <ParamField path="reasoningEffort" type="&#x22;none&#x22; | &#x22;minimal&#x22; | &#x22;low&#x22; | &#x22;medium&#x22; | &#x22;high&#x22; | &#x22;xhigh&#x22; | &#x22;max&#x22;" />
      </Expandable>
    </ParamField>

    <ParamField path="languages" type="array<string>" />

    <ParamField path="allowPartial" type="boolean" default="false">
      Return unresolved pages explicitly instead of failing. Default false.
    </ParamField>

    <ParamField path="cache" type="boolean" default="false" />
  </Expandable>
</ParamField>

<ParamField path="input" type="string" required>
  File reference or template expression for the document
</ParamField>

## Output reference

<ResponseField path="document" type="object" required>
  <Expandable title="document properties">
    <ResponseField path="filename" type="string" required />

    <ResponseField path="mimeType" type="string" required />

    <ResponseField path="size" type="number" required>
      File size in bytes
    </ResponseField>

    <ResponseField path="storageRef" type="string">
      Storage reference for the original file
    </ResponseField>
  </Expandable>
</ResponseField>

<ResponseField path="usage" type="object" required>
  <Expandable title="usage properties">
    <ResponseField path="pageCount" type="number" required />

    <ResponseField path="processingTimeMs" type="number" />

    <ResponseField path="ocrPagesProcessed" type="integer" />

    <ResponseField path="visionPagesProcessed" type="integer" required />

    <ResponseField path="visionPromptTokens" type="integer" />

    <ResponseField path="visionCompletionTokens" type="integer" />

    <ResponseField path="cached" type="boolean" />
  </Expandable>
</ResponseField>

<ResponseField path="pages" type="array<object>" required>
  <Expandable title="pages properties">
    <ResponseField path="spreadsheet" type="object">
      Opt-in original spreadsheet cell evidence

      <Expandable title="spreadsheet properties">
        <ResponseField path="dateSystem" type="&#x22;1900&#x22; | &#x22;1904&#x22;" required />

        <ResponseField path="declaredRange" type="string" />

        <ResponseField path="cells" type="array<object>" required />
      </Expandable>
    </ResponseField>

    <ResponseField path="pageIndex" type="number" required>
      0-based page index
    </ResponseField>

    <ResponseField path="text" type="string" required>
      Extracted page text (markdown/HTML/plain, or spatial layout text)
    </ResponseField>

    <ResponseField path="pageName" type="string">
      Page/sheet name (e.g., Excel sheet name)
    </ResponseField>

    <ResponseField path="layoutElements" type="array<object>">
      Semantic layout elements in reading order

      <Expandable title="layoutElements properties">
        <ResponseField path="type" type="&#x22;paragraph&#x22; | &#x22;title&#x22; | &#x22;sectionHeading&#x22; | &#x22;pageHeader&#x22; | &#x22;pageFooter&#x22; | &#x22;pageNumber&#x22; | &#x22;footnote&#x22; | &#x22;table&#x22; | &#x22;figure&#x22; | &#x22;formulaBlock&#x22;" required />

        <ResponseField path="content" type="string" required />

        <ResponseField path="table" type="object" />

        <ResponseField path="caption" type="string" />

        <ResponseField path="figureId" type="string" />

        <ResponseField path="boundingRegion" type="object" />
      </Expandable>
    </ResponseField>

    <ResponseField path="words" type="array<object>">
      Word-level positions

      <Expandable title="words properties">
        <ResponseField path="text" type="string" required />

        <ResponseField path="confidence" type="number" />

        <ResponseField path="boundingRegion" type="object" />
      </Expandable>
    </ResponseField>

    <ResponseField path="lines" type="array<object>">
      Line-level positions

      <Expandable title="lines properties">
        <ResponseField path="text" type="string" required />

        <ResponseField path="confidence" type="number" />

        <ResponseField path="boundingRegion" type="object" />
      </Expandable>
    </ResponseField>

    <ResponseField path="tables" type="array<object>">
      Extracted tables

      <Expandable title="tables properties">
        <ResponseField path="cells" type="array<object>" required />

        <ResponseField path="rowCount" type="number" required />

        <ResponseField path="columnCount" type="number" required />

        <ResponseField path="boundingRegion" type="object" />
      </Expandable>
    </ResponseField>

    <ResponseField path="width" type="number">
      Page width
    </ResponseField>

    <ResponseField path="height" type="number">
      Page height
    </ResponseField>

    <ResponseField path="unit" type="&#x22;pixel&#x22; | &#x22;inch&#x22; | &#x22;point&#x22; | &#x22;normalized&#x22;">
      Unit for width/height and bounding regions
    </ResponseField>

    <ResponseField path="confidence" type="number">
      Overall page confidence
    </ResponseField>

    <ResponseField path="nativeTextQuality" type="object">
      Per-page diagnosis of native-extracted text. Reports objectively detectable anomalies (empty, U+FFFD, forbidden controls, unassigned/noncharacter code points, heavy private-use). Does not certify semantic correctness, valid-looking wrong mappings and literal "?" cannot be distinguished from legitimate text.

      <Expandable title="nativeTextQuality properties">
        <ResponseField path="status" type="&#x22;clean&#x22; | &#x22;empty&#x22; | &#x22;suspect&#x22;" required />

        <ResponseField path="reasons" type="array<&#x22;empty&#x22; | &#x22;replacement-character&#x22; | &#x22;lone-surrogate&#x22; | &#x22;forbidden-control&#x22; | &#x22;unassigned&#x22; | &#x22;noncharacter&#x22; | &#x22;private-use&#x22;>" required />

        <ResponseField path="counts" type="object" required />

        <ResponseField path="profile" type="object" required />
      </Expandable>
    </ResponseField>

    <ResponseField path="provenance" type="object" required>
      <Expandable title="provenance properties">
        <ResponseField path="source" type="&#x22;native&#x22; | &#x22;ocr&#x22; | &#x22;vision&#x22;" required />

        <ResponseField path="provider" type="string" />

        <ResponseField path="model" type="string" />

        <ResponseField path="reason" type="string" required />

        <ResponseField path="textFormat" type="&#x22;plain&#x22; | &#x22;markdown&#x22; | &#x22;html&#x22; | &#x22;djot&#x22;" required />

        <ResponseField path="capabilities" type="array<&#x22;wordCoordinates&#x22; | &#x22;tables&#x22;>" required />
      </Expandable>
    </ResponseField>

    <ResponseField path="status" type="&#x22;complete&#x22; | &#x22;blank&#x22; | &#x22;no-text&#x22; | &#x22;unresolved&#x22;" required />

    <ResponseField path="warnings" type="array<string>" default="[]" />
  </Expandable>
</ResponseField>

<ResponseField path="text" type="string" required>
  Combined text from all pages
</ResponseField>

<ResponseField path="parserType" type="&#x22;plaintext&#x22; | &#x22;office&#x22; | &#x22;llm-vision&#x22; | &#x22;ocr&#x22;" required />

<ResponseField path="parserVersion" type="string" required />

<ResponseField path="model" type="string">
  Model used (for LLM/OCR parsers)
</ResponseField>

<ResponseField path="processingStrategy" type="&#x22;native&#x22; | &#x22;ocr&#x22; | &#x22;vision&#x22; | &#x22;hybrid&#x22;">
  How this result was produced: `native` (local PDF text only), `ocr`, `vision`, or `hybrid` (native pages merged with OCR on fallback pages).
</ResponseField>

<ResponseField path="structured" type="record<string, unknown>">
  Canonical structured document with ordered blocks, regions, bounding boxes, tables, figures, and chunks when supported by the parser
</ResponseField>

<ResponseField path="completeness" type="object" required>
  <Expandable title="completeness properties">
    <ResponseField path="status" type="&#x22;complete&#x22; | &#x22;partial&#x22;" required />

    <ResponseField path="unresolvedPageIndexes" type="array<integer>" required />
  </Expandable>
</ResponseField>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.