ai.parse-v2 reads supported PDF, image, Office, and text files. Start with just the input:
blank without a provider call. Vision can explicitly report no-text for a visual page without readable text. Empty provider responses, missing pages, and missing required evidence remain unresolved. Ambiguous multi-page vision results are retried as individual pages.
Vision selection prefers the deployment parsing default or vision role, then another registered image-capable client. A vision-capable model with a text role can still recover scans; a text-only model cannot. Configuration readiness uses the same client capability inference and requires a live probe to verify custom endpoints.
Policy and provider selection
Omitted providers use deployment parsing defaults. Vision otherwise uses the configured vision role, then an available vision-capable registered client. V2 does not inherit the workflow’s text-generation
defaultModel or the tenant’s general text default.
On-prem administrators can configure defaults and enforce provider boundaries once:
allowedProviders applies to provider overrides and figure enrichment, too. With no OCR, a configured allowed vision provider reads scans automatically. With neither capability, scans cannot be read; configure an approved OCR/vision endpoint. Eigenpal does not bundle a new local OCR service in this release.
Recoverable provider outages use client retry budgets and may proceed to the next allowed backend. Authentication, authorization, cancellation, quota, and validation failures propagate. A sparse or empty OCR result can recover only unresolved pages through vision.
Output and evidence
textFormat supports markdown, plain, html, and djot. OCR/vision Markdown is converted locally to the requested format. Use nativeWhitespace: spatial with textFormat: plain to preserve native PDF column spacing. Spatial whitespace is PDF-only and does not invent coordinates for OCR/vision output.
output.require accepts wordCoordinates and tables. Every nonblank text page must provide the requested evidence. Vision transcription cannot satisfy these requirements. A native page lacking required evidence is eligible for OCR recovery. Table capability does not mean a table exists on every page.
The familiar text, pages, document, and processingStrategy fields remain available for downstream extraction, splitting, and review. V2 adds:
parserVersion: "2".completeness.status:completeorpartial.completeness.unresolvedPageIndexes: original 0-based page indices.- Per-page
status:complete,blank,no-text, orunresolved. - Per-page
provenance: source, provider/model when applicable, reason, actual text format, and available evidence capabilities. - Per-page warnings and native-text diagnostics.
usage.ocrPagesProcessedandusage.visionPagesProcessed.
Spreadsheet fidelity
XLS and XLSX parse locally into one page per nonempty sheet. Cells retain addresses and their displayed text, formatted using recorded number formats and the workbook’s 1900/1904 date system. Formulas use cached values and are never calculated. Empty declared trailing extent does not generate text. Setoutput.includeCellMetadata: true for ambiguous or mixed date/identifier columns. Each page adds spreadsheet: {dateSystem, declaredRange, cells} with raw/displayed values, storage types, formats, addresses, and formulas. The same evidence is annotated in page text, so downstream extraction using only text can still see it. A valid date-formatted serial also exposes dateValue without a timezone; Excel’s fictitious 1900-02-29 is explicitly marked. This describes Excel formatting, not the value’s business meaning.
The parser preserves identifiers and strings even under misleading headers. It does not infer dates from unformatted numbers, normalize names, or classify birth dates versus identifiers. Let the extraction model interpret source evidence and domain rules; flag conflicting or ambiguous evidence. For row objects from one selected sheet, use transform.xlsx-to-json with valueMode: displayed and includeCellMetadata: true, passing the full output to extraction.
Enrichment and advanced controls
advanced.allowPartial: true returns unresolved pages with explicit diagnostics. The default fails whenever a page remains unresolved. An empty provider response is never accepted as completed parsing.
advanced.cache: true reuses only complete results. Cache identity includes v2 policy/version, file identity, settings, and resolved provider IDs. Cache hits do not incur new parsing usage. Cache JSON remains in tenant storage without automatic expiry.
Readiness and migration
parse-document workflow:
HEADLESS_API_TOKEN for authenticated headless deployments.
Legacy ai.parse keeps its existing behavior. Migration writes a new local workflow file and never changes or publishes the source:
native-or-ocr preserves its OCR-only restriction: it still cannot use vision. Choose automatic policy to fix the OCR-absent/vision-available case. Migration reports broadened backend use, formatting changes, and document-wide nativeText behavior changes. Unknown settings require review rather than being silently dropped. The builder offers the same two reversible migration choices.
Validate and compare representative documents before publishing, including downstream fields, evidence, page coverage, latency, and costs. V2 completeness checks can expose failures that legacy parsing previously accepted. Migration is not a byte-identical output promise.
Figure enrichment is independent of text transcription policy and must be explicitly enabled. It may call vision even with policy.native: require; keep it disabled when all document processing must remain local. Requested enrichment failures fail the step.
Configuration reference
object
required
object
required
object
required
object
required
object
required
string
required
File reference or template expression for the document
Output reference
Combined text from all pages
Model used (for LLM/OCR parsers)
How this result was produced:
native (local PDF text only), ocr, vision, or hybrid (native pages merged with OCR on fallback pages).Canonical structured document with ordered blocks, regions, bounding boxes, tables, figures, and chunks when supported by the parser