- Multipart upload sends the file to the Eigenpal app. Maintained clients use it while the complete request fits the deployment’s configured body limit.
- Storage-direct upload is selected for larger files when enabled and sends bytes directly to object storage instead of through a web function.
file_... id. Use it in run JSON as { "$fileId": "file_..." }.
Parseable documents support files up to 100 MiB. ZIP archives are also accepted
as durable workflow inputs and default to a separate 5 GiB ceiling. Set
EIGENPAL_ARCHIVE_MAX_BYTES or deployment.archiveMaxBytes to match your
deployment’s storage and ingress policy.
Archive capability matrix
ZIP is the only archive format in this version. Encrypted ZIP entries are rejected. A ZIP may contain at most 10,000 central-directory entries. Each extracted member is capped at 64 MiB inflated and compressed. One ZIP reader may inflate at most 512 MiB across reads; parallel extracts share a 512 MiB process-wide concurrent inflation budget (
ZIP_PROCESS_INFLATION_BUDGET_BYTES). Nested ZIP files stay opaque application/zip objects until you pass them to another list, extract, or search step.
File references: runs vs dataset archives
These look similar but serve different paths:
Dataset
$file paths are resolved when you push or import the archive. Run
$fileId refs and multipart files.<field> parts are resolved when the run
starts.
A completed durable { "$fileId": "file_..." } becomes a private run-owned
reference to the original stored bytes. Eigenpal does not duplicate that object
per run, so multi-gigabyte ZIP inputs stay a single copy. Deleting the reusable
file leaves past runs readable until those runs themselves are deleted. Direct
multipart uploads on POST /api/v1/runs and temporary purpose: "run-input"
pre-uploads still materialize as run-owned copies.
Storage-direct upload
First declare the file before sending its bytes:transport: "presigned-put" with:
uploadIdand the reservedfileId- an expiring
url - exact
headersto include in thePUT expiresAtandmaxFileSizeBytes
PUT succeeds, complete the upload:
Resumable multipart (multi-GB ZIP)
When the declared size needs more than a single PUT, hosted S3/R2 deployments returntransport: "presigned-multipart" instead of a single url. The session lasts up to 24 hours. Part URLs are minted on demand and expire in minutes — list completed parts, then presign only the gaps:
bundle.bin with contentType: "application/zip" stays a ZIP. Maintained SDKs and the CLI do this automatically.
Start a run from the reusable id. The platform keeps a private run-owned alias of the original bytes:
Multipart upload
Small files and deployments with direct transfer disabled returntransport: "multipart". Existing clients can also call that operation directly:
/api/v1/files and at https://studio.eigenpal.com/api/v1/files.
Vercel rejects function bodies above approximately 4.5 MB before Eigenpal can return its
structured error. Use storage-direct upload for larger cloud files. Self-hosted ingress can accept
larger multipart bodies when its proxy is configured accordingly. Parseable documents remain
limited to 100 MiB; ZIP archives use the deployment’s separate archive ceiling.
Deployment behavior
Maintained Eigenpal Files clients negotiate before transferring bytes:- A file within
deployment.multipartMaxBytesreturnstransport: "multipart". - A larger file returns a signed upload when direct storage is enabled and reachable.
- A local or private-storage deployment returns
transport: "multipart"regardless of size when direct storage is disabled. - A failed signed transfer is retried as a signed transfer; clients do not automatically send the same large file through multipart.
PUT, Content-Type, signed x-amz-meta-* headers, and the ETag response header. Eigenpal cloud configures this automatically. Self-hosted deployments default to multipart; operators can configure equivalent Ceph/S3 CORS, set EIGENPAL_BROWSER_STORAGE_ORIGINS to the comma-separated storage origins in the app container environment (each a bare origin such as https://s3.example.com, lowercase, https or http on localhost only; other values are ignored), and set EIGENPAL_DIRECT_FILE_UPLOADS=1 to opt in without changing clients.
The deployment limit is optional and defaults to 4718592 bytes (4.5 MiB). Configure it in the mounted eigenpal.config.yaml:
multipartMaxBytes: null to disable size-based direct transfer and keep uploads and downloads on the app server. This is the default in the supplied on-prem manifests. EIGENPAL_MULTIPART_MAX_BYTES overrides the YAML value for container/platform configuration; it accepts a non-negative byte count or none.
Cloud applies the same reusable-pool storage quota to multipart and storage-direct completion, including temporary run-input and builder-attachment files. Delete reusable files when they are no longer needed. Self-hosted deployments are unlimited by default and can set EIGENPAL_REUSABLE_FILES_MAX_BYTES to enforce a per-tenant byte limit.
Run pre-uploads (purpose: run-input)
Maintained SDKs and the CLI automatically pre-upload large run inputs through the Files API so the run request stays under the configured multipart limit. Their client-side default is also 4.5 MiB:
- TypeScript:
new EigenpalClient({ multipartMaxBytes: 20 * 1024 * 1024 }) - Python:
EigenpalClient(multipart_max_bytes=20 * 1024 * 1024) - All maintained clients and the CLI:
EIGENPAL_MULTIPART_MAX_BYTES=20971520
null in an SDK constructor, or EIGENPAL_MULTIPART_MAX_BYTES=none, to disable automatic pre-uploads and keep every run file on multipart. Ensure the target proxy and app can accept the resulting aggregate request. Client configuration controls run-request preparation; the deployment YAML controls Files negotiation and server download redirects, so self-hosted operators should set both consistently.
Studio workflow and quick-run forms use the deployment’s deployment.multipartMaxBytes value directly. Files that would exceed the run multipart budget are uploaded through the negotiated Files API first; null keeps every Studio run file in the run multipart request.
Automatic uploads set purpose: "run-input". The server retains these temporary pool files so a lost run-start response or concurrent retry can safely reuse the same $fileId, then removes abandoned files after 24 hours.
Explicit SDK client.files.upload(...) or POST /v1/files multipart uploads omit purpose, enforce supported document MIME types, and remain durable reusable files until you delete them.
On versioned cloud buckets, deletes (pending uploads, temporary pool files, builder intermediaries, and normal file deletion) leave noncurrent versions for a recovery window — 7 days in production (aligned with bucket backups) and 1 day in staging — then lifecycle removes those prior versions. Current durable objects are not expired by lifecycle.
Studio builder attachments (purpose: builder-attachment)
The agent builder uploads attachments through the Files API with purpose: "builder-attachment". These intermediaries accept any MIME (ZIP, application/octet-stream, etc.) while still enforcing size and filename rules. They are retained briefly for handoff retries and reaped after 24 hours if abandoned. Do not set this purpose on explicit durable files.upload calls.
Git-backed Agent automations do not run ai.search-files. ZIP attachments land in the sandbox workspace (input/ or session uploads) and the agent inspects them with filesystem tools (find, read, unzip -l, bash). Builder attachments stay capped at 100 MiB — that is not the hosted YAML-workflow archive ceiling, and the worker ZIP inspector is not imported into the sandbox.
Builder uploads use the same deployment negotiation described above: small/on-prem files go to the app’s multipart endpoint; files above the configured limit use a signed storage PUT when direct transfer is enabled.
Cleanup operation
The 24-hour rule is database-driven, not an S3 scan and not a long-lived AWS/ECS task. The existing/api/cron/reap-stale-executions maintenance request also processes one bounded batch of at most 100 temporary file records:
- production has one Vercel Cron HTTP invocation every 15 minutes;
- staging reuses the existing daily reset workflow and makes one request before replacing its database;
- each invocation ends after its request, and creates no background task or service;
- one failed object does not abort other rows; its persisted
cleanupAttemptedAtmoves it out of the next batch for 15 minutes; - a failed invocation leaves rows in the database for the next invocation, while bucket lifecycle removes old noncurrent versions as defense in depth.
Downloads
GET /v1/files/{id}/content and run artifact downloads authenticate in the API, then:
- When direct transfer is enabled, responses above
deployment.multipartMaxBytesredirect to a short-lived exact-key signed storage GET that preserves safeContent-TypeandContent-Disposition. - On local/on-prem deployments, the API streams the bytes as before.
- Dynamically generated ZIP downloads (run artifact zip, dataset export) materialize to a cleanup-safe temporary object when needed, then redirect the same way.
- Cloud run-artifact ZIP assembly is capped at 32 MiB of uncompressed input so the Vercel function stays within memory and duration. The API checks cumulative object sizes (known metadata and/or storage
HEAD) before downloading bodies; oversized sets return413withartifacts_zip_too_large— download files individually viaGET /v1/runs/{id}/artifacts/{path}, or pass a smallerfiles=subset. On-prem/local keep a higher in-process ceiling. - Cloud dataset export assembly uses the same 32 MiB source-byte budget and returns
413 dataset_too_largebefore downloading an oversized storage object. Export a selected example subset when the full dataset exceeds the limit.
GET/HEAD. Eigenpal cloud configures this automatically.
API origins
New cloud integrations should usehttps://api.eigenpal.com/v1. Staging uses https://sapi.eigenpal.com/v1. The existing https://studio.eigenpal.com/api/v1 and https://staging.eigenpal.com/api/v1 routes remain compatible and execute the same handlers. Self-hosted SDK and CLI base URLs remain configurable and continue using /api/v1.
To smoke-test upload + download against local or staging: