Parse a document
Send a file or a URL to a parser and get structured data back in the same response.
/dapi/v1/parser/{slug}/parseProcesses the document and responds when the data is ready.
The sync endpoint is the simplest way to integrate: one request in, structured data out. Use it for single documents and interactive flows, such as a user uploading an invoice and waiting to see the extracted values.
For large files, batches or clients with short timeouts, use the async endpoint instead. It takes exactly the same request.
Request
Send multipart/form-data with either a file or a file_url:
Prop
Type
Upload a file
curl -X POST "https://api.simplyparse.com/dapi/v1/parser/$PARSER_SLUG/parse" \
-H "Authorization: Token $SIMPLYPARSE_API_TOKEN" \
-F "file=@invoice.pdf" \
-F "environment=production"Don't set Content-Type yourself
Let your HTTP client build the multipart body and its Content-Type header, which includes the boundary. A hand-written Content-Type: multipart/form-data without a boundary breaks the upload.
Send a URL instead
If the file is already in cloud storage, pass its URL and skip the upload. SimplyParse downloads it with a plain GET and no credentials, so use a public or pre-signed URL (for example an S3 pre-signed URL valid for a few minutes).
curl -X POST "https://api.simplyparse.com/dapi/v1/parser/$PARSER_SLUG/parse" \
-H "Authorization: Token $SIMPLYPARSE_API_TOKEN" \
-F "file_url=https://my-bucket.s3.amazonaws.com/invoices/inv-20418.pdf?X-Amz-Signature=..."The URL must point straight at the file, not at a viewer page, and should return the right Content-Type (for example application/pdf).
Response
{
"status": "success",
"code": "document_processed",
"message": "Document processed successfully",
"data": [
{
"document_id": "8f14e45f-ceea-467f-a8e6-3c1d2f0b9a51",
"status": "completed",
"is_valid": false,
"parsed_data": {
"invoice_number": "INV-20418",
"vendor_name": "Northwind Supplies",
"total_amount": null
},
"validation_errors": [
{
"field": "total_amount",
"value": null,
"rule": "required",
"message": "This field is required",
"is_valid": false,
"parser_field_id": "1f0e3dad-9990-4cb7-a1c8-3a7c3fa0e1b2",
"path": { "key": "total_amount", "depth": 0, "index": 3 }
}
]
}
]
}Each item in data is one entry. Its document_id is the entry's ID, which you pass to Correct parsed values. For every field and status, see Responses.
Files and limits
- Formats: PDF, PNG, JPG/JPEG and WEBP. Images and other non-PDF files count as one page.
- Page limit: each parser has a Maximum Pages setting (default 100). Longer documents are rejected with a
doc:unknownerror that names the limit, before you are charged. Change it in the parser's settings. - Cost: pages × the parser's per-page rate, debited from your balance. See What counts as a page.
- Webhooks: documents processed through the sync endpoint do not trigger webhooks. The response is the result. Use the async endpoint if you want webhook delivery.
Timeouts
The connection stays open while the document is processed, and longer documents take longer. Set a generous client timeout (the examples use 180 seconds) and make sure any proxy or load balancer in front of your code allows it too.
If you can't hold a request open that long, for example on serverless functions with a short limit, use async processing.
A timeout doesn't cancel the work
If your client gives up, the document may still be processed and charged. Before retrying a timed-out request, check the parser's View data tab, or switch to async so every submission has a document_id you can look up.
Multiple records in one file
Some files hold many records: a PDF with twenty invoices, or a statement with one block per customer. Send has_multiple_entries=true to get one entry per record instead of a single merged entry.
curl -X POST "https://api.simplyparse.com/dapi/v1/parser/$PARSER_SLUG/parse" \
-H "Authorization: Token $SIMPLYPARSE_API_TOKEN" \
-F "file=@march-invoices.pdf" \
-F "has_multiple_entries=true"This needs at least one unique field on the parser, such as the invoice number, so records can be told apart. Uniqueness also protects against duplicates across files: if an entry's unique value already exists in the parser, the existing entry is updated and the new one is marked duplicate.
Common mistakes
| Symptom | Likely cause |
|---|---|
parser_not_found | Wrong slug, an archived parser, or a template's slug. Use your own parser's slug from Integrations. |
file_not_found | The form field isn't named file (or file_url), or the body isn't multipart. |
401 with detail | See Authentication errors. |
Every field is null | The document doesn't match the parser, for example a receipt sent to an invoice parser. Test it in the dashboard. |