SimplyParseDocs

Async processing

Queue a document, get a document_id straight away, then poll or receive the result by webhook.

POST/dapi/v1/parser/{slug}/parse/async

Queues the document and responds immediately with a document_id.

GET/dapi/v1/parser/{slug}/document/{document_id}

Returns the document's status and, once processed, its entries.

The async endpoint takes exactly the same request as the sync endpoint, but it doesn't wait for processing. You get a document_id back as soon as the file is received, and the result arrives later.

Sync or async?

Sync /parseAsync /parse/async
ResponseThe parsed dataA document_id
Best forSingle documents, interactive uploadsLarge files, batches, background jobs
Client timeoutMust cover the whole processing timeNot a concern
WebhooksNot sentSent to matching webhooks
After a dropped connectionUnclear whether it was processedLook it up by document_id

If you're unsure, use async. It's the more robust default for anything that runs unattended.

1. Submit the document

The Python and Node.js examples on this page reuse the API, HEADERS/headers and slug setup from Parse a document.

curl -X POST "https://api.simplyparse.com/dapi/v1/parser/$PARSER_SLUG/parse/async" \
  -H "Authorization: Token $SIMPLYPARSE_API_TOKEN" \
  -F "file=@annual-report.pdf"
Response
{
  "status": "success",
  "code": "document_queued",
  "message": "Document queued for processing",
  "data": {
    "document_id": "3c59dc04-8a4e-4d6b-9f3e-7a1f0e2b5c66",
    "job_id": "6512bd43-d9ca-4e6f-b1d2-0c7a8e3f4b21"
  }
}

Store the document_id with your own record (for example next to the upload in your database). It's how you find the result, and how you match the webhook when it arrives.

The page limit and your balance are checked before the document is queued, so insufficient_balance, currency_mismatch and page-limit errors come back on this call rather than later.

2. Get the result

You can poll for the result, receive it by webhook, or both.

Option A: Poll

Call Get a document until data.status is final:

data.statusMeaningWhat to do
pendingReceived, not started yetWait and poll again
processingBeing processedWait and poll again
completedDone. entries holds the resultsRead entries
failedProcessing failedCheck the document in the dashboard; resubmit if appropriate

While the document is in progress, the response code is no_parsed_data and entries is empty. That's expected, not an error.

Poll with a growing interval and an overall deadline:

import time


def wait_for_document(slug: str, document_id: str, max_wait: float = 600) -> dict:
    """Poll until the document is completed or failed. Returns `data`."""
    deadline = time.monotonic() + max_wait
    delay = 2.0
    while True:
        response = requests.get(
            f"{API}/{slug}/document/{document_id}", headers=HEADERS, timeout=30
        )
        response.raise_for_status()
        result = response.json()
        if result["status"] != "success":
            raise RuntimeError(f"{result['code']}: {result['message']}")

        data = result["data"]
        if data["status"] in ("completed", "failed"):
            return data
        if time.monotonic() + delay > deadline:
            raise TimeoutError(f"{document_id} still {data['status']} after {max_wait}s")
        time.sleep(delay)
        delay = min(delay * 1.5, 30)  # 2s, 3s, 4.5s ... capped at 30s


data = wait_for_document(slug, document_id)
if data["status"] == "completed":
    for entry in data["entries"]:
        print(entry["id"], entry["is_valid"], entry["parsed_data"])

Option B: Webhook

Add a webhook to the parser and SimplyParse sends each entry to your server as soon as it's ready. You don't need to poll. See Webhooks.

Webhooks give you results with no delay, and a periodic poll catches anything your server missed, for example while it was being deployed:

  1. Submit with async and store the document_id with status pending in your database.
  2. When the webhook arrives, save the result and mark the record done.
  3. Every few minutes, poll Get a document for records still pending after, say, 10 minutes.

Failures

A failed document has no entries. Open the parser's View data tab in the dashboard to see the document and why it failed, for example an unreadable or corrupt file. Some transient errors, such as a brief network problem during processing, are retried automatically before a document is marked failed.

Resubmit only after fixing the cause; resubmitting the same broken file fails the same way.

On this page