SimplyParseDocs

Errors and retries

Every error code the API returns, what causes it, and whether retrying will help.

How errors are reported

There are three shapes to handle:

  1. The error envelope. Most errors use the normal envelope with "status": "error" and a code. Several of them are sent with HTTP 200, so don't rely on the HTTP status alone.

    { "status": "error", "code": "parser_not_found", "message": "Please provide a valid parser slug", "data": null }
  2. Authentication and lookup errors. HTTP 401 (and 404 on the export endpoint) with a detail message:

    { "detail": "Invalid token." }
  3. Unexpected server errors. HTTP 5xx, with a body that may not be JSON.

A request succeeded only when the HTTP status is 2xx and status is success.

Error codes

codeHTTPCauseRetry?
file_not_found200Neither file nor file_url was sent, or the form field has a different name.No. Fix the request.
parser_not_found200The slug is wrong, the parser is archived, or it's a library template.No. Use your own parser's slug.
file_upload_error200The uploaded file couldn't be stored.Yes, with backoff.
doc:unknown200The document was rejected before processing. The message gives the reason, for example Document exceeds maximum page count limit of 100.No, unless the message says to try again.
doc:<document_id>200Processing of that document failed. The message gives the reason.Only if the message says to try again.
insufficient_balance402Your balance doesn't cover this document.After topping up.
currency_mismatch400The parser's pricing currency differs from your account's base currency.No. Contact support.
document_not_found200No document with that ID belongs to this parser slug.No. Check the ID and slug pair.
entry_not_found200No entry with that ID belongs to this parser slug.No.
no_fields_provided200update_fields was empty or missing.No.

Auth errors (401): see Authentication errors. Retrying won't help until the token is fixed.

Export errors: the export endpoint returns 400 with a plain-text message for bad dates, and 404 {"detail": "Parser not found."} for an unknown parser ID.

5xx: a temporary problem on our side, or a file_url that couldn't be downloaded. Retry with backoff. If it keeps happening with a file_url, check that the URL returns the file to an unauthenticated GET.

Retry safely

Retry only errors that can succeed on a second attempt: 5xx, network errors, timeouts and file_upload_error. Use exponential backoff with jitter and a cap on attempts.

Retrying a parse can charge twice

If a parse request times out or the connection drops, the document may already have been processed and charged. Prefer the async endpoint for unattended work: the submit call returns quickly, and once you have a document_id you can look the result up instead of resubmitting. Unique fields on the parser also stop a resubmitted document from creating a second entry (it is marked duplicate).

A small client wrapper that does this for you:

import random
import time

import requests


class SimplyParseError(Exception):
    def __init__(self, code: str, message: str, retryable: bool = False):
        super().__init__(f"{code}: {message}")
        self.code = code
        self.retryable = retryable


RETRYABLE_CODES = {"file_upload_error"}


def call(method: str, url: str, *, attempts: int = 4, **kwargs) -> dict:
    """Make an API call. Returns `data`, or raises SimplyParseError.

    For uploads, pass the file's bytes so a retry can resend them:
        call("POST", f"{API}/{slug}/parse/async",
             files={"file": ("invoice.pdf", open("invoice.pdf", "rb").read())})
    """
    kwargs.setdefault("timeout", 180)
    for attempt in range(1, attempts + 1):
        try:
            response = requests.request(method, url, headers=HEADERS, **kwargs)
        except (requests.ConnectionError, requests.Timeout) as exc:
            error = SimplyParseError("network_error", str(exc), retryable=True)
        else:
            if response.status_code >= 500:
                error = SimplyParseError(f"http_{response.status_code}", response.text[:200], True)
            elif response.status_code in (401, 404) and "detail" in response.text:
                raise SimplyParseError(f"http_{response.status_code}", response.json()["detail"])
            else:
                body = response.json()
                if body.get("status") == "success":
                    return body["data"]
                code = body.get("code", f"http_{response.status_code}")
                error = SimplyParseError(code, body.get("message", ""), code in RETRYABLE_CODES)

        if not error.retryable or attempt == attempts:
            raise error
        time.sleep(min(2 ** attempt, 30) + random.random())  # backoff with jitter

Retrying uploads

In Python, an open file handle is consumed by the first attempt, so a retry would send an empty file. Pass the file's bytes (as in the docstring above) or reopen the file per attempt. In Node.js, a FormData built from a Blob can be sent again as is.

Log enough to debug

For every failed call, log the endpoint, HTTP status, code, message, your own record ID and, if you have one, the document_id. With those, support can trace exactly what happened.

On this page