SimplyParseDocs

Build a review queue

Auto-approve entries that pass validation, send the rest to a person, and write their corrections back to SimplyParse.

Most documents extract cleanly. The goal of a review queue is to let those flow straight through, while the few that fail your validation rules wait for a person, who sees exactly which values to check.

webhook ─► is_valid? ── yes ─► post to your system
                     └─ no ──► review queue ─► person fixes values
                                                   │
                         PUT corrections ◄─────────┘
                                │
                                └─► post to your system

1. Make validation do the sorting

On the parser, add rules for every field your system depends on: required for must-have values, ranges for amounts, patterns for IDs. Give each rule a clear custom message. It's what the reviewer will read.

Leave Block Invalid Documents off. You want invalid entries delivered to your webhook so they can enter the queue.

2. Route each entry

Extend the handler from Receive results with a webhook:

def process_entry(data: dict) -> None:
    if data["is_valid"]:
        post_to_erp(data["entry_id"], data["parsed_data"])
        mark(data["entry_id"], "posted")
    else:
        create_review_task(
            entry_id=data["entry_id"],
            document_id=data["document_id"],
            values=data["parsed_data"],
            # Show each message next to its field in your review UI.
            problems={e["field"]: e["message"] for e in data["validation_errors"]},
        )
        mark(data["entry_id"], "needs_review")

In your review screen, show the extracted values with the failing fields highlighted, and link the reviewer to the original file. The document's url is available from Get a document.

3. Write corrections back

When the reviewer saves, send their changes to SimplyParse so the dashboard and exports match what your system received. Then post the corrected record:

import os
import requests

API = "https://api.simplyparse.com/dapi/v1/parser"
SLUG = os.environ["PARSER_SLUG"]
HEADERS = {"Authorization": f"Token {os.environ['SIMPLYPARSE_API_TOKEN']}"}


def approve(entry_id: str, corrected: dict, original: dict) -> None:
    # Send only top-level values that changed. Lists such as line_items
    # can't be updated through the API.
    changes = {
        key: value
        for key, value in corrected.items()
        if not isinstance(value, (dict, list)) and original.get(key) != value
    }
    if changes:
        response = requests.put(
            f"{API}/{SLUG}/parsed-data/{entry_id}",
            headers=HEADERS,
            json={"update_fields": changes},
            timeout=30,
        )
        body = response.json()
        if body.get("status") != "success":
            raise RuntimeError(f"{body.get('code')}: {body.get('message') or body.get('detail')}")

    post_to_erp(entry_id, corrected)
    mark(entry_id, "posted")

Validation isn't re-run on corrections

is_valid stays as it was at processing time. Record approval in your own system (the mark(...) calls above) rather than reading it back from SimplyParse.

Measure it

Track the share of entries that need review per parser. If it's high, look at the most common rule and field in validation_errors: often a field needs a clearer definition, or a rule is stricter than real documents allow.

On this page