Validation rules
Check every entry against your business rules, so your code knows which data it can trust.
Extraction answers "what does the document say?" Validation answers "can my system trust this value?" Rules run on every entry, and the result travels with the data:
{
"is_valid": false,
"validation_errors": [
{ "field": "total_amount", "value": -120, "rule": "min_value", "message": "Must be at least 0" }
]
}is_valid is true only when every rule passes. Each failing field contributes one error: the first rule it failed. See Responses for the full error format.
How integrations use it
Validation turns a stream of documents into two lanes:
- Valid entries flow straight into your system with no human in the loop.
- Invalid entries go to a person, with
validation_errorspointing at exactly which values to check. After review, write corrections back with Correct parsed values.
To keep invalid entries out of a webhook-fed system entirely, turn on the parser's Block Invalid Documents setting. Invalid entries then aren't sent to webhooks and stay in the dashboard for review. See Build a review queue.
Add rules
- Open the parser and go to Validation Rules.
- Pick a field and add one or more rules.
- Optionally set a custom error message. It's returned as
messageinvalidation_errors, so write it for the person who will fix the value. - Save, then run Test Parser again to see the effect.
Mark a field Required in Configure Fields to reject empty values.
Rule reference
Rule (rule in errors) | Checks | Applies to | Default message |
|---|---|---|---|
required | The value is present and not empty | All types | This field is required |
min_length / max_length | Text length | Text | Must be at least / no more than N characters |
pattern | Matches a regular expression | Text, ID, address | Does not match the required pattern |
allowed_chars | Only uses permitted characters | Text | Contains characters that are not allowed |
min_value / max_value | Numeric range | Number, integer, currency | Must be at least / no more than N |
precision | Number of decimal places | Number, currency | |
date_range | Date between two bounds | Date, date-time | Date must be on or after / on or before D |
date_format | Date written in a given format | Date, date-time | Date must be in format F |
min_items / max_items | Number of rows in a list | List of Items | Must have at least / no more than N items |
unique_items | No duplicate rows in a list | List of Items | All items must be unique |
is_true / is_false | A flag's value | True / False | Value must be true / false |
lookup | The value exists in one of your lookup tables | Text, number, ID |
A practical strategy
- Get extraction right first. Rules on a schema that's still changing just create noise.
- Protect what your code depends on. Start with the fields your integration reads: document number, date, party, totals, currency.
- Add ranges and formats where real values are predictable, such as
total_amount >= 0or an ID pattern like^INV-\d+$. - Watch real failures for a week, then tighten or relax rules based on what you see.
When an entry fails, ask: is the extraction wrong, is the rule too strict, or is the document genuinely unusual? The fix isn't always in the parser.
Validation is a policy check, not proof
A valid entry passed your rules. It doesn't guarantee every extracted value matches the source document. For high-stakes fields, combine validation with spot checks or review.
Plan availability
Access to validation features can depend on your plan. If the Validation Rules tab isn't available, check your plan in the dashboard.