Create a parser
Design a parser for one document family, from the sample document to fields and settings.
Create a parser from scratch when you need full control over the extraction schema or when no template is a good fit.
What a parser should represent
A parser should represent one document family that is similar in layout and meaning.
Examples:
- invoices from one supplier group
- a recurring internal form
- one style of statement
- one type of logistics document
Avoid making one parser handle many unrelated layouts. A narrower parser is usually easier to test, maintain, and trust.
Before you create the parser
Prepare:
- one representative sample document
- the list of fields you need
- a rough idea of whether any fields are repeated, nested, or optional
If you are processing documents that contain repeated rows, such as line items, think about that structure before you start building fields.
Step 1: Add a parser
Go to Parsers and choose Add parser.
You will configure the parser around a sample document and a schema.
Step 2: Choose a processor
Select the processor that best matches your extraction needs.
If you are unsure which one to pick, start with the option that most closely aligns with your document type and test early.
Step 3: Upload a representative sample
Your sample document is the foundation for the first version of the parser.
Choose a sample that is:
- readable and complete
- typical of the documents you will actually process
- not an unusual edge case
- rich enough to include the fields you care about
Supported browser uploads include PDF, PNG, JPG, JPEG, and WEBP, with a 10 MB limit per file.
Step 4: Design the schema
In Configure Fields, define the output you want.
Start small. Aim for the minimum useful structure, not the final perfect one.
Examples of good initial fields:
- document number
- document date
- supplier name
- customer name
- currency
- subtotal
- tax
- total
If your document includes repeated rows, add a list structure for those entries instead of flattening everything into top-level fields.
How to choose good fields
Good fields are:
- meaningful to your business process
- stable across documents of the same type
- easy to verify in testing
- useful to downstream systems or users
Avoid adding too many low-value fields in version one. Extra fields increase review effort and usually slow down early iteration.
Step 5: Add a clear name and description
Name the parser based on the document family it serves.
Good names are specific, for example:
- Supplier Invoice Parser
- Purchase Order Parser
- Employee Onboarding Form Parser
The description should explain what the parser is intended to process.
Step 6: Save the parser
Save the parser once the initial schema is ready.
This unlocks the rest of the workflow:
- Validation Rules
- Test Parser
- Integrations
Recommended first-pass workflow
Use this sequence for the best results:
- Build the smallest useful schema.
- Save the parser.
- Test on the sample document.
- Correct obvious field or structure issues.
- Add validation for critical fields.
- Process a small batch.
Common parser design mistakes
Trying to support too many document variants
If two document layouts differ significantly, use separate parsers.
Adding too many fields too early
Start with the fields that matter most. Expand after the core parser works.
Using a poor sample document
A low-quality or unrepresentative sample leads to unnecessary rework later.
Designing for edge cases first
Start with the normal case. Handle exceptions after you have a working baseline.
Signs your parser is ready for broader use
Your parser is usually ready for the next stage when:
- key fields are extracted consistently
- output structure is stable
- testing results are easy to verify
- the parser works across a small sample batch
Once you reach that point, move into Validation rules and Test and process documents.