SimplyParseDocs

Create a parser

Design a parser for one document family, from the sample document to fields and settings.

Create a parser from scratch when you need full control over the extraction schema or when no template is a good fit.

What a parser should represent

A parser should represent one document family that is similar in layout and meaning.

Examples:

  • invoices from one supplier group
  • a recurring internal form
  • one style of statement
  • one type of logistics document

Avoid making one parser handle many unrelated layouts. A narrower parser is usually easier to test, maintain, and trust.

Before you create the parser

Prepare:

  • one representative sample document
  • the list of fields you need
  • a rough idea of whether any fields are repeated, nested, or optional

If you are processing documents that contain repeated rows, such as line items, think about that structure before you start building fields.

Step 1: Add a parser

Go to Parsers and choose Add parser.

You will configure the parser around a sample document and a schema.

Step 2: Choose a processor

Select the processor that best matches your extraction needs.

If you are unsure which one to pick, start with the option that most closely aligns with your document type and test early.

Step 3: Upload a representative sample

Your sample document is the foundation for the first version of the parser.

Choose a sample that is:

  • readable and complete
  • typical of the documents you will actually process
  • not an unusual edge case
  • rich enough to include the fields you care about

Supported browser uploads include PDF, PNG, JPG, JPEG, and WEBP, with a 10 MB limit per file.

Step 4: Design the schema

In Configure Fields, define the output you want.

Start small. Aim for the minimum useful structure, not the final perfect one.

Examples of good initial fields:

  • document number
  • document date
  • supplier name
  • customer name
  • currency
  • subtotal
  • tax
  • total

If your document includes repeated rows, add a list structure for those entries instead of flattening everything into top-level fields.

How to choose good fields

Good fields are:

  • meaningful to your business process
  • stable across documents of the same type
  • easy to verify in testing
  • useful to downstream systems or users

Avoid adding too many low-value fields in version one. Extra fields increase review effort and usually slow down early iteration.

Step 5: Add a clear name and description

Name the parser based on the document family it serves.

Good names are specific, for example:

  • Supplier Invoice Parser
  • Purchase Order Parser
  • Employee Onboarding Form Parser

The description should explain what the parser is intended to process.

Step 6: Save the parser

Save the parser once the initial schema is ready.

This unlocks the rest of the workflow:

  • Validation Rules
  • Test Parser
  • Integrations

Use this sequence for the best results:

  1. Build the smallest useful schema.
  2. Save the parser.
  3. Test on the sample document.
  4. Correct obvious field or structure issues.
  5. Add validation for critical fields.
  6. Process a small batch.

Common parser design mistakes

Trying to support too many document variants

If two document layouts differ significantly, use separate parsers.

Adding too many fields too early

Start with the fields that matter most. Expand after the core parser works.

Using a poor sample document

A low-quality or unrepresentative sample leads to unnecessary rework later.

Designing for edge cases first

Start with the normal case. Handle exceptions after you have a working baseline.

Signs your parser is ready for broader use

Your parser is usually ready for the next stage when:

  • key fields are extracted consistently
  • output structure is stable
  • testing results are easy to verify
  • the parser works across a small sample batch

Once you reach that point, move into Validation rules and Test and process documents.

On this page