Documents

Turn every document into structured, model-ready data.

A professional platform for annotating text, forms, and scanned documents into clean training datasets for OCR and NLP models that extract information without errors and automate document workflows.

Applications

financial-tax-documents.mp4 Live preview

Financial & tax documents

Annotate invoices, tax returns, and receipts so models auto-extract details, amounts, and dates and reduce manual entry and accounting errors.

Technical edge

Multi-modal versatility

Annotate scanned pages, photos of forms, and native digital documents in one place.

AI-powered automation

Pre-label recurring fields and layouts so OCR and NLP models learn document structure faster.

Quality control

Reviewers and validation rules catch misread fields and mislabeled regions before the dataset reaches training.

Collaborative workflow

Route large document batches to annotators and reviewers, each with the right permissions.

Customization & integration

Define custom fields, entities, and tags, then export to your OCR or NLP pipeline via API.

Enterprise-grade security

Keep sensitive records confidential, with on-premise deployment for regulated data.

Annotation types

Bounding box

Rectangular boxes that select text blocks, stamps, signatures, or characters for OCR.

Oriented bounding box

Rotated rectangles for text or stamps on angled or skewed scans.

Polygon

Flexible shapes that outline handwritten notes, damaged text, or logos.

Points

Single points that mark checkboxes (OMR) and reference points to verify form completion.

Lines & multilines

Trace underlines, table separators, and handwritten rules in forms.

Skeletons

Connected keypoints that model handwritten characters or graphics like diagrams and charts.

Cuboid

3D boxes for annotating packing slips or documents in 3D space.

3D point cloud

Spatial labeling that assesses the physical condition of documents or archives via 3D scanning.

Semantic segmentation

Pixel-level classification of document regions: text fields, images, tables, and headers.

Instance segmentation

Pixel-wise separation of individual words or characters.

Bitmap

Raster masks that capture fine details like watermarks or paper texture.

Mesh

Polygonal meshes that create digital copies of documents, preserving bends and surface deformation.

Custom

A mix of annotation methods built around your specific document-processing task.

Put your document data to work

Book a demo and we'll scope your document dataset with you.

Custom solution? hello@keylabs.ai