AI functions in business
AI document processing: extract, classify, summarise
What is AI document processing?
AI document processing – known internationally as Intelligent Document Processing (IDP) – describes software that automatically turns unstructured documents into structured, machine-usable data. Instead of fixed templates like classic OCR, modern systems use layout understanding, machine learning and large language models. This lets them recognise fields even when every invoice, contract and form looks different.
Three core tasks are central: extraction (reading out relevant fields such as amount, date, contracting party), classification (identifying the document type, e.g. invoice, reminder or contract) and summarisation (condensing long text to its essentials). The goal is not to replace people, but to automate repetitive capture and free specialists for review and exceptions.
How processing works: from OCR to summary
A typical IDP pipeline runs in several steps. Each step can be measured and improved separately:
- Capture and OCR: scanned or photographed paper is converted into machine-readable text; digital PDFs provide the text directly.
- Classification: the system assigns the document to a type (invoice, contract, delivery note, form) and routes it to the right process.
- Extraction: layout and language models read out fields and line items – e.g. invoice number, VAT rate, IBAN, due date or contract term.
- Validation: rules and lookups check the data (totals check, master-data match, IBAN or Swiss QR-bill format).
- Summarisation and analysis: large language models condense contracts into key points, extract deadlines or flag clauses – ideally with a reference back to the source passage.
- Handover: the structured data flows via interfaces into ERP, accounting, DMS or CRM – without manual re-keying.
Invoices, contracts, forms: typical use cases
- Invoices (accounts payable): automatic reading of supplier, amount, VAT, IBAN and line items, matching against purchase order and goods receipt, posting with human approval. Swiss specificity: the QR-bill with its structured QR code makes reliable extraction easier.
- Contracts: extract clauses, notice periods, terms and parties, classify contracts and summarise them for legal review. AI provides the draft – the legal judgement stays with professionals.
- Forms and applications: digitise onboarding, claims or application forms, validate fields and pass them into business systems – including handwritten entries of varying quality.
- Multilingual documents: in Switzerland, documents in German, French, Italian and English arrive in the same inbox. Modern models handle these languages, removing the need for manual pre-sorting.
Accuracy, confidence and human review
Accuracy is not measured in the aggregate but per field: a date may be right 99 percent of the time, a handwritten note far less often. What matters is the confidence score the model outputs for each extraction. It governs which documents pass through untouched (straight-through processing) and which go to people for review (human-in-the-loop).
Good systems therefore rely on a threshold: below a certain confidence, or when validation fails, the document lands in a review screen. Control stays with people, while most of the routine runs automatically. Caution is especially warranted with language-model summaries: models can produce plausible-sounding but false statements (hallucinations). Source references and spot checks are therefore mandatory.
Data protection and compliance in Switzerland
Documents often contain personal data and trade secrets. In Switzerland the revised Federal Act on Data Protection (revFADP) applies, supervised by the FDPIC; for customers in the EU the GDPR may apply in addition. Purpose limitation, data minimisation, a record of processing activities and – for higher-risk projects – a data protection impact assessment all matter.
- Data residency: clarify where processing and storage happen. Swiss or EU hosting, encrypted transfer and a data processing agreement with the provider are central.
- Training on your data: contractually ensure that submitted documents are not used to train third-party models without consent.
- Retention: business records and vouchers must generally be kept for ten years under the Code of Obligations (CO Art. 958f) – archiving must remain audit-proof even when capture is automated.
- Professional secrecy: health, legal or banking records carry additional confidentiality duties that may call for an on-premises or strictly isolated solution.
Introduce it in five steps
- Pick one clearly bounded, high-volume process – such as incoming invoices – rather than automating everything at once.
- Define success criteria: target accuracy per field, acceptable straight-through rate and maximum processing time.
- Test with real, representative documents – including poor scans, foreign languages and edge cases.
- Set up the review workflow and confidence thresholds so uncertain cases reliably reach people.
- Clarify data protection, contracts and data residency early, and continuously monitor and tune accuracy in production.
Frequently asked questions
What is the difference between OCR and AI document processing?
OCR turns image into text but does not understand its meaning. AI document processing builds on top: it classifies the document, recognises fields regardless of layout, validates values and can summarise content. OCR is thus one building block; IDP is the whole process from capture to structured data handover.
How accurate is AI at data extraction?
It depends heavily on document quality, field type and variety, and cannot be captured in a single percentage. Clearly structured fields such as amounts are usually recognised very reliably; handwritten or poorly scanned content much less so. That is why confidence scores decide which documents pass automatically and which are reviewed.
Can language models reliably summarise contracts?
They produce good drafts and quickly find deadlines or clauses, but can confuse details or generate plausible falsehoods. For legally relevant statements you need references to the source passage and a human final check. As an assistant they are strong; as the sole basis for a decision they are unsuitable.
Is AI document processing compliant with the revFADP?
It can be, if the framework is right: a clear processing purpose, data minimisation, a suitable data location, encryption, a data processing agreement and a record of processing activities. For sensitive data, a data protection impact assessment is advisable. Compliance lies not in the tool alone but in how it is configured and used.
How long must documents be retained in Switzerland?
For business records and accounting vouchers, the Code of Obligations (CO Art. 958f) generally requires a ten-year retention period. Storage must be unalterable and traceable. Automated capture does not change this duty – audit-proof archiving must be part of the solution.
On-premises or cloud – which fits better?
Cloud solutions are quick to deploy and powerful but require clarity on data location and contracts. On-premises or Swiss hosting suits cases involving professional secrecy, especially sensitive data or strict internal rules. What decides is the protection need, volume and existing IT – not the technology itself.
Practical AI for your business
From idea to implementation – we show you what is concretely possible in your case.
Request a demo