Case study
Quote preparation from days to minutes
Figures from manufacturers' certificates and specifications were retyped into spreadsheets by hand, taking one person a full working day. The software now pulls them out, flags what it is unsure about, and a person confirms the result before it is used.
- Industry
- Chemical raw materials distribution
- From signature to production
- under two months
- Document types
- product certificates and specifications
- Document languages
- various, no template per supplier
- What we delivered
- a web application for data extraction and human review
- Where the data sits
- In the EU, including AI processing
A full day of retyping, a quote in days
The client supplies chemical raw materials to manufacturers. Every quote needs figures from the producer's documents, and nobody on the sales team has a full day to spare for retyping them.
To issue a quote they had to retype figures from manufacturers' product certificates and specifications into their own spreadsheet templates. Every manufacturer uses a different layout and often a different language.
One person spent a full working day on it. A quote for thirty items stretched to several days. Manual retyping also produced slips in figures that could reach the customer's quote.
The model pulls the data, a person confirms it
We built a web application that pulls the required fields out of a PDF with a language model. It handles different layouts and languages without a separate template per supplier.
The application works out whether it is a certificate or a product specification, pulls the data into one schema and gives every field a confidence score. A person then checks and confirms. The decision stays with them, not with the model.
Company data stays in the EU. Documents and extracted figures sit in Frankfurt, and the language model that reads them runs in the EU too. The model does not learn from the client's data. Both are guaranteed in the contract.
Extraction by a language model
The model reads the document and writes it into one schema. It handles different languages and layouts without a template per supplier.
Document type detection
The application decides whether it is a certificate or a product specification and processes the data accordingly.
A confidence score per field
Every extracted field carries a confidence score. Low-confidence fields are highlighted on the review screen, so a person knows where to look.
Human review
The review screen shows the PDF beside the extracted data. A person corrects the unclear fields, approves the result and the system records the changes.
Export into their own template
Approved data exports straight into the client's spreadsheet template. No further retyping.
Concurrent processing
The application handles several documents at once. On failure it retries or falls back.
What the acceptance tests showed
- 100%
- document type detected correctly in 860 documents
- over 97%
- mandatory fields extracted correctly
- ~40 s
- p95 per document
Specification table rows came out around 97 percent at confidence 0.8 and above. The figures describe acceptance tests on a sample of the client's documents, not long-run operation. The application replaced a full-time retyping job and cut multi-day quote preparation to minutes.
Is there a process where somebody retypes data all day?
Describe it to us. We will show you whether software can take it over and where a person should stay.
Talk about your problem