[ CLASSIFICATION AT SCALE ]
Sort ten thousand items without reading ten thousand items.
Shared inboxes, archives and ticket queues are sorted by hand, and the categories drift. Compact embedding-based classifiers label items in bulk, cheaply and at speed, and are measured against your labelled sample before anything goes live.
[ 01 — WHERE IT GOES WRONG ]
Why classification projects disappoint.
The model is rarely the problem. The categories and the yardstick are.
- [ 01 ]If two people label the same mail differently, no model can be right. The label scheme has to be settled on paper first.
- [ 02 ]
A large model for everything
It works at 500 items and gets slow and costly at 50,000. Volume changes the right tool. - [ 03 ]
A number without a yardstick
A vendor says 95 percent accurate. On whose data, and per category? Only a held-out sample from your experts answers that. - [ 04 ]
Guessing on the uncertain cases
A classifier that forces an answer on every item puts wrong labels quietly into your process. Uncertain items need a route to a person.
[ 02 — WHAT YOU GET ]
A classifier you can test.
Large language models can classify, but at volume they are slow and costly. Compact models do the bulk, and a large model or a person handles the uncertain rest.
- [ 01 ]
Label scheme
A clear set of categories with definitions and edge cases, agreed with the people who use the result. - [ 02 ]
Labelled baseline
A sample labelled by your experts. It is the yardstick for every version of the classifier. - [ 03 ]
Compact classifier
An embedding-based model, such as a JEPA-style model, open-source and run locally, or via a provider's API, whichever fits your data rules. - [ 04 ]
Confidence and routing
Each item gets a score. Low-confidence items go to a person or a larger model instead of being guessed. - [ 05 ]
Measured accuracy
Precision and recall per category on held-out data, and a check on fresh samples every month.
[ 03 — HOW WE DE-RISK IT ]
How niivo takes the risk out.
You see the accuracy before you commit to a rollout.
- [ 01 ]
A threshold you set
We only go live if accuracy meets the level you define, per category. No number is promised up front. - [ 02 ]
Local when it must be
Open-source embedding models run on your own servers, so content stays with you. An API is used only where your data rules allow it. - [ 03 ]
A review queue by design
Low-confidence items go to a person, and their corrections feed the next version. The error rate goes down, visibly. - [ 04 ]
Monthly re-checks
Categories change and data drifts. A fresh sample is checked every month, and versions and results are kept for comparison.
[ 04 — PROCESS ]
Label, train, measure, run.
The labelled sample decides what is good enough.
- 01
Define categories
We work with your team on a label scheme and resolve the ambiguous cases on paper.1 week - 02
Build the baseline
Your experts label a sample. We support with tooling and check agreement between labellers.1 to 2 weeks - 03
Train and evaluate
We compare candidates on held-out data and report accuracy per category, including where it fails.2 weeks - 04
Run and monitor
Batch or live classification with confidence thresholds, a review queue and monthly re-checks.Ongoing
[ 05 — USE CASES ]
Volumes that fit.
Typical scenarios. None of these describe a specific client.
Shared inbox sorting
- Today
- Staff read every mail to decide which team handles it.
- With the workflow
- The classifier assigns a team and a topic. Low-confidence mails land in a review queue.
Document archive tagging
- Today
- Years of scans and PDFs have no consistent type or metadata.
- With the workflow
- A batch run tags document types once, and new documents are tagged as they arrive.
Ticket categorisation
- Today
- Categories are picked by hand and drift over time.
- With the workflow
- The classifier suggests a category with a score. Agents confirm or correct, and corrections feed the next version.
[ 06 — SCOPE & PRICE ]
A proof of concept with a measured result.
Manual sorting costs a little on every item, forever. Compare the proof of concept with your current volume times the minutes per item. The result is an accuracy table, so the go decision is arithmetic, not faith.
Classification proof of concept · 4 to 6 weeks
Running costs depend on volume and hosting and are stated before rollout. Extensions: CHF 2,200 per day.
- Label scheme and labelled baseline
- Candidate classifiers compared on held-out data
- Accuracy per category, including failures
- Confidence thresholds and review queue design
[ 07 — QUESTIONS ]
What people ask about classification.
How accurate will it be?
We do not promise a number up front. We measure it on your labelled sample and only go live if it meets the threshold you set.
How much labelled data do we need?
Often a few hundred examples per category are enough to start, fewer for simple schemes. We tell you after a first test.
Why not just use a large language model?
You can, for small volumes. At tens of thousands of items, compact classifiers are faster and cheaper, and we can still use a large model for the uncertain cases.
Can it run without sending documents to an external provider?
Yes. Open-source embedding models run on your own servers, so the content stays with you.
What happens when categories change?
We update the scheme, add labelled examples and re-evaluate. Version history and results are kept so you can compare.
What if it mislabels something important?
Low-confidence items go to a person instead of being guessed, and you can set stricter thresholds for critical categories. Errors are measured per category, so you know where the risk sits.
Bring a sample of what needs sorting.
We test a small set and show you the accuracy before you decide.