Take a pile of free text nobody has labelled and end up with a classifier that picks a label or honestly declines — small enough to run on a CPU or a phone.
The whole pipeline, in the order you would really do it: get the text, beat a no-model baseline, compare four families of encoder on your own data and write down why you picked one, convert it to something a phone can run, let clustering propose the labels, turn scores into a decision that can say nothing, and prove the shipped build gives the same answers your laptop did.
What you end up with
- A 25 MB sentence encoder that takes text and returns one vector
- A label set you derived from your own data, not one you guessed
- A 100 KB index of prototypes, and a scorer that abstains on a tie
- A parity fixture proving the device agrees with Python
What you will type
Pythonscikit-learnONNX RuntimeTypeScript
The practical end of Classifying with small models. It restates what it needs as it goes, so it stands on its own.
#classification#embeddings#evaluation#deployment