Project · a full practical build

Build a text classifier

Start with a folder of short text that people typed and nothing else — no labels, no training set. End with something that reads a new sentence and says what it is about, or honestly admits it does not know, in a few megabytes, on a CPU, offline.

23 steps · 6 parts A weekend, or a few evenings Python to prepare · any language to ship

The idea in one page

The problem. You have a few thousand short texts that people wrote themselves — support messages, survey answers, bug reports, search queries, notes in a free-text box. "my card still hasn't arrived", "why was I charged twice", "can't log in on the new phone". You want to know what each one is about so something downstream can act on it. Nobody has labelled any of it, and nobody is going to.

The approach. Three ingredients, none of them trained by you:

Part B & C
An encoder

A frozen, off-the-shelf model that turns any sentence into a fixed list of numbers. Similar meanings land on similar numbers. You choose it by measuring, not by reputation.

Part D
Labels from your data

Encode your real text, let clustering find the groups that are actually there, read them and name them. Each group's average becomes its prototype.

Part E & F
Nearest prototype

For new text: encode, compare against every prototype, closest label wins — unless the top two are too close to call, in which case say nothing.

typed text→sub-word IDs→one vector→compare to every prototype→best label, or nothing

Why this design

  • No labelled dataset to build or maintain. The expensive, boring, never-finished part of supervised learning is gone. A new label is a new prototype, not a retrain.
  • It runs anywhere. A 25 MB file and a dot product. No GPU, no API key, no network, no per-call cost, and the text never leaves the device.
  • Every part is inspectable. You can print the prototypes, read the texts behind them, and explain exactly why a given sentence got a given label. That is rare and worth a lot when someone asks.

Why not just call an LLM for every input?

Often you should — if the volume is low and the network is reliable, one API call with a good prompt will beat this on day one and cost you nothing to build. Reach for this instead when you need per-call cost at zero, a tens-of-milliseconds answer, operation with no signal, or data that cannot leave the device. The honest division of labour: an LLM is an excellent labeller and an expensive classifier. Use one in Step 7 to produce the few hundred referee labels this project grades against, then ship the small model.

What it can't do

It only knows what the encoder understands. A small English model cannot read other languages, has no idea about your product's internal jargon, and — the one that will bite you — barely notices negation: "I was not charged a fee" looks almost identical to "I was charged a fee". When similarity isn't enough covers what to do about that.

What to follow along with

Every step below works on your own data. But if you want to actually run this today, with a number at the end you can compare against the rest of the world, use a public dataset instead.

OptionWhat it isWhy
BANKING77
PolyAI/banking77
13,083 real customer-service queries (10,003 train / 3,080 test) across 77 fine-grained intents. Average query is about 55 characters. CC-BY-4.0.Short, genuinely human-written, and already labelled — so you can pretend it isn't, build the whole thing blind, and still grade yourself honestly at the end.
Your own exportA CSV or Parquet file of id, text, created_at from wherever your free-text box writes to.The real thing. Everything here is written so you can swap it in, and Step 2 covers getting it out safely.
The ceiling, so you know what good looks like

Published models fine-tuned on all 10,003 labelled BANKING77 examples reach roughly 93% accuracy, with the best reported results around 94%. You are building something that uses no labels and fits in 25 MB. Knowing the gap you are accepting is the point of Step 4 — not a reason to be discouraged by it.

Part A

The text — get it, clean it, and find the floor

Before any model: a workspace, a dated snapshot of real text with the labels sealed away, an honest split, and the cheapest possible baseline so you know what the clever version has to beat.

1

Set up a Python workspace

An isolated Python environment with pinned library versions.

All the preparation — downloading, comparing, converting, clustering — uses Python libraries with no real equivalent elsewhere. Python is the workshop; only the finished files reach whatever you ship.

Do
brew install uv                      # or: pip install uv
mkdir text-classifier && cd text-classifier
uv venv --python 3.12
source .venv/bin/activate

uv pip install torch transformers huggingface-hub datasets \
               onnx onnxruntime numpy scikit-learn umap-learn pandas
uv pip install sentence-transformers model2vec   # convenience, for the bake-off
uv pip freeze > requirements-lock.txt            # freeze the exact versions
PackageJob
huggingface-hub / datasetsdownload models and public datasets
transformersload a model and its tokenizer
torchrun the model, and export it to ONNX
onnx / onnxruntimethe portable model format, and a way to run it locally
scikit-learn / umap-learnthe baseline, the clustering and the metrics
sentence-transformers / model2vecload candidate encoders in two lines during the bake-off
Watch out

Pin versions. A different torch can export an ONNX file that loads fine but produces slightly different numbers — and you would only notice much later, when the shipped build disagrees with your tests.

Check

python -c "import torch, transformers, onnxruntime, sklearn; print('ok')"

2

Get the text — and seal the labels away

A local, dated snapshot of text only, with any labels you happen to have locked in a separate file.

Following along with BANKING77
from datasets import load_dataset
import pandas as pd

ds = load_dataset("PolyAI/banking77")
train = pd.DataFrame(ds["train"])      # 10,003 rows: text, label
test  = pd.DataFrame(ds["test"])       #  3,080 rows

# The whole point: hide the labels from yourself.
train[["text"]].to_parquet("work/texts.parquet")        # what you may look at
train[["label"]].to_parquet("sealed/referee.parquet")   # only for grading
Using your own data
# Read-only, into a dated folder, scrubbing as you read.
import pandas as pd, glob
df = pd.concat(pd.read_parquet(f) for f in glob.glob("raw/2026-10/**/*.parquet", recursive=True))
df = df[df["source"] == "<the one free-text field you care about>"]
df["text"] = df["text"].map(scrub)      # see the table below
df[["id", "text"]].to_parquet("work/texts.parquet")
Remove personal information while reading, not later
FindReplace with
emails · links · phone numbers · card and account numbers · @handles<EMAIL> <URL> <PHONE> <NUMBER> <HANDLE>
"I'm Sarah", "my name is Tom"<NAME>
Why seal the labels

If you can see the answers while you cluster, you will tune until the clusters match them, and learn nothing about whether the method works on data where you don't have answers — which is the only situation you are building for. Open the sealed file in Step 4, Step 7 and the measuring section, to grade. Never to choose.

Rules for real user text
  • Read-only — never write back to the source.
  • Scrub before saving, so no unscrubbed copy ever lands on a laptop.
  • Real user text never goes in git. Add the raw folder to .gitignore before the first download, not after.
  • Check your retention and consent position before a snapshot leaves its system of record. This is the step where a project quietly becomes someone else's problem.
3

Clean it, and split it honestly

A list of unique, genuinely typed texts, each with a fair weight, split so the test set is actually hard.

  1. Drop junk — empty, under 3 characters, or fewer than 2 letters.
  2. Set aside text nobody typed. Preset replies, template phrases, autofill, dropdown options and form defaults are not free text. Answer them with a lookup table and keep them out of everything downstream — otherwise a handful of fixed phrases dominate every cluster you find.
  3. Normalise for comparing only — lowercase, straighten quotes, collapse whitespace. Keep the original wording for the model to read.
  4. Merge duplicates and near-duplicates (about 92% character similarity catches typo variants without merging genuinely different messages).
  5. Cap any one person's influence: weight = min(times_seen, distinct_users × 3). One prolific user should not define a label.
  6. Split train / validation / test by group, not at random. Cluster near-duplicates first and send a whole group to one side. Random splitting puts "my card hasn't arrived" in train and "my card still hasn't arrived" in test, and your test score becomes a measure of how well you memorised.
Field note

In the production build this recipe came from, about 60% of all rows were not typed at all — they were a handful of preset phrases from buttons. The genuinely typed remainder shrank by roughly a further quarter once near-duplicates merged. Count this before you assume you have a data volume problem.

Check

Print the 20 most frequent texts. If any of them is a phrase a human would not spontaneously type, it belongs in the lookup table, not the model.

4

Beat nothing first — the baseline with no model in it

A number the clever pipeline has to beat, produced in about fifteen lines.

Word counts with a linear model on top are a genuinely strong text classifier, and they take a minute to train. If you skip this you will have no idea whether your 25 MB encoder is earning its keep.

from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import make_pipeline
from sklearn.metrics import f1_score

baseline = make_pipeline(
    TfidfVectorizer(sublinear_tf=True, ngram_range=(1, 2), min_df=2),
    LogisticRegression(max_iter=2000, class_weight="balanced"),
)
baseline.fit(train_texts, train_labels)                  # needs labels — see below
pred = baseline.predict(test_texts)
print(f1_score(test_labels, pred, average="macro"))      # macro: every label counts equally
You now have three reference points, and they mean different things
Reference pointWhat it tells youYour number
TF-IDF + linear, trained on labelsWhat a cheap supervised model gets. The floor.fill in
A fine-tuned transformer, publishedThe ceiling with a full labelled set. On BANKING77, ~93%.~0.93
This project: no labels at allWhat you can ship without a labelling budget.by Step 19

The engineering decision this project exists to inform is the gap between the first and the third. If it turns out small, ship the small thing. If it turns out large, you have just built the strongest possible argument for a labelling budget — and the clusters from Part D are what you would label.

Watch out

This step needs labels, which Step 2 told you to seal. That is deliberate and it is not cheating: the sealed labels are a referee, used to grade, never to fit the unsupervised pipeline. If you have no labels at all, hand-label 300 texts now — that is about two hours and it is the single highest-value two hours in the project. Step 7 cannot be done without them.

What TF-IDF cannot do

It has no notion of synonyms. "card hasn't arrived" and "still waiting on my card" share almost no words and look unrelated to it. So TF-IDF tends to win on jargon-heavy text with consistent vocabulary, and lose badly where people paraphrase. Note which of those your data is — it predicts the rest of this project.

Part B

The model — four families, and how to pick between them

This is the part most write-ups skip by naming one model and moving on. Four families are worth your time, they differ by about 25× in size, and the only evidence that matters is a bake-off on your own text.

5

Shortlist: four families, not twenty models

Four candidates on the bench, one from each family, chosen for what they are rather than for a leaderboard position.

Hugging Face has thousands of embedding models. They collapse into a handful of kinds, and the kind decides far more than the individual model does. Pick one from each row and let Step 7 settle it.

FamilyA good representativeSizeVectorPick it when
Static embeddings
no transformer at all
minishlab/potion-base-8M
Model2Vec, MIT
7.6M params
~30 MB
256You need microseconds per text, or no model runtime at all. It is a lookup table per token plus an average — distilled from a real encoder. Its own card claims ~92% of MiniLM's benchmark quality, orders of magnitude faster.
Small bi-encoder
the default
sentence-transformers/all-MiniLM-L6-v2
6 layers, Apache-2.0
or BAAI/bge-small-en-v1.5, MIT
23M / 33M
~23 MB int8
384Almost always. Small enough for a phone, trained specifically for sentence similarity, and the quality difference to the next row up is usually smaller than you expect.
Base bi-encodersentence-transformers/all-mpnet-base-v2
or BAAI/bge-base-en-v1.5
109M
~110 MB int8
768The bake-off says the gain is real and you have the budget. Roughly 4–5× the compute per text and double the vector, so double the index too.
Cross-encoder / NLI
reads text and label together
MoritzLaurer/deberta-v3-base-zeroshot-v2.0
184M, MIT
184Mn/aYou need the hard cases right and can afford one model pass per candidate label. No prototypes, no clustering — you hand it a sentence and a hypothesis and it scores entailment. Far too slow for every input; exactly right for the 10% that are ties. See When similarity isn't enough.
Why a bi-encoder is the default

A bi-encoder embeds each text once, independently. That is what makes everything downstream cheap: you encode your whole corpus once, you encode a new text once, and comparing it to 200 prototypes is 200 dot products — microseconds. A cross-encoder has to re-read the text alongside every label you are considering, so cost scales with the number of labels. The whole architecture of this project depends on choosing the cheap one for the common path.

The raw-BERT trap

Do not reach for bert-base-uncased, or any plain language model, and average its outputs. It was never trained to put similar sentences near each other, and the vectors you get are surprisingly bad at similarity. Every model in the table above has been through contrastive training on sentence pairs specifically for this — MiniLM on over a billion pairs. That training, not the architecture, is what you are buying.

6

Read each model card for the four things that will break you

For every candidate, four facts written down before you run anything.

These are not performance details. Get any of them wrong and the model silently returns worse vectors, with no error anywhere.

1 · Pooling — how per-token vectors become one sentence vector

The model emits one vector per sub-word. You need one per sentence. Two conventions: mean (average them) or CLS (take the special first token's vector). It is not your choice — it is how the model was trained, and the wrong one quietly degrades everything. Read it from 1_Pooling/config.json, whichever flag is true:

all-MiniLM-L6-v2      pooling_mode_mean_tokens: true   → mean
bge-small-en-v1.5     pooling_mode_cls_token:   true   → CLS
e5-small-v2           pooling_mode_mean_tokens: true   → mean

When averaging, exclude the padding tokens (multiply by the attention mask first) or short sentences get diluted by a long batch. Then normalise to length 1, so comparing two sentences is a single dot product.

2 · Required prefixes — some models demand them

The E5 family is trained with literal prefixes on every input and degrades without them. BGE's v1.5 models accept an optional retrieval instruction but are documented as fine without it. MiniLM wants nothing.

e5-small-v2     tok("query: my card hasn't arrived")    # mandatory, even for similarity
bge-small-v1.5  tok("my card hasn't arrived")           # instruction optional for v1.5
MiniLM-L6-v2    tok("my card hasn't arrived")           # nothing

If you pick a prefixed model, the prefix becomes part of the contract: it has to be applied identically in Python and in whatever ships, forever. That is a real reason to prefer a model that needs none.

3 · Maximum tokens — where your text gets cut off
ModelMax tokensConsequence
all-MiniLM-L6-v2256Trained at 128; anything past 256 sub-words is silently truncated.
bge-small-en-v1.5 · e5-small-v2512Comfortable for a paragraph.

Check the 99th percentile length of your own texts in sub-words, not characters. For short free text, 128 is usually generous — and a smaller cap makes every batch faster.

4 · Licence and vector width

Licence is at the top of the model page, and you need to look before you ship commercially. MiniLM is Apache-2.0; BGE, E5, Model2Vec and the DeBERTa zero-shot models above are MIT. Plenty of strong models are not, and some good ones are explicitly non-commercial.

Vector width is your index size and your per-comparison cost, both linear in it. 384 versus 768 is double the index file and double the scoring work, forever, for a quality difference you have not measured yet.

On the ready-made ONNX files in model repos

Many repos ship their own onnx/ folder. Look at the output signature before using one: they usually emit last_hidden_state [batch, tokens, dim] — one vector per sub-word — which leaves you to implement pooling and normalising in whatever language you ship, where it can silently drift from Python. Steps 9–10 build a file that outputs one finished sentence vector instead, which is the single most effective thing you can do to prevent cross-language bugs.

7

Bake them off on your own text

One table, same protocol for every candidate, measured on your data rather than someone else's benchmark.

Leaderboard rankings are averages over tasks that are not yours. The differences between the models in Step 5 on your text are routinely smaller than the differences between runs — which is itself the most useful thing you can learn, because it means you should take the small one.

The protocol — and it must be identical for all of them
  1. One referee label set, made without any encoder. Human-labelled, or LLM-labelled and spot-checked. A few hundred texts is enough.
  2. One split, decided once. Same texts in train and test for every candidate.
  3. For each candidate: embed everything, build one prototype per referee label from the train half, score the test half by nearest prototype, report macro-F1.
  4. Five random splits, report mean ± spread.
  5. Also record what it costs: file size after int8, and p95 encode time for one text on the slowest hardware you intend to support.
from sentence_transformers import SentenceTransformer
import numpy as np, time

CANDIDATES = [
    "minishlab/potion-base-8M",                 # static
    "sentence-transformers/all-MiniLM-L6-v2",    # small, mean pooling
    "BAAI/bge-small-en-v1.5",                   # small, CLS pooling
    "sentence-transformers/all-mpnet-base-v2",   # base
]

for name in CANDIDATES:
    m = SentenceTransformer(name)                # handles pooling per its card
    E = m.encode(texts, normalize_embeddings=True, batch_size=64)
    scores = [proto_macro_f1(E, referee_labels, seed=s) for s in range(5)]
    t = time.perf_counter(); m.encode([texts[0]]); ms = (time.perf_counter()-t)*1000
    print(f"{name:48s} {np.mean(scores):.3f} ± {np.std(scores):.3f}   {E.shape[1]:4d}d  {ms:6.1f} ms")

sentence-transformers applies each model's own pooling, so use it for the bake-off even though Part C exports ONNX by hand. Comparing candidates and shipping one are different jobs.

Record your own — this table is the deliverable of Part B
CandidateMacro-F1 ± spreadVectorSize (int8)p95 encode
TF-IDF + linear (Step 4)————
potion-base-8M—256——
all-MiniLM-L6-v2—384——
bge-small-en-v1.5—384——
all-mpnet-base-v2—768——
The mistake that invalidates the whole table

Never grade a model against labels that a model produced. If you cluster with encoder A, name the clusters, and then score encoder B against those names, you are measuring how much B agrees with A — and A wins by construction. In the build this recipe came from, doing exactly that inflated one model's apparent score by about 0.12, which was larger than every real difference in the table. The referee labels must come from outside every candidate.

Watch out

Split the test set into typical and unusual once, using one reference model, and reuse that division for every candidate. Texts far from their group's centre are the closest thing you have to "phrasing we have never seen", and almost all of the difference between models shows up there. If each candidate defines its own split, they are being graded on different exams.

8

Decide — and write down why

One model chosen, with the reasoning recorded somewhere a future colleague will find it.

The rule that holds up: take the smallest candidate whose score is inside the noise band of the best one. If mpnet-base scores 0.81 ± 0.02 and MiniLM scores 0.79 ± 0.02, those are the same number, and one of them is a fifth of the size.

The decision record — four lines in the repo, next to the model file
Encoder: sentence-transformers/all-MiniLM-L6-v2  (Apache-2.0)
Chosen:  2026-10-05, from the Step 7 bake-off on 400 referee-labelled texts.
Because: within one standard deviation of all-mpnet-base-v2 (the best scorer)
         at 1/5 the parameters, 1/2 the vector width and no required prefix.
         Rejected potion-base-8M: clearly below the band, though 40x faster.
Revisit: if p99 text length exceeds 256 sub-words (its cap), or macro-F1 on
         the monthly fresh sample drops more than 0.05 below this baseline.
Why bother writing it down

Six months on, someone will ask why this isn't using whatever is currently fashionable. Without the record, the honest answer is "no idea", and you will re-run the whole bake-off to find out. With it, the answer is a diff: re-run Step 7 with one more row. The Revisit line is the valuable part — a decision with no expiry condition is a decision nobody can ever safely revisit.

Watch out

Write down the vector width too, not just the name. It is baked into the index file, the scoring code and every stored embedding. Changing encoders later is a migration, not a config change — Step 17 puts the width in the index file precisely so a mismatch fails loudly instead of returning confident nonsense.

Part C

The vectors — make the chosen model portable

Convert the model you picked into a single file that takes IDs and returns one finished sentence vector, shrink it, and prove the conversion did not change its answers. Done once, reused for everything after.

9

Convert to ONNX, with pooling built in

One file that takes sub-word IDs and returns one normalised sentence vector.

ONNX is a portable model format that runs almost anywhere via onnxruntime — browsers, phones, servers, embedded. Wrap the model in a small class that does the pooling and normalising inside the graph, then export the whole thing, so nothing downstream has to reimplement that arithmetic.

import torch
from transformers import AutoModel, AutoTokenizer

NAME = "sentence-transformers/all-MiniLM-L6-v2"   # whatever Step 8 chose
POOLING = "mean"                                   # from Step 6 — not a guess

tok   = AutoTokenizer.from_pretrained(NAME)
model = AutoModel.from_pretrained(NAME)

class SentenceEncoder(torch.nn.Module):
    def __init__(self, m, pooling): super().__init__(); self.m, self.pooling = m, pooling
    def forward(self, input_ids, attention_mask):
        words = self.m(input_ids=input_ids, attention_mask=attention_mask).last_hidden_state
        if self.pooling == "cls":
            sent = words[:, 0]                                       # first token
        else:
            mask = attention_mask.unsqueeze(-1).float()
            sent = (words * mask).sum(1) / mask.sum(1).clamp(min=1e-9)  # mean, skip padding
        return torch.nn.functional.normalize(sent, dim=1)            # length 1

sample = tok("a sample sentence", return_tensors="pt",
             padding="max_length", truncation=True, max_length=128)

torch.onnx.export(
    SentenceEncoder(model, POOLING).eval(),
    (sample["input_ids"], sample["attention_mask"]),
    "model.onnx",
    input_names=["input_ids", "attention_mask"],
    output_names=["embedding"],
    dynamic_axes={"input_ids": {0: "batch", 1: "seq"},
                  "attention_mask": {0: "batch", 1: "seq"},
                  "embedding": {0: "batch"}},
    opset_version=17,
    dynamo=False,
)
The settings that matter
  • sample — any sentence; the exporter traces the arithmetic through it
  • dynamic_axes — without it the file only accepts text exactly as long as the sample
  • dynamo=False — at the time of writing the newer exporter's output breaks the int8 step below
  • POOLING — read from the model card in Step 6. There is no error if it is wrong, only worse answers.
Check

Open the file at netron.app (also available offline as an app). The output should be a single embedding [batch, dim] — not last_hidden_state with a token axis.

If your model needs a prefix

A required query: prefix cannot live in the graph — it is applied to text, before tokenizing. Put it in the settings file in Step 11 so the shipped tokenizer applies it, and add a parity case for it in Step 23. This is the hidden cost of choosing a prefixed model.

10

Shrink it to 8-bit

About 4× smaller, with no measurable change in which label wins.

Weights are stored as 32-bit decimals — 4 bytes each. Rounding each to a 1-byte integer is called quantisation, and for an encoder like this it is close to free.

from onnxruntime.quantization import quantize_dynamic, QuantType
quantize_dynamic("model.onnx", "model_int8.onnx", weight_type=QuantType.QInt8)
# then delete model.onnx — only the small one ships
Field note

On a real corpus of a few thousand texts: 90 MB → 22.9 MB, with accuracy of 0.651 before and 0.655 after — well inside run-to-run noise. The same held for a 109M-parameter model. Shrinking is effectively free; file size comes from model capacity, not from quantisation, so pick the small model in Step 8 rather than hoping to compress a big one.

Check

Don't take the above on trust for your model — Step 12 measures it. The number that matters is not vector similarity but whether any decision changed.

11

Save the vocabulary and settings

Everything whatever-you-ship needs to turn text into exactly the IDs the model expects.

The model reads IDs, not text. The target has to do that conversion itself, so it needs the vocabulary (one sub-word per line, line number = ID) and every rule that was applied on the way.

vocab   = tok.get_vocab()
ordered = [t for t, _ in sorted(vocab.items(), key=lambda kv: kv[1])]
open("vocab.txt", "w").write("\n".join(ordered))

meta = {   # read every value off the tokenizer object — never type them by hand
  "model":   NAME,
  "pooling": POOLING, "dims": 384, "max_len": 128,
  "prefix":  "",      # e.g. "query: " for E5 — see Step 6
  "special_tokens": {"cls_id": tok.cls_token_id, "sep_id": tok.sep_token_id,
                     "pad_id": tok.pad_token_id, "unk_id": tok.unk_token_id},
  "tokenizer": {"lowercase": tok.do_lower_case, "strip_accents": True,
                "continuing_subword_prefix": "##"},
}
What a BERT-style tokenizer actually does
"my card is missing"  →  [CLS] my card is missing [SEP]   # markers wrap every text
"overdraft"           →  over ##draft                     # ## = glued to the previous piece
"Café"                →  cafe                             # lowercased, accent stripped
"thanks 🙏"           →  thanks [UNK]                      # unknown symbol
Watch out

If one rule differs between Python and the target — say accents are not stripped — the same sentence becomes different IDs and a different answer, with no error anywhere. Generate this file from the tokenizer object; do not write it from memory or copy it from another project.

12

Check the conversion, stage by stage

Proof that the converted file answers like the original — with each stage checked separately.

Encode the same sentences with the original model and with your files, and compare with a dot product (1.0 = identical).

StageExpectedIf not
Original → ONNX (Step 9)1.0000a real bug: pooling, the mask, or export settings
ONNX → int8 (Step 10)≈ 0.96–0.99normal rounding — not a bug

Then the check that actually matters: run a few hundred real texts through both and confirm which label wins does not change. Vectors drift by a few percent; decisions should not.

Watch out

Check the stages separately. Looking only at the final int8 file, 0.97 looks like a bug and sends you hunting for one. Looking only at end-to-end accuracy, you would never discover that the export itself was exact and the drift was entirely rounding.

13

Embed the whole corpus — with the file that ships

One vector per text, saved alongside a matching list of ids.

import onnxruntime as ort, numpy as np
sess = ort.InferenceSession("model_int8.onnx")
vecs = []
for i in range(0, len(texts), 32):
    b = tok(texts[i:i+32], padding=True, truncation=True, max_length=128, return_tensors="np")
    vecs.append(sess.run(None, {"input_ids": b["input_ids"],
                                "attention_mask": b["attention_mask"]})[0])
E = np.vstack(vecs)                         # (n_texts, dims)
np.save("embeddings.npy", E)                # + save the ids in the same order

Similarity is now a dot product. Absolute values stay modest even for obvious matches — what matters is the gap between related and unrelated pairs:

0.71   "my card hasn't arrived"   vs  "still waiting on my new card"
0.15   "my card hasn't arrived"   vs  "how do I change my address"
Watch out

Embed with the int8 file that actually ships, not the original model. They differ by a few percent, and everything downstream — clusters, prototypes, thresholds — should live in the same vector space the target will. And assert that ids and vectors still line up every time you load them; if they slip out of order, every later step labels the wrong text and nothing complains.

Part D

The labels — derive them from the data instead of guessing

Let clustering propose the groups that are actually in your text, read them, name them, and pack each one's average meaning into a small file. This is the part where you learn what your users are really saying.

14

Cluster — find the natural groups, then freeze them

Groups of "the same kind of thing", found without any labels, and reproducible forever after.

import umap, joblib
from sklearn.mixture import GaussianMixture

reducer = umap.UMAP(n_components=10, n_neighbors=15, min_dist=0.0,
                    metric="cosine", random_state=42).fit(E)       # 384 → 10 numbers
gmm     = GaussianMixture(n_components=40, covariance_type="full",
                          random_state=42).fit(reducer.embedding_)
cluster = gmm.predict(reducer.embedding_)

joblib.dump(reducer, "frozen/v1/umap.joblib")       # freeze — never refit v1
joblib.dump(gmm,     "frozen/v1/gmm.joblib")

Why squash to 10 dimensions first? In 384 dimensions, distances between points concentrate — almost everything looks roughly equally far from everything else — and clustering struggles. Reducing first gives it a space with real structure in it.

How many groups? Do not trust the textbook criterion (BIC). On text data it tends to keep asking for more groups indefinitely. Ask the question you actually care about instead: if each group became a label, how often would nearest-prototype get it right? Sweep the count, score each against your referee labels, and take the peak.

Record your own
Number of groupsHeld-out macro-F1
Half what you expect—
Roughly what you expect—
BIC's choice—
Double what you expect—

The curve is usually flat-topped and then falls away. Take the smallest count on the plateau: fewer labels is less to name, map and maintain.

Field note

On a real corpus the peak sat at 40 groups (0.768 held-out), against 0.733 at 28 and 0.756 at the 52 that BIC asked for. BIC was both wrong and expensive — those extra twelve groups would each have needed naming and mapping.

Why freeze the fitted models

A fixed seed makes a run repeatable, but the groups depend on every text that went in. Removing 1% of the corpus reshuffled about a quarter of the group assignments in one measured case. New data arrives continuously, so refitting silently renames everything — and every label name, mapping and threshold you built on "cluster 12" now points somewhere else. Save the fitted objects, reuse them for new text, and create v2 deliberately when you actually want new groups.

15

Name the clusters — they become your labels

A readable name and a written description for every group worth keeping.

  1. Pull keywords — words common inside the group and rare outside it (c-TF-IDF).
  2. Draft a name from the top keywords, or have an LLM draft one from twenty sampled texts.
  3. Test each group's coherence: build its prototype from half its texts and check how often the other half lands back on it. Below about 0.6, the group is two things wearing one name — split it or drop it.
  4. Reject invented groups: if you added any synthetic or templated text, drop groups that are mostly synthetic. Those are labels for things no real person said.
  5. Read the texts and rename by hand. Twenty per group. There is no shortcut and this is the step that makes the project good.
Keyword names lie — always read the texts
Auto-generated nameWhat the group actually contained
great_good_thanks"didn't feel this was good enough" — the opposite sentiment
transfer_pendingmostly fee disputes that happened to mention a transfer
nicht_qvjxknon-English messages, plus keyboard mashing

In one real build, 14 of 37 auto-generated names were actively misleading — not vague, but pointing at the wrong thing. Naming from keywords alone would have mapped a seventh of the traffic to the wrong action.

Check

Export a sheet of name · description · top keywords · size · coherence, grouped into themes, and have someone who knows the domain but not the code read it. If they cannot tell two labels apart, neither can the model.

16

Give each label several prototypes

Cover the genuinely different ways people say the same thing.

A label's prototype is the average of its texts' vectors. But one label often holds several distinct phrasings — "card not delivered" and "tracking says delivered but nothing arrived" — and a single average of both lands between them, close to neither. So cluster again inside each label and keep one average per sub-group. A label's score is then the best of its prototypes.

from sklearn.cluster import KMeans
X  = E[label_texts]                          # this label's REAL texts only
m  = min(5, len(X) // 4)                     # at least 4 texts behind every vector
km = KMeans(n_clusters=m, random_state=42).fit(X)
prototypes = [normalise(X[km.labels_ == j].mean(0)) for j in range(m)]
Field note

Measured on a real corpus: 1 prototype per label → 0.742; 3 → 0.790; 5 → 0.803; 12 → 0.813. A flat 5 was the knee of the curve. Letting each label pick its own count by silhouette score did worse (0.792) than the flat 5 — the extra cleverness cost accuracy. Start at 5.

The one that catches people: more variety in a single average does not help

Adding paraphrases into one prototype made it measurably worse, because averaging more distinct phrasings pulls the result towards the middle of all of them. Multiple prototypes is the fix; a richer single prototype is not.

A privacy property worth knowing

An average of many sentences cannot be turned back into any one of them. A single sentence's vector partly can. The len(X) // 4 floor — at least four texts behind every shipped vector — is what keeps individual users' words out of a file you distribute. Keep it even when a label is small; drop the label instead.

17

Pack everything into one small index file

Labels, prototypes and decision rules in one file the target reads at startup.

{
  "labels":           ["card_not_arrived", "fee_dispute", …, "none"],
  "proto_label":      [0,0,0,0,0, 1,1,1,1,1, …],   // which label each prototype belongs to
  "prototypes_int8":  "<base64>",                 // the vectors, 1 byte per number
  "prototype_scales": [0.153, …],
  "min_margin":       0.03,                        // tie → say nothing
  "tau":              0.05,                        // only shapes the confidence %
  "encoder":          {"model": "all-MiniLM-L6-v2", "dims": 384,
                       "pooling": "mean", "prefix": ""}
}
scale = abs(P).max(axis=1, keepdims=True)
Q     = round(P / scale * 127).astype(int8)       # target decodes: byte / 127 * scale
  • none has zero prototypes — it can never win on similarity. It is only where "say nothing" goes, so that downstream code has one label vocabulary rather than a label plus a null case.
  • encoder lets the target refuse an index built for a different model. A mismatched model-and-index pair produces confident nonsense, not an error, so this field is load-bearing.
Field note

185 prototypes: 1.55 MB as plain JSON floats, 97 KB as int8 with per-row scales. Zero decisions changed across 1,500 real texts. Quantising the index is as free as quantising the model.

Part E

The decision — from scores to an answer, or to silence

A vector and a list of prototypes are not yet a decision. This is the part that turns them into one, and the part that decides to keep quiet — which matters more than any accuracy number.

18

Score, with the rules in a fixed order

From a sentence vector to a label, by a sequence you can read off and explain.

  1. Known phrase? Exact match in the lookup table from Step 3 → that label, no model involved.
  2. Junk? Too short, or no letters → none.
  3. Score — dot product against every prototype; keep the best per label.
  4. Too close to call? Top two labels within min_margin → none.
  5. Confidence — softmax over the per-label scores. For display only; it cannot change the winner.
const raw  = protos.map(row => dot(vec, row));
const best = new Array(labels.length).fill(-1);
protoLabel.forEach((l, i) => { if (raw[i] > best[l]) best[l] = raw[i]; });

const sorted = [...best].sort((a, b) => b - a);
if (sorted[0] - sorted[1] < minMargin) return { label: "none", source: "margin" };

const probs = softmax(best.map(x => x / tau));
const top   = argmax(probs);
return { label: labels[top], confidence: probs[top], margin: sorted[0] - sorted[1] };
Watch out

Decode the int8 prototypes once at startup, not per call. Bytes from a base64 string arrive unsigned, so values above 127 must become b − 256. And put the margin test on the cosine scores, not on the softmax percentages — otherwise changing tau, which is supposed to be cosmetic, silently moves your abstention rate.

19

Decide when to stay silent

One tunable dial, chosen on evidence, trading how often you answer against how often you are right.

Saying nothing is almost always better than a confident wrong answer. Two ways to decide, and they are not equivalent:

RuleAsksCatches
Confidence floor"is the best match good enough in absolute terms?"text unlike anything you have seen
Margin"is the best clearly ahead of the runner-up?"genuine ambiguity — text that really does say two things
Field note

Measured on real traffic, holding accuracy at 91% for the answers given: a confidence floor answered 48% of inputs; a margin rule answered 74%. The margin targets coin-flips rather than unusual phrasing, so it declines far less traffic for the same reliability. Both were tried; the margin won clearly.

Record your own — this is a product decision, not a technical one
min_margin% answeredAccuracy when answering
0.02——
0.03——
0.05——
0.08——

Sweep it, put the table in front of whoever owns the feature, and let them pick the row. "Right 88% of the time on 80% of messages" versus "right 95% of the time on 57%" is a question about what a wrong answer costs your users, and that is not an engineering judgment.

Two things about the confidence number
  • tau only stretches or flattens the percentages. It can never change which label wins. Treat it as display formatting.
  • tau is model-specific. A wider model packs its scores into a narrower band, so the value that makes percentages look sensible differs per encoder — measured at roughly 0.033 for one base model where a small one needed 0.079. Re-tune it if you change encoders, and never port it between projects.
  • The margin cannot tell a weak winner from a strong one. If "nothing here resembles anything we know" is a case you need to catch, add a floor on cosine as well (e.g. best ≥ 0.50) — on the cosine, not the percentage.

Part F

Shipping — make the target agree with Python

Five files, a small amount of logic reimplemented in the target language, and a fixture that fails loudly the moment the two disagree. Whether the target is a phone, a browser, a Go service or an embedded device, the shape is the same.

20

Get the tokenizer into the target language

Text → the exact same IDs Python produces. Exactly, not approximately.

Look for a real library first. Hugging Face's tokenizers has Rust bindings and a WASM build; there are solid BERT tokenizer ports for most languages. If one exists for your target, use it and load tokenizer.json directly — that is strictly better than what follows. Hand-porting is the fallback, and it is only reasonable because BERT's WordPiece is small:

  1. Normalise — strip control characters, unify whitespace, put spaces around CJK characters, decompose accents and drop the combining marks, lowercase.
  2. Pre-split — on whitespace, then make every punctuation character its own piece.
  3. WordPiece — for each word, repeatedly take the longest vocabulary match; every piece after the first carries ##. No match, or a word over 100 characters → [UNK].
function wordPiece(word: string, vocab: Map<string, number>): string[] {
  const out: string[] = [];
  let start = 0;
  while (start < word.length) {
    let end = word.length, piece = "";
    while (start < end) {
      const cand = (start > 0 ? "##" : "") + word.slice(start, end);
      if (vocab.has(cand)) { piece = cand; break; }
      end--;
    }
    if (!piece) return ["[UNK]"];
    out.push(piece);
    start = end;
  }
  return out;
}
// finally: prefix + [CLS] + ids + [SEP], truncated to max_len
The five files the target needs
FileFromTypical size
model_int8.onnxStep 10~23 MB
vocab.json — {"text": "<vocab.txt>"}Step 11~270 KB
encoder_meta.jsonStep 111 KB
index.jsonStep 17~100 KB
lookup.json (known phrases)Step 3small

Wrapping the vocabulary as JSON rather than shipping a .txt is worth it: bundlers import .json natively, in both the app and its tests, with no asset-loader configuration.

Check — do not skip this one

Run several hundred real and deliberately awkward strings through both tokenizers and require identical ID arrays: accented text, emoji, full-width characters, curly apostrophes, mixed scripts, strings of punctuation, the empty string, and something longer than max_len. This is where the large majority of cross-language bugs live, and they are all silent.

21

Wrap it in one function

A single call for the rest of the codebase, with the expensive setup done once.

const r = await classify("my card still hasn't arrived");
// { label, confidence, margin, source, top3 }
  • Load once — build the model session on first call and share it. Concurrent callers must await the same in-flight promise, not start their own load.
  • Do not cache failure — clear the cached promise on error, or one transient bad load breaks every subsequent call for the process lifetime.
  • Preload — start loading when the input box opens, not when the user submits, so the result is not a spinner.
  • Refuse mismatches — if the index's dims, pooling or model disagree with the loaded encoder, throw at startup. This is the Step 17 field earning its place.
  • Never block the user — on any error, return none and let them carry on. A classifier is an enhancement; it should not be able to break the form it is attached to.
22

Keep the label and the action apart

A separate, explicit table from label to whatever the product does about it.

The model predicts a label — what the person said. A separate map decides the action — what happens next. Keep them apart: several labels often share one action, some labels have no useful action at all, and the product will change its mind about actions far more often than the model changes its mind about labels.

export const LABEL_ACTION: Record<string, string | null> = {
  card_not_arrived: "card_tracking_flow",
  card_damaged:     "card_tracking_flow",   // several labels → one action
  not_english:      null,                    // understood, nothing to offer
  none:             null,
};
export const actionFor = (l: string) => LABEL_ACTION[l] ?? null;  // unknown → null, never throw
Check

Add a test that fails if the map and the shipped index.json disagree in either direction — a label with no entry, and an entry for a label that no longer exists. Label names change whenever you build a v2 in Part D, and this is the test that catches it before your users do.

23

Prove the target gives Python's answers

Identical decisions at every level, from your laptop to the real hardware.

Generate a fixture in Python: a list of inputs with the expected label, source, margin and confidence — deliberately covering known phrases, junk, ordinary cases, near-ties, text over the length cap, and non-English. Then check it at three levels, because each proves something the others cannot:

LevelWhat runsWhat it proves
Off-device, target languageyour port + onnxruntime for the desktopthe tokenizer and scorer ports are correct, end to end
Unit testthe scorer alone, fed the fixture's own vectorsthe decision arithmetic, inside your normal test suite
Real targetthe actual runtime build on the actual devicethat the hardware you ship to agrees
The third level is not optional

A desktop ONNX Runtime build is not a mobile or WASM one: different kernels, different threading, sometimes different accumulation order. Passing off-device proves nothing whatsoever about the device. Run the fixture on the target — a hidden debug screen that runs it on demand is enough — and require a real tolerance, e.g. 1e-4 on confidence and exact equality on the label.

You are done when

The same 300 sentences produce the same 300 labels in Python, in your port, and on the device; the margin dial is set to a row somebody chose deliberately; and the decision record from Step 8 is in the repo next to the model file.

Reference

Measuring, limits, and what goes wrong

The parts that are not a single step but shape every result you get.

Measuring honestly

Six rules. Each of them exists because breaking it produces a number that is too good and completely useless.

  • Hold data out.Build prototypes from one half, score the other. Scoring on the texts you built the prototypes from proves only that averaging works.
  • Split the test set into typical and unusual.Texts near their group's centre are easy by construction. Texts far from it are the closest proxy you have for "phrasing we have never seen", which is what production is. Report both; most of the difference between any two methods lives in the second number.
  • Compare models only against labels no model made.If the clusters came from encoder A, scoring encoder B against them rewards A for agreeing with itself. Measured inflation from this mistake in one real build: about 0.12 — larger than every genuine difference being measured.
  • Give every candidate the same test set.Decide the typical/unusual division once, with one reference model, and reuse it. Otherwise each model is graded on a different exam.
  • Report the spread.Run five random splits and publish mean ± standard deviation. A gap smaller than the spread is not a result, and treating it as one is how teams end up shipping a 110 MB model for nothing.
  • Test on genuinely new data.Fetch text written after your snapshot and drop anything already seen. In one measured case, 79.6% of genuinely new texts were covered by the existing labels against 87.6% of training-era text — an eight-point generalisation cost that no internal split would ever have shown.

When similarity isn't enough

Nearest-prototype on a frozen encoder has a hard ceiling, and it is worth knowing exactly where it is before you spend a week trying to tune past it.

The limit you will hit first: negation

Small encoders barely represent negation. "I was not charged a fee" and "I was charged a fee" land close together, because almost every word is shared and the model's training signal was topical similarity, not logical polarity. In one measured set of twelve deliberately opposite pairs, eight landed on the same label. This is not fixable with more prototypes or a keyword rule, and it matters most in exactly the domains where people state problems using negation — "not working", "no confirmation", "didn't arrive".

Three things that actually help, in increasing cost

OptionWhat it needsWhen it is the right answer
A trained head on the same vectorsA few hundred labels per class. Logistic regression on the frozen embeddings — minutes to train, kilobytes to ship, same encoder.You have labels, or Step 4 showed the gap is worth getting them. This is the cheapest real upgrade and it keeps every other part of the pipeline unchanged.
A cross-encoder or NLI model on the hard casesOne extra model, and one pass per candidate label. Reads the text and the label description together, so word order and negation are visible to it.You need the ambiguous ones right. Use the prototypes to answer the easy majority and route only the within-margin cases to it — the two-stage pattern, which keeps the average cost near the cheap path's.
Fine-tune the encoderA labelled set, a training loop, and a new parity exercise for every retrain.Last. It is the biggest commitment and the gain over a trained head on frozen vectors is usually much smaller than the effort difference. On BANKING77 it is roughly the difference between the mid-80s and ~93% — real, but measure it before assuming you need it.

The other hard limit: small English models are English-only. Other languages cluster together by being foreign rather than by topic, which is at least easy to detect: leave that group unmapped, or translate before encoding. Do not map it to a real label.

Traps

  • Grading a model against labels another model made.The single most expensive mistake available in this project, because the resulting number looks plausible. Referee labels must come from outside every candidate.
  • Keyword-derived cluster names.Misleading often enough to matter — 14 of 37 in one build. Read the texts before naming or mapping anything.
  • Trusting BIC for the cluster count.It keeps asking for more. Choose the count by whether the resulting labels are useful.
  • Refitting the clustering on new data.Silently renumbers every group, invalidating every name, mapping and threshold built on the old numbering. Freeze, and version deliberately.
  • Mixing synthetic text into prototypes.It can mint labels for things no user ever said. Track the real/synthetic share per cluster, reject the mostly-synthetic ones, and keep synthetic text out of the prototype averages entirely.
  • Expecting a richer single prototype to help.Averaging more distinct phrasings pulls the result to the middle of all of them and measured worse. Use more prototypes instead.
  • Embedding with the unquantised model.Everything downstream then lives in a slightly different vector space from the one you ship. Use the int8 file from Step 10 for Step 13 onwards.
  • Putting the margin on the softmax percentage.Makes tau, which is meant to be cosmetic, silently change your abstention rate. Margin goes on the cosine.
  • Assuming a desktop runtime predicts the device.Different ONNX Runtime builds, different kernels. Run the parity fixture on the real target.
  • Bundlers that exclude test directories.Several — React Native's Metro among them — refuse to resolve anything under __tests__/ into the app bundle, so app code cannot import a fixture that lives there. Keep test-only files in the test folder and a separate copy of anything the app itself reads in assets.
  • Development builds writing to production.Check where your dev build sends user input before you test with made-up sentences. Otherwise every test string lands in the real corpus you will cluster next month.
  • The scorer existing in two places.A research implementation and a shipped one will drift. Keep a parity test that runs the same fixture through both, in both test suites.

Where this maps onto the track

This project is the practical end of Classifying with small models. Each module explains a piece of it properly; the project is all of them as one thing you run. It restates what it needs as it goes, so either order works.

ModuleSteps that use it
A · Framing — abstention, rules before modelsOverview · 3 · 18 · 19
B · Labels and representation — TF-IDF, choosing an encoder, model cards and bake-offs4 · 5 · 6 · 7 · 8
C · Making a decision — prototypes, classifier heads, softmax and temperature, thresholds16 · 18 · 19 · When similarity isn't enough
D · Knowing it works — precision and recall, macro-F1, slicing, splits, where labels come from3 · 7 · Measuring honestly
E · Exploring data by hand — dimensionality reduction, clustering, Gaussian mixtures, labelling workflow14 · 15
F · Beyond bi-encoders, and shipping — NLI and cross-encoders, retrieve-then-rerank, ONNX and int89 · 10 · 11 · 12 · 20 · 23 · When similarity isn't enough

Glossary

Encoder
A model that turns a piece of text into a fixed-length list of numbers.
Embedding / vector
That list of numbers. Similar meanings get similar vectors.
Bi-encoder
Embeds each text independently, so comparisons are cheap. The default here.
Cross-encoder
Reads two texts together and scores the pair. Much better at hard cases, much more expensive.
NLI
Natural language inference: does this text entail this statement? Any classification can be posed this way, which is how zero-shot classifiers work.
Static embeddings
A fixed vector per token, averaged — no transformer at runtime. Far faster, somewhat weaker.
Cosine similarity
How alike two vectors are, from −1 to 1. For length-1 vectors it is just the dot product.
Token / sub-word
A chunk of text the model knows — a word, or part of one (##draft).
Pooling
Combining per-token vectors into one text vector: mean, or the CLS token.
TF-IDF
Word counts weighted so rare words matter more. No notion of synonyms, and a strong baseline anyway.
Macro-F1
F1 computed per label and then averaged, so a rare label counts as much as a common one. The right default for many labels.
Referee set
Labels produced independently of every model being compared, used only to grade.
Bake-off
Running every candidate through one identical protocol on your own data.
ONNX
A portable model file format that runs almost anywhere.
Quantisation
Storing numbers in fewer bits (here 8) to shrink files.
UMAP
Squashes many dimensions down to a few while keeping neighbours close.
GMM
Gaussian mixture model — finds blob-shaped groups, with a soft assignment per point.
K-means
Splits points into k groups around k centres. Used here for sub-prototypes.
Prototype
The average vector of a label's texts; new text is compared against it.
Margin
The gap between the best and second-best label score. The abstention dial.
Abstention
Returning no label on purpose. Usually better than a confident wrong answer.
Softmax / tau
Turns scores into percentages; tau controls how sharp they are. Display only.
Fixture
Inputs with known expected outputs, used to prove two implementations agree.
Freezing
Saving a fitted model and never refitting it, so its numbering stays stable.