Why another error rate
Plain WER treats every mismatch as the same mistake. For Indic languages and domain-heavy transcription that hides the story: an agglutinated word split in two, a date written with different separators, and a misrecognized legal term are three very different failures. SCRIBE decomposes tokens into base classes and an optional domain class, aligns them with a category-aware algorithm, and reports one error rate per class over a shared denominator.
| Category | Label | What lands here |
|---|---|---|
| LEXICAL | ER_LEX | General words, Indic and English |
| NUMERAL | ER_NUM | Numbers, dates, times — 302, 22.05.2023, 10:30 |
| PUNCT | ER_PUNCT | Punctuation marks |
| LEGAL / MEDICAL / TECH / custom | ER_DOMAIN | Domain terminology, shielded from incorrect splitting |
WER_SCRIBE is the composite: all category errors over the combined reference-token denominator. CER_SCRIBE is a character error rate computed on normalized token streams — format variants cost nothing, and it needs no sandhi machinery to be robust to agglutination.