SCRIBE

the scribe-eval Python package · import scribe

Diagnostic evaluation for Indic & domain-specific ASR — sandhi-aware alignment, domain-shielded tokenization, and error metrics that say what went wrong, not just how much.

Why another error rate

Plain WER treats every mismatch as the same mistake. For Indic languages and domain-heavy transcription that hides the story: an agglutinated word split in two, a date written with different separators, and a misrecognized legal term are three very different failures. SCRIBE decomposes tokens into base classes and an optional domain class, aligns them with a category-aware algorithm, and reports one error rate per class over a shared denominator.

CategoryLabelWhat lands here
LEXICALER_LEXGeneral words, Indic and English
NUMERALER_NUMNumbers, dates, times — 302, 22.05.2023, 10:30
PUNCTER_PUNCTPunctuation marks
LEGAL / MEDICAL / TECH / customER_DOMAINDomain terminology, shielded from incorrect splitting

WER_SCRIBE is the composite: all category errors over the combined reference-token denominator. CER_SCRIBE is a character error rate computed on normalized token streams — format variants cost nothing, and it needs no sandhi machinery to be robust to agglutination.

Try it — live in your browser

Loading Python engine · runs entirely in your browser
Examples

Sandhi awareness

Agglutination makes word boundaries unstable in Indic text: the same speech can be written as one word or two. SCRIBE detects such two-word merges and splits at alignment time and scores them as matches instead of errors. Each row below is real evaluation data — and each would count as 100% WER on its phrase without detection.

LanguageReferenceHypothesisJunction
Malayalamകഥാപാത്രമായിട്ടാണ്കഥാപാത്രം ആയിട്ടാണ്compound splitTry it ↗
Malayalamഎനിക്ക് അറിയാംഎനിക്കറിയാംvowel elisionTry it ↗
Kannadaಪ್ರಧಾನ ಮಂತ್ರಿಗಳಪ್ರಧಾನಮಂತ್ರಿಗಳcompound mergeTry it ↗
Kannadaಮಿತ್ರರಾಷ್ಟ್ರಗಳುಮಿತ್ರ ರಾಷ್ಟ್ರಗಳುcompound splitTry it ↗
Hindiभाई साहबभाईसाहबcompound spacingTry it ↗
Hindiउस मेंउसमेंpostposition mergeTry it ↗

On an internal benchmark of 48 Malayalam legal dictations, sandhi detection recovers 3.0 percentage points of WER_SCRIBE across 377 events that would otherwise masquerade as recognition errors. Detection is an orthographic heuristic, not a linguistic analysis — deliberately lenient, with documented limitations: see scope and limitations before relying on sandhi counts.

How this demo works

The demo runs the published scribe-eval wheel — the same code you get from pip — inside your browser via Pyodide (CPython on WebAssembly), with a small pure-Python stand-in for the Levenshtein C extension. Nothing you type leaves your browser. The page itself is static files served from GitHub Pages; the source lives in site/.