Ox challengerNemotron EN 0.6B native c8 · public Nemotron 3.5 reference
vsReference stackElevenLabs Scribe v2 Realtime
Ox demonstrated a lower raw-compute denominator on the English checkpoint. Separately, AA shows the newer open-weight Nemotron 3.5 at 80 ms finishing 1.93× faster than Scribe after speech end, but with 2.33× the WER. These are distinct experiments—not a combined parity result.
Nemotron raw compute / audio minOx modeled
$0.00294Modeled from measured A10 c8 realtime density; excludes idle, HA, networking, and operations.
Native c8 artifact ↗Modeled active-compute gapOx modeled
2.21× lowerAt eight continuously occupied streams. Scribe list price divided by raw active compute; the modeled advantage disappears below four occupied streams and is not an invoice comparison.
Native c8 artifact ↗Nemotron synthetic entity recallOx measured
10 / 10Exact normalized entity matches across five single-voice synthetic medical utterances; not a customer-corpus result.
Native c8 artifact ↗Nemotron synthetic corpus WEROx measured
10.1%Five single-voice synthetic medical utterances; not comparable to Scribe's third-party corpus.
Native c8 artifact ↗Nemotron first text p95Ox measured
809 msServer-side native emission at 80 ms cadence and A10 c8; excludes network and is not directly comparable to Scribe's published boundary.
Native c8 artifact ↗Nemotron c8 compute cadenceOx measured
0 misses42.9 ms p95 compute step against the 80 ms realtime deadline; eight submitted streams.
Native c8 artifact ↗Parakeet first text p95Ox measured
353 msServer-side native emission at c8 and 80 ms cadence; excludes network and cannot be compared with AA's after-speech boundary.
Parakeet c8 artifact ↗Parakeet tagged-string hitsOx measured
7 / 10Favorable substring scoring across five synthetic utterances; medical smoke gate failed.
Parakeet c8 artifact ↗Scribe realtime latencyVendor reported
~150 msPublished model latency; not yet reproduced on the Ox corpus.
ElevenLabs models ↗Scribe realtime list priceVendor reported
$0.39 / hrPublished API price before negotiated volume terms.
ElevenLabs pricing ↗Scribe realtime WERThird-party
3.64%Third-party streaming benchmark; use the paired customer corpus for the decision.
AA streaming benchmark ↗Scribe final after speech endThird-party
~140 msThird-party streaming result; timing definitions must match in the Ox run.
AA streaming benchmark ↗Next proof gateUse AA to avoid a broad Eleven baseline run. Only after domain tuning, run identical customer audio through the surviving candidate and Scribe to establish medical-term accuracy and end-of-speech latency.