Hitting sub-100ms: building real-time transcription UIs
Aug 18, 2025 · 7 min read
Hitting sub-100ms: building real-time transcription UIs
Real-time transcription is one of those features that feels like magic when it works and broken the instant it lags. On Neura I spend a lot of time keeping the perceived latency under ~100ms. Here's the playbook.
1. Separate "incoming data" from "what the user sees"
The network and the model don't run at 60fps; your UI should. I buffer incoming partial transcripts and flush them to the DOM on an animation frame, so rendering never blocks on the stream and the stream never blocks on rendering.
// Coalesce partials; paint once per frame.
let pending = "";
function onPartial(text: string) {
pending = text;
scheduleFlush();
}
const scheduleFlush = rafThrottle(() => {
setTranscript(pending); // one render per frame, not per token
});
2. Keep the hot path off React state
Updating React state on every token is a re-render storm. For the live caret I write directly to a ref / DOM node and only commit to state when a segment finalizes. The "boring" finalized text lives in React; the fast-moving tip does not.
3. Make it look instant even when it isn't
Optimistic UI buys you headroom: show the user's word the moment audio is captured, then reconcile when the model confirms. A confident-but-correctable UI beats a slow-but-perfect one.
4. Measure perceived latency, not server latency
Server time is necessary but not sufficient. I instrument capture → first paint and watch the p95, because that's the number the user actually feels.
Takeaway
Sub-100ms isn't one trick — it's a discipline: coalesce updates, keep the hot path out of React, render optimistically, and measure what the user perceives.