Speech & Audio
How machines turn sound into text, text into sound, and, increasingly, sound directly into sound in real time.
Almost every audio model works on a spectrogram instead of the raw waveform. A spectrogram is a picture of how much energy sits at each frequency at each moment in time, and that is a far more learnable representation than a wall of raw amplitude samples.
Speech recognition then has one specific hard problem: you don't know which audio frames correspond to which output characters. CTC (2006) solves it by training a single network end-to-end, summing over every frame-alignment consistent with the correct transcript instead of needing frame-level labels. Attention-based encoder-decoders and RNN-T follow for streaming. The field settles on the Conformer (2020) as the standard encoder: convolution for local acoustic patterns, self-attention for global context.
Whisper (2022) then pulls a different lever entirely. Skip the architectural cleverness, train one big transformer on hundreds of thousands of hours of messy, weakly-labeled multilingual audio, and you get an ASR system that generalizes zero-shot to accents, noise, and languages it was never explicitly tuned for.
Text-to-speech runs a parallel arc: text to spectrogram, then spectrogram to waveform, as two specialized models rather than one, because those are different enough problems that splitting them wins.
Chain VAD, ASR, an LLM, and TTS together and you get a real-time voice agent. There the hard problem isn't any single model's speed. It is the sum of every stage's latency, plus handling the fact that people interrupt each other.
One sentence for an interview: modern speech AI is end-to-end neural networks eating what used to be separate signal-processing and language-modeling pipelines, one stage at a time.
01Intuition
A spectrogram is a photograph of sound.
A waveform is just amplitude over time: one number per sample, 16,000 of them per second for typical speech audio. (The sampling rate has to be at least twice the highest frequency you care about, or you lose information you can never get back. That's the sampling theorem in one sentence.)
Stare directly at that sequence of numbers and almost nothing useful jumps out. Two acoustically very different sounds can produce raw waveforms that look like similar squiggles. What distinguishes a "sh" sound from an "ah" sound is which frequencies carry energy, and that information is smeared across the waveform in a way raw amplitude doesn't expose directly.
So you transform it: chop the waveform into short, overlapping windows, roughly 20–25 milliseconds each, which is about the timescale over which speech is locally stable. Run a Fourier transform on each window to get "how much energy is present at each frequency, right now."
Stack thousands of these windows side by side in time and you get a spectrogram: time on one axis, frequency on the other, brightness (or color) is energy. A mel-spectrogram warps the frequency axis to match how human hearing resolves pitch, with much finer resolution at low frequencies than high. That happens to also be a more efficient representation for a model to learn from.
This representation shift, working in frequency space instead of raw amplitude, is why almost every audio architecture from GMM-HMM's hand-built MFCCs through Whisper's input features is built on some form of spectrogram. That default is starting to shift, though, since models like wav2vec 2.0 learn useful representations directly from the raw waveform instead.
Now the second idea, specific to speech recognition: the alignment problem. Say the audio for a two-second utterance gets sliced into 200 frames, and the correct transcript is 15 characters long. You don't know, and the training data doesn't tell you, which of those 200 frames correspond to which of those 15 characters.
Different speakers say the same words at different speeds, so the boundaries shift constantly. Hand-aligning them for every training example is exactly the kind of expensive, brittle labeling step you want a neural network to make unnecessary. CTC's trick, in one sentence: let the network propose a full frame-by-frame labeling and blame-average over every proposal that's consistent with the correct transcript, instead of insisting on one "true" alignment.
Speech systems rarely work on raw waveforms. They work on which frequencies are present, and the Fourier transform is the tool that finds them.
Try it Find the pitch of a tone with a Fourier transform
import numpy as np
sr = 16000 # 16,000 samples per second
t = np.arange(sr) / sr # one second of time stamps
wave = np.sin(2 * np.pi * 440 * t) # a pure 440 Hz tone
spectrum = np.abs(np.fft.rfft(wave))
freqs = np.fft.rfftfreq(len(wave), 1 / sr)
print("samples in one second:", len(wave))
print("loudest frequency:", freqs[spectrum.argmax()], "Hz")
440.0 Hz. The waveform is 16,000 numbers that look like nothing by eye. The transform turns them into a strong component at 440 Hz. Speech recognition builds on the same step, applied to short overlapping windows of audio.02Timeline
For three decades speech recognition was a pipeline of separately trained parts, GMM-HMM acoustic models over hand-engineered MFCC features plus a pronunciation lexicon and an n-gram language model, joined at decode time with no component trained end to end.
CTC (Graves et al., 2006; popular in deep ASR from about 2014–2016) let one network train directly on (audio, transcript) pairs by marginalizing over alignments, the Conformer (2020) became the standard encoder, and Whisper (2022) showed that massive weakly supervised multilingual training beats careful architecture tuning for real-world robustness.
Self-supervised pretraining (wav2vec 2.0, HuBERT) plus large-scale weak supervision (Whisper) is now the default recipe, and the frontier has moved to real-time speech-to-speech agents that listen, reason and speak back with the latency and turn-taking of a conversation.
03Architecture: how the pipeline evolved
Everything below is one continuous story: replace hand-engineered, separately-trained pipeline stages with single networks trained end-to-end, then replace careful architecture design with scale once scale becomes affordable.
The ASR pipeline, one generation at a time
GMM-HMM was the classical baseline for decades. Three components did three separate jobs:
- MFCCs. Hand-engineered features approximating human auditory perception, extracted per frame.
- A Gaussian Mixture Model. Scores how well those features match each phoneme sub-state.
- A Hidden Markov Model. Handles temporal sequencing: which state comes after which, and for how long.
That got combined at decode time with a separate pronunciation lexicon and an n-gram language model. Three components, three different training objectives, glued together by a search algorithm rather than trained jointly against transcript accuracy. Deep neural acoustic models were the first big neural win here, swapping the GMM's emission scores for a DNN's. But the HMM alignment machinery, and the separately-trained language model, stayed exactly where they were.
CTC removed the HMM entirely. (See 07 for the full mechanics of the collapsing rule; this page builds on it rather than repeating it.) A single network outputs a distribution over the vocabulary plus a blank symbol at every frame, trained directly on (audio, transcript) pairs with no frame-level labels at all.
It was originally RNN-based, later CNN or Conformer-based. Popularized in deep ASR around 2014–2016, with Baidu's Deep Speech line of work as the canonical example, it was the first architecture where alignment stopped being a separate subproblem.
Attention-based encoder-decoder ASR (the Listen-Attend-Spell family, mid-2010s) took a different route: treat transcription like translation. Encode the audio, then let an attention-based decoder generate characters or subwords one at a time, attending back over the entire encoded audio at every output step.
This drops CTC's assumption that each frame's output is independent of the others, since the decoder conditions on its own previous outputs, which tends to help accuracy. But the decoder wants to see the whole utterance before it can attend well, which makes streaming awkward without extra engineering.
RNN-T, the Transducer (Graves, 2012; adopted widely for ASR from around 2017 on), was built specifically to fix that. It combines three pieces: an audio encoder, a label "prediction network" that behaves like a small language model over labels emitted so far, and a joint network that combines the two frame by frame.
Two properties follow. Because it only ever needs audio decoded up to the current point, it is naturally streaming. And because the prediction network already models dependencies between output labels, it leans less heavily on an external language model than CTC typically does. RNN-T became the default recipe for on-device and other latency-sensitive streaming ASR.
The Conformer (Gulati et al., 2020) is an encoder architecture and not a full ASR system: it slots into a CTC, RNN-T, or attention-decoder head interchangeably. Its idea is to stop choosing between convolution and self-attention and use both in the same block:
- Convolution is excellent at local, translation-invariant acoustic patterns: the shape of a particular phoneme, roughly independent of exactly which frame it starts on.
- Self-attention is excellent at global context: using the rest of the utterance to resolve what an ambiguous or noisy sound probably was.
Combining them in one block made the Conformer the dominant ASR encoder architecture within a couple of years of publication.
Whisper (Radford et al., 2022) then changes the axis of improvement entirely. It is a fairly ordinary encoder-decoder transformer, with no CTC and no exotic alignment trick beyond ordinary cross-attention.
What is unusual is the data: roughly 680,000 hours of audio scraped from the web, paired with whatever transcripts already existed for it. That is weak supervision, with no carefully curated labels. It spans dozens of languages and several tasks trained jointly, including transcription, translation, and language identification.
The bet is that scale and diversity of real, messy, real-world audio buys more robustness than any of the architectural refinements above. Empirically it generalizes remarkably well zero-shot to accents, background noise, and languages it was never specifically tuned for.
Self-supervised speech representations
wav2vec 2.0 (Baevski et al., 2020) asks a different question. Can you learn a useful representation of speech from raw, unlabeled audio, the way BERT learns a representation of text from unlabeled text (see 03)? The recipe has four steps:
- A CNN feature encoder turns the raw waveform into a sequence of latent audio representations.
- Those latents get quantized into a learned discrete codebook.
- Spans of the latent sequence get masked, and the masked sequence goes through a transformer context network.
- At each masked position the model identifies which quantized latent, out of a set of distractors, was there, via a contrastive loss.
This is the speech analogue of BERT-style masked prediction: no transcripts needed at all during pretraining, just audio. Fine-tune the pretrained model on a comparatively small amount of labeled (audio, transcript) data afterward, and it reaches accuracy that would otherwise need far more labeled data from scratch.
HuBERT takes a related but distinct route to the same goal. It clusters audio features into discrete pseudo-labels first, then trains masked prediction of those cluster assignments. The underlying bet is the same: most of what a speech model needs to learn doesn't require a single transcript.
Text-to-speech runs a parallel arc, in reverse
TTS went through the same three eras as ASR, roughly on the same clock.
- Concatenative synthesis (pre-2010s) spliced together short units of pre-recorded human speech. Intelligible, but with audible seams at every splice point.
- Parametric synthesis replaced recordings with a statistical model predicting acoustic parameters, fed into a vocoder. More flexible, but the output tended to sound flat and robotic: the classic "text-to-speech voice."
- Neural TTS (2016 onward) replaced both stages with learned models, and is the reason synthetic voices now sound close to human.
Modern neural TTS still typically uses two specialized models rather than one end-to-end system. It is worth being clear on why.
An acoustic model predicts a mel-spectrogram from text. Tacotron-style systems do this autoregressively with an attention mechanism, which can occasionally skip or repeat words if the attention alignment goes wrong. FastSpeech-style systems predict a duration for each phoneme up front and generate the whole spectrogram in parallel, non-autoregressively. That is faster, and structurally immune to that failure mode.
A vocoder then converts that mel-spectrogram into an actual waveform. WaveNet generates raw audio samples autoregressively via dilated causal convolutions, one sample at a time: extremely high quality but slow. HiFi-GAN generates the whole waveform in a single parallel forward pass using a GAN-based objective, dramatically faster, and became the standard production vocoder.
Predicting a mel-spectrogram is a text-alignment and prosody problem, running at maybe 80–100 frames a second.
Synthesizing the waveform means producing tens of thousands of raw amplitude samples a second that match that spectrogram.
Different output rates, different failure modes, different problems. Specializing each one wins over forcing a single model to do both, and it lets you swap a vocoder without retraining the acoustic model.
VITS-style systems, which train a single end-to-end model to do both jobs at once, show the two-model split is a good engineering default and not a hard requirement. It is still the default for a reason.
Zero-shot voice cloning follows directly from this: instead of training a dedicated model per speaker, condition the acoustic model (and sometimes the vocoder) on a speaker embedding extracted from a short reference clip. A few seconds of someone's voice is enough to synthesize new speech in that voice, with no additional training for that specific speaker.
The rest of the speech stack, briefly
Three related but distinct problems get lumped together as "speaker stuff."
- Speaker identification. Picks one identity out of a known, enrolled set.
- Speaker verification. A binary same/different check against one claimed identity.
- Speaker diarization. Who spoke when, across a whole recording, with no prior enrollment.
The first two typically work by embedding a few seconds of audio into a fixed-size speaker vector and comparing embeddings by similarity against an enrolled reference. Diarization instead segments the audio and clusters segment embeddings into speaker groups.
Overlapping speech breaks that segment-and-cluster assumption, since a segment can no longer be cleanly assigned to one speaker. Handling it well needs either source separation up front, or a diarization model explicitly built to output more than one label at once.
Voice activity detection is the gatekeeper in front of nearly every speech system. It is a small, fast classifier deciding, frame by frame, whether a chunk of audio contains speech at all, which lets a pipeline avoid running expensive ASR on silence, and it is the first ingredient in endpointing (did the user just finish talking, or just pause?).
Keyword spotting is VAD's narrower cousin: an always-on, low-power model listens for one specific wake word so the rest of the system can stay asleep until it is needed, running under a much tighter compute and battery budget than a full ASR model ever has to.
Real audio rarely arrives clean. Three tools address that, and they complement each other:
- Speech enhancement. Predicts and removes noise, typically by estimating a mask over the spectrogram (which time-frequency bins are noise versus speech) and reconstructing a cleaner waveform from the masked result. This measurably improves both perceived quality and downstream ASR accuracy.
- Source separation. The harder version of the same problem when there is no single clean target, recovering several overlapping signals (multiple speakers, or speech mixed with music) instead of one.
- Beamforming. Given a microphone array, use the small timing and phase differences between microphones to spatially filter audio, boosting sound arriving from one direction and suppressing the rest.
Not all audio understanding is about words. Sound event detection identifies and timestamps non-speech events such as glass breaking, a dog barking, or a smoke alarm. Audio tagging does the coarser version, labeling a whole clip.
Music and general audio generation extend the same generative toolkit covered in 12 to sound: autoregressive or diffusion-based models operating over learned discrete audio tokens, in the spirit of wav2vec 2.0's quantized latents, trained to produce new, coherent audio rather than recognize existing audio.
Chaining it into a real-time voice agent
Wire a speech recognizer, a reasoning model, and a speech synthesizer together and you get a voice agent. Past this point, the interesting engineering problem stops being any individual model. It becomes the pipeline itself.
Each stage has its own job in the latency budget.
- VAD gates the whole pipeline and drives endpointing: the decision that the user has finished a turn and is not merely pausing.
- Streaming ASR emits partial hypotheses continuously, so the "final" transcript step is fast rather than starting from nothing.
- The reasoning stage ideally starts generating and streaming tokens the moment enough transcript is available, rather than waiting for every downstream step to fully settle.
- TTS mirrors that on the way out, streaming audio sentence by sentence as the response text arrives, instead of waiting for the entire reply before synthesizing anything.
Barge-in, shown as the dashed loop above, is the part that's easy to leave out and expensive to leave out. VAD has to keep listening even while the agent is talking. The instant it detects real user speech, the system needs to stop audio playback and cancel or redirect the in-flight generation immediately.
That's a full-duplex design requirement as well as a speed problem. A pipeline can hit every individual latency target and still feel broken if it can't be interrupted the way a person would be.
04The equation, stated in words first
CTC's collapsing rule, as introduced in 07: given one frame-level output path, first merge adjacent repeated labels into a single instance, then remove every blank. What's left is the final predicted transcript for that one path. CTC training doesn't pick one "correct" path. It sums the probability of every path that collapses to the ground-truth transcript.
- π one specific frame-level path: a label, or the special blank symbol, chosen at every one of the T encoder frames.
- P(π∣x) the network predicts each frame's label independently given the audio x, so a path's probability is just the product of every frame's chosen-label probability.
- ℬ(·) the collapse rule: merge adjacent repeated labels into one, then drop every blank. A blank inserted between two identical labels is exactly what lets CTC output the same character twice in a row (like the two t's in "letter") without the collapse rule merging them away.
- ℬ⁻¹(y) the set of every path that collapses to the target transcript y. Summing all of their probabilities means the network never has to be told which exact frame-to-character alignment is "the" right one, only that the total probability mass over every alignment consistent with y should be high.
- computing it naively, that set is enormous. But a forward-backward dynamic program, the same family of trick as HMM forward-backward, computes the full sum and its gradient in time linear in T. You don't need to derive it in an interview; know that it's tractable and isn't brute-forced.
05Why this beat the alternatives
GMM-HMM's acoustic model, pronunciation lexicon, and language model were trained on three different objectives and combined only at decode time. An error made by the acoustic model had no mechanism to be corrected by a later stage, because no stage was ever trained against the thing you care about: getting the final transcript right.
CTC, attention-based seq2seq, and RNN-T all fix this the same way: train the entire network directly against the real objective, end-to-end, so the whole system can jointly compensate for weaknesses anywhere in the pipeline instead of having errors compound silently through a chain of independently-optimized components.
Whisper's "scale over cleverness" bet works for a related reason. Careful architecture tuning historically bought accuracy on whatever benchmark or domain it was tuned against, but real-world audio has enormous unmodeled variability that any curated, clean training set under-represents by construction: accents, background noise, code-switching, cheap microphones, non-native speakers.
Training on hundreds of thousands of hours of diverse, messy, real internet audio exposes the model to that variability directly during training instead of hoping it generalizes from a clean benchmark distribution. It's the same lesson the LLM world learned around the same time (see 10): scale and data diversity often beat architectural cleverness for robustness. It just costs far more compute and data-collection effort to pull off.
06Tradeoffs and what breaks
| CTC | RNN-T | Attention seq2seq | |
|---|---|---|---|
| Streaming-capable | Yes, with a causal encoder | Yes, natively — built for it | Not naturally — decoder wants the full encoded utterance |
| Alignment assumption | Monotonic; per-frame labels treated as independent | Monotonic; label dependencies modeled by the prediction network | None enforced — decoder can attend anywhere, which is more flexibility than speech needs |
| Typical latency | Low — single forward pass, frame-synchronous | Low — designed for incremental, on-device decoding | Higher — best accuracy historically came from seeing the whole utterance |
| Needs an external LM for best accuracy | Often — frame independence leaves output-level language structure mostly unmodeled | Less so — the prediction network already behaves like a small LM | Less so — the decoder conditions on its own previous outputs |
"Lower ASR word-error-rate always means a better real-time voice agent." Not necessarily. A system with excellent transcription accuracy but poorly tuned endpointing, cutting people off mid-sentence or waiting too long before responding, or with no barge-in handling at all, will feel worse to talk to than a system with a slightly higher error rate and well-tuned turn-taking.
Once accuracy is reasonably good, latency and turn-taking UX usually dominate how "smart" a voice agent feels far more than the last few points of WER do.
What shows up in production: ASR accuracy degrades on accents, domains, and languages that were underrepresented in training data. A model can look excellent on an aggregate benchmark and still fail badly for a specific accent or noisy real environment it barely saw, which is why per-subgroup error rates matter more than one blended number.
TTS mispronounces rare words, proper nouns, and numbers. Text normalization (turning "$42.50" or "3/14" or "Dr." into the words you'd say) is a real, unglamorous failure surface that sits entirely outside the neural models themselves.
And real-time agents built by chaining VAD, ASR, an LLM, and TTS sequentially, without explicitly designing for interruption, will talk over the user or ignore them, and that is a full-duplex gap and not a model-quality gap.
Fully end-to-end speech-to-speech models (audio in, audio out, no intermediate text) are an active frontier. collapsing the ASR → LLM → TTS pipeline into one model can cut latency and preserve paralinguistic information (tone, emphasis, hesitation) that a text bottleneck throws away.
As of 2026, cascaded ASR → LLM → TTS pipelines are still the more controllable, more widely deployed default in production, mainly because each stage stays independently debuggable, swappable, and auditable. Treat "speech-to-speech has already replaced cascaded pipelines" as an overclaim if you see it stated flatly.
07Build this
CTC's sum over alignments is easy to state and hard to believe. One plot of the network's per-frame output settles it in a way no derivation can.
Train a small CTC recognizer on spoken digits, transcribing each clip as the spelled-out word: "three", "seven", "zero". Ten words, a couple of dozen output symbols, clips under a second. The task is trivial on purpose. What you are here to look at is the per-frame output distribution, because that is where the alignment the training data never supplied shows up.
- Get a small spoken-digit corpus, a few thousand short recordings. Label each with its spelled-out word. Your output alphabet is the handful of letters those ten words use, plus one blank symbol.
- Convert each clip to a log-mel spectrogram with
torchaudio. That is the representation from section 01, and it is the one part you should let a library handle. - Encoder: a couple of 1D convolutions and a small bidirectional GRU, then a linear layer to one logit per symbol per frame. Nothing Conformer-sized is needed: a few hundred thousand parameters is plenty.
- Before training, implement the CTC forward recursion by hand for one (audio, transcript) pair. Check your log-probability against what
nn.CTCLossreports on that same pair. That is the marginalization in section 04, computed rather than recited. - Train with
nn.CTCLoss. Then plot the per-frame softmax as a heatmap: symbols down one axis, frames across the other, one plot per clip. - Now break it. Delete the blank from the output alphabet and retrain, keeping everything else fixed.
Where this runs in production
A caller finishes a sentence. VAD and endpointing decide whether they're done talking or only pausing. Streaming ASR, which has been emitting partial hypotheses the whole time, finalizes the transcript quickly since it isn't starting from scratch.
The reasoning stage begins generating a response. What matters here is time-to-first-token, not total generation time, since the response can start being spoken before it's fully generated. Streaming TTS begins synthesizing audio for the first sentence or chunk as soon as enough text exists, rather than waiting for the complete reply, so audio starts playing well before the LLM has finished.
The budget across the whole loop is additive: a fast LLM doesn't help if endpointing alone waits an extra beat to be sure the caller is done, and a fast TTS model doesn't help if it's sitting idle waiting for the entire response text instead of being fed sentence by sentence.
Building a system that feels responsive means attacking every stage's latency at once, rather than optimizing any one model in isolation: streaming ASR partials, low time-to-first-token, sentence-level streaming TTS, and fast, well-tuned endpointing. Pipeline each stage so it starts consuming partial output from the stage before it the moment that output exists.
08The four things that decide real speech systems
Everything above is the standard tour. These four are what separate someone who has read about speech from someone who has shipped it.
Streaming or offline: the axis everything else hangs off
The first question to ask of any speech system is not how accurate it is. It is this: can it emit output before the audio finishes?
That single constraint eliminates most of the design space, because it decides what the encoder is allowed to see.
- Causal encoder. Each frame attends only to frames before it. Zero added latency, worst accuracy. A phoneme is often only disambiguated by what comes after it.
- Limited lookahead. Each frame sees a bounded window of future frames, usually via chunked attention, which is the working compromise in production.
- Full bidirectional. Every frame sees the whole utterance. Best accuracy, and unable to stream by construction.
Lookahead is the amount of future audio the model waits for before committing to an output. It is not an abstract knob. It is paid directly in wall-clock latency, and the user feels every millisecond of it.
Whisper is not merely tuned for offline use. Its encoder has a fixed 30-second input window, with positional embeddings sized to match. Audio shorter than that is zero-padded up to 30 seconds before the model sees it.
So the model always processes a 30-second block. Transcribing an hour-long recording means sliding that window and stitching the pieces together.
It is a fine design for batch transcription of a podcast and the wrong shape for an agent that has to answer in 300 milliseconds. See the model card for the input spec.
RNN-T sits at the opposite end, and its dominance on-device is not an accident. It consumes audio frame by frame, its memory footprint stays bounded no matter how long the utterance runs, and its prediction network supplies language-model-like behaviour without a separate LM to ship and fuse.
Why the Conformer's two halves match speech's two structures
Speech carries information at two very different timescales, and they want different machinery.
- Local, on the order of tens of milliseconds. Phoneme identity lives in the formant pattern of a short window. It is also translation-invariant: an /s/ is an /s/ wherever in the utterance it lands.
- Global, on the order of seconds. Speaker identity, prosody, and sentence-level context. Whether "read" is past or present tense may only be resolvable from a clause several words away.
Convolution is the natural fit for the first. Weight sharing across time is the translation-invariance assumption, stated as an architecture.
Self-attention is the natural fit for the second, because it reaches any distance in one hop.
Run only a transformer and it must learn local acoustic structure from scratch, with no inductive bias pointing at it. Run only a CNN and long-range context needs a very deep stack of dilations to reach. The Conformer stops treating this as a choice.
Neural audio codecs: the move that turned sound into tokens
Most treatments skip this piece, and it is the reason audio joined the LLM era at all.
A language model predicts the next token from a finite vocabulary. Audio is continuous, and arrives as tens of thousands of floating-point samples per second. The two do not obviously fit together.
A neural audio codec closes that gap: it learns to compress a waveform into a short sequence of discrete codes, and to reconstruct listenable audio from those codes alone.
The standard mechanism is residual vector quantization, and the idea is a cascade:
- The first codebook quantizes the encoder output. Coarse, and wrong by some residual.
- The second codebook quantizes that residual, correcting the first stage's error.
- The third corrects what is still left, and so on down the stack.
- Summed back together, the stages converge on the original representation.
Why a cascade and not one enormous codebook? High fidelity needs a very large number of distinct codes. A single flat codebook of that size is impractical to store, search and train. Several small codebooks compose to cover the same space.
SoundStream (Zeghidour et al., 2021) established the pattern, running at variable bitrates from 3 kbps to 18 kbps, streamable, and fast enough for real time on a smartphone CPU. EnCodec (Défossez et al., 2022) followed the same RVQ recipe.
Once audio is a sequence of discrete tokens, every trick from language modelling transfers directly: next-token prediction, scaling laws, and a shared vocabulary with text.
AudioLM made the split explicit: take discretized activations from a masked language model pretrained on audio to carry long-term structure, and codec codes to carry fidelity. Then treat generation as a language modelling problem over that combined sequence.
This is also what makes end-to-end speech-to-speech tractable instead of aspirational. Moshi is a speech-text foundation model built for real-time dialogue, with no text bottleneck sitting in the middle of the loop.
The evaluation trap: what word error rate does not measure
WER counts substitutions, deletions and insertions against the reference, divided by the number of reference words. It is the field's default number, and it is useful. It is also narrower than most people reporting it seem to realise.
- It usually ignores punctuation and casing. Standard WER is computed on normalized text, lowercased and stripped. "lets eat grandma" and "let's eat, Grandma" can score identically.
- Every word weighs the same. Dropping "not" and dropping "the" cost exactly one error each. Only one of them inverts the meaning of the sentence.
- It says nothing about who spoke. A perfect transcript with every speaker label wrong scores a perfect WER. Diarization has its own metric, diarization error rate, for exactly this reason.
- One aggregate number hides the distribution. Strong average WER is compatible with failing badly on one accent or one noise condition.
The practical consequence: pick the metric that matches the thing you are shipping. For a medical or legal transcript, error rate on entities and numbers matters far more than the blended figure. For a voice agent, intent accuracy after the LLM reads the transcript is closer to the truth than WER ever gets.
09Where you meet this in the wild
Speech is unusual in that the same handful of models get deployed under wildly different constraints. The model matters less than the latency budget it has to live inside.
The pattern to notice is that the deployment constraint picks the architecture. Streaming or batch, on-device or cloud, general vocabulary or specialised: answer those three and the shortlist is usually down to one.
- Voice is the natural interface. Hands or eyes are busy, or typing is not available to the user at all.
- The audio already exists and is unsearchable. Transcription converts a dead archive into something you can index, search and analyse.
- You need what text throws away. Tone, hesitation, emphasis and who spoke are all lost the moment audio becomes a transcript.
- Latency is generous. Batch transcription lets you use the largest, most robust model available and ignore the streaming constraint entirely.
- A form field would do. For structured input, typing is faster and a dropdown has no error rate at all.
- Your vocabulary is mostly rare proper nouns and you have no adaptation data. Generic ASR will mangle exactly the words you care about, and headline WER will hide it.
- No transcription error is tolerable and there is no human in the loop to catch one.
- Audio cannot leave the device for privacy or regulatory reasons and you have no on-device compute budget to spend.
10Interview questions
BeginnerWhat's the difference between a waveform and a spectrogram, and why do most audio models use the spectrogram?
A waveform is amplitude over time — one number per sample. A spectrogram is energy per frequency, over time, produced by running a Fourier transform over short overlapping windows of the waveform (the short-time Fourier transform). Models learn far more easily from the spectrogram because it exposes frequency-domain structure — the stuff that actually distinguishes phonemes and pitch — directly, instead of forcing the network to rediscover that structure from raw amplitude on its own.
BeginnerWhy does CTC solve the alignment problem in ASR?
Because you don't know, and the training data doesn't tell you, which audio frames correspond to which output characters — speakers vary in speed, so there's no fixed frame-to-character mapping. CTC sidesteps needing that mapping by having the network output a label (or a blank) at every frame, defining a collapsing rule (merge adjacent repeats, drop blanks) that maps many different frame-level paths to the same final transcript, and training by summing the probability over every path that collapses to the correct transcript. The network only ever needs (audio, transcript) pairs — never frame-level alignment labels.
IntermediateHow does wav2vec 2.0's pretraining objective work, and why is it called the BERT of speech?
A CNN feature encoder turns raw audio into a sequence of latent representations, which get quantized into a learned discrete codebook. Spans of that latent sequence are masked, the masked sequence goes through a transformer, and the model has to identify which quantized latent — out of a set of distractors — was actually at each masked position, using a contrastive loss. Like BERT, it learns a representation purely from unlabeled data by predicting masked content from context, letting you fine-tune on far less labeled (audio, transcript) data afterward than training from scratch would need.
IntermediateWhy does neural TTS typically use two separate models instead of one end-to-end model?
An acoustic model predicts a mel-spectrogram from text — a text-alignment and prosody problem running at maybe 80–100 frames a second. A vocoder converts that spectrogram into a raw waveform — tens of thousands of amplitude samples a second that have to actually match it. Different output rates, different failure modes, different enough problems that specializing each one, and being able to swap a vocoder without retraining the acoustic model, tends to beat forcing one model to do both. Single-stage systems like VITS show the split isn't mandatory, but it's still the more common production default.
IntermediateCTC, RNN-T, and attention-based seq2seq are all valid ways to build a neural ASR system — how do you choose?
If you need streaming and low latency — on-device, real-time transcription — RNN-T is usually the strongest default, since it's built to process audio incrementally and its prediction network already captures label dependencies without leaning as hard on an external language model. CTC also streams reasonably well with a causal encoder and is simpler to implement, but often needs an external LM fused in for the best accuracy. Attention-based seq2seq historically gave the best offline accuracy by letting the decoder attend over the full utterance, but that same flexibility makes it the most awkward of the three to make truly streaming.
DeepDesign a low-latency real-time voice agent — what are the pieces, and where does the latency actually come from?
VAD gates the pipeline and drives endpointing — deciding the user is actually done talking, not just pausing. Streaming ASR emits partial hypotheses continuously so the final transcript is fast rather than computed from scratch. The reasoning stage should start streaming tokens as soon as it has enough transcript, and TTS should synthesize sentence by sentence as those tokens arrive rather than waiting for the full response. The latency budget is additive across every stage, so optimizing any single model in isolation — a faster LLM, say — barely moves the needle if endpointing or non-streaming TTS is the actual bottleneck. On top of that, the system needs full-duplex barge-in: VAD has to keep listening even while the agent is speaking, and detecting real user speech mid-playback needs to immediately cancel audio output and redirect the in-flight generation. Get every latency number right and skip barge-in, and the agent will still feel broken the first time someone interrupts it.
DeepWhy did Whisper's "scale over cleverness" approach outperform more architecturally sophisticated but narrowly-trained ASR systems on real-world robustness?
Curated, clean ASR training sets under-represent the variability real-world audio actually has — accents, background noise, code-switching, cheap microphones, non-native speakers. No amount of architecture tuning against a clean benchmark teaches a model to handle variability it never saw during training. Training on roughly 680,000 hours of diverse, messy, weakly-labeled real internet audio exposes the model directly to that variability, so robustness comes from what the model was trained on rather than how cleverly it was built. It's the same lesson the LLM world converged on around the same time: scale and data diversity often beat architectural sophistication for generalization, at the cost of far more compute and data-collection effort.
11Go deeper
Other techniques for this problem
A scoped slice of the full Technique Map — every technique this page covers, grouped by what it solves.