←Home KnowML
Model AtlasChapter 34

Model Atlas: Speech & Audio

Transcription, streaming, synthesis, isolation and noise suppression are five different problems with five different model families. Picking one model for all of them is the usual mistake.

Reference 28 min read Snapshot: 6 September 2026 Assumes: speech basics (13)
Start reading
TL;DR

An audio product is a pipeline of specialists. Noise suppression, voice activity detection, transcription, synthesis and separation are all separate models, and the strongest model in each slot comes from a different lab.

Two numbers govern every choice: word error rate for accuracy and latency for whether it feels alive. They trade against each other directly, because a streaming model has less future audio to condition on than a batch one.

Licences bite harder here than anywhere else in ML. Several of the best-sounding open speech models are research-only, and the good ones are frequently not the free ones.

Read this before you trust a number below

Assembled on 6 September 2026. Speech leaderboards turn over faster than language ones because the top of the board is separated by fractions of a percentage point.

Use the tables to build a shortlist and the leaderboard links in section 13 to check it. The metric definitions and the pipeline structure will last longer than the rankings.

01The pipeline, and where the milliseconds go

Almost every audio product is the same chain. Understanding where each stage adds delay is what makes the model choices obvious.

mic in
→
denoise / isolate
→
VAD
→
ASR
→
LLM
→
TTS
→
speaker out

Every stage is optional except the one your product is named after, and every stage you add costs time. In conversation, humans start to read a gap as hesitation somewhere around 300 milliseconds, so that is the budget the whole chain has to fit inside.

A worked budget for a voice agent makes the pressure concrete.

  • Endpointing. Deciding the user has stopped talking. Often 200 to 500 ms on its own, and usually the largest single item.
  • ASR finalisation. A streaming model has already emitted most of the words, so the cost here is the emission delay and not the whole utterance.
  • The model turn. Time to first token, which page 31 covers in full.
  • TTS first audio. This is the time to the first audio chunk and does not include finishing the sentence. Streaming synthesis starts speaking before it has planned the whole utterance.
  • Network and jitter buffer. Small, constant, and easy to forget until you deploy outside your office.
Why the endpointer, not the model, is usually the villain

Teams chase a faster LLM when the perceived lag is coming from waiting to be sure the user finished a sentence. Shortening that wait makes the agent interrupt people. Lengthening it makes the agent feel slow.

This is a product decision disguised as a hyperparameter, and no model upgrade removes it. The systems that feel best cheat: they start generating on a partial transcript and cancel if the user keeps talking.

02Batch transcription: the accuracy tier

Audio you already have, transcribed as accurately and as cheaply as possible. Nobody is waiting, so the only metrics are error rate and throughput.

ModelSizeLicenceLanguagesNotable numbers
Parakeet TDT 0.6B v3600MCC-BY-4.025 European6.34% avg WER, RTFx 3,333; LibriSpeech test-clean 1.93%
Whisper large-v31.55BMIT99The multilingual baseline; RTFx around 69
Whisper large-v3-turbo809MMIT99Distilled decoder, several times faster than v3
Canary Qwen 2.5B2.5BCC-BY-4.0English5.63% avg WER, RTFx 418
Voxtral Small—Apache 2.0multilingualMost accurate open-weight system on Artificial Analysis, 2.8% AA-WER
Granite Speech 5.0 470M TurboCTC470MApache 2.0EnglishConformer CTC, non-autoregressive, built for edge
Qwen3-ASR-1.7B1.7BApache 2.0multilingualLLM-decoder ASR, one of the most-downloaded open models

The table splits cleanly along one architectural line, and the split is the main thing to take away.

  • Conformer encoder plus an LLM decoder gives the lowest error rates. The decoder brings language knowledge, so it repairs ambiguous audio using context. It is autoregressive, so it is slow.
  • CTC and TDT decoders give enormous throughput: Parakeet's inverse real-time factor is measured in thousands against Whisper's double digits, a 40-fold gap on the same hardware, at the price of a little accuracy.

For a large batch job that difference is the whole decision. Transcribing ten thousand hours at RTFx 3,000 versus RTFx 70 is the difference between an afternoon and a fortnight.

Whisper remains the default for one reason that has nothing to do with accuracy: 99 languages under MIT. When the language set is unknown or wide, nothing else is as safe a starting point.

03Reading word error rate

The same system scores 2.8% on one public leaderboard and 6.3% on another. Neither is wrong, and understanding why is the difference between using these numbers and being misled by them.

Word error rate is edit distance against a reference, normalised by reference length:

$$\text{WER} = \frac{S + D + I}{N}$$

Substitutions, deletions and insertions over the number of reference words. Insertions mean WER can exceed 100%.

Three choices sit underneath every published WER, and changing any of them moves the number by more than the gap between competing models.

  • Which datasets. An average over read audiobooks flatters everyone. An average that includes telephone speech, accented speech and meeting cross-talk does not. The Open ASR Leaderboard averages across roughly a dozen datasets and reports separate English, long-form and multilingual tracks for exactly this reason.
  • Text normalisation. Before scoring, both sides get lowercased, stripped of punctuation, and have numbers and contractions rewritten. An aggressive normaliser forgives real errors; a strict one punishes cosmetic ones. Whisper ships a famously aggressive one.
  • Long-form handling. A model with a 30-second window has to chunk a two-hour recording, and errors at chunk boundaries come from the chunking code and not from the model.
The error rate that decides your product is not WER

For most applications, the words that matter are a tiny subset: names, product codes, amounts, addresses. Getting "the" wrong costs nothing. Getting an account number wrong costs a support ticket.

Measure entity error rate on the terms your users say. A system with worse overall WER and better proper-noun accuracy is the better system, and no leaderboard will tell you which one that is.

There is one more practical gap. Benchmark audio is recorded near a microphone in a quiet room, while your audio may be a phone call from a car with a second person talking over the first. Expect real-world WER to be a multiple of the published figure and not a small increment above it.

04Streaming ASR: the latency tier

The moment a human is waiting, accuracy stops being the only axis. Streaming models trade error rate for emission delay, and the exchange rate is the thing to understand.

A batch model sees the whole utterance before committing to a word. A streaming model must emit as audio arrives, so it decides using only the audio so far plus a small window of future audio. That window is the emission delay, sometimes called right context, and it is the single dial that sets the trade.

ModelSizeLicenceDelayNotable numbers
Voxtral Mini 4B Realtime3.4B LM + 970M encoderApache 2.0240 ms – 2.4 s, configurableAt 480 ms: 8.72% FLEURS, 7.35% GigaSpeech, 5.05% Meanwhile long-form
Nemotron streaming (int4)0.67 GB quantisedopen weights0.56 s algorithmic8.20% avg streaming WER across eight benchmarks, runs on CPU
VibeVoice-ASR-Streaming 1.5B / 7B1.5B / 7BMITstreamingSpeaker-attributed transcription with hotword support, 10 languages
Parakeet TDT 0.6B v3600MCC-BY-4.0chunkedSame weights as the batch entry, run with a chunked context window

Voxtral Mini Realtime is the clearest illustration of the dial, because it exposes it directly. The same weights run anywhere from 240 milliseconds to 2.4 seconds of delay, and Mistral reports that at 480 milliseconds it matches leading offline systems. Below that, accuracy starts paying for the speed.

Chunked Whisper is not streaming

The common workaround is to slice audio into windows and run a batch model on each one. It produces output that arrives progressively, which looks like streaming, and it is not.

Two things give it away. Your latency floor is the chunk length, so a two-second chunk means a two-second delay no matter how fast the GPU is. And every boundary cuts a word in half, so accuracy degrades in a way a natively streaming model does not suffer. Use it as a prototype. Do not ship it as a real-time feature.

Note the Nemotron result carefully, because it points at where this field is going. A quantised 0.67 GB model achieving 8.2% streaming WER with half a second of delay, on CPU, means real-time transcription no longer requires a GPU or a network round trip. That changes what is possible on a phone.

05Text to speech

Synthesis is judged by ear, so the metrics are unusual. Three of them are objective and the one that decides adoption is not.

  • Time to first audio (TTFA). is how long before the first sample plays. It is the TTFT of speech, and it makes a voice agent feel responsive.
  • Real-time factor (RTF). is the seconds of compute per second of audio produced. A value below 1.0 means faster than real time, so a single stream keeps up.
  • Intelligibility. Transcribe the generated audio with a good ASR model and compute WER against the input text. It catches mispronunciation and dropped words objectively.
  • Speaker similarity. Cosine similarity between speaker embeddings of the reference and the clone. The number that matters for voice cloning.
  • Human preference Elo. Blind A/B votes on which sample sounds better. The only metric that tracks whether people will accept the voice.
ModelSizeLicenceLatencyNotes
Breeze TTS 23BResearch / non-commercial<40 ms TTFA, RTF 0.32 on H100English and Chinese, voice cloning, top open-weight Elo on the Artificial Analysis speech arena
Fish Audio S2 Pro5B (4B slow AR + 400M fast AR)Research / non-commercial~100 ms TTFA, RTF 0.195 on H20080+ languages, RVQ codec with 10 codebooks at ~21 Hz
Kokoro-82M82MApache 2.0very fast, CPU-capableStyleTTS 2 with an ISTFTNet vocoder, 8 languages and 54 voices, no cloning
VoxCPM22Bopen weights—Widely used open synthesis model
OmniVoice (k2-fsa)0.6Bopen weights—From the Kaldi lineage, built for deployment

The architectural split matters as much as it did for ASR. Autoregressive codec models like Fish Audio S2 Pro predict discrete audio tokens and sound the most natural, especially on emotion and emphasis. Non-autoregressive models like Kokoro run a fixed number of passes and are dramatically cheaper.

Fish Audio's dual-autoregressive design is a neat answer to the cost problem. A 4B model predicts only the primary semantic codebook, and a 400M model fills in the nine residual codebooks. The expensive network runs at the low rate and the cheap one handles the detail.

The licence is the constraint here, not the quality

The two best-sounding open-weight TTS models in the table are both research and non-commercial. Breeze TTS 2 requires a paid subscription for commercial use, and Fish Audio S2 Pro requires a separate commercial licence.

If you need a permissive licence, the field narrows sharply and Kokoro's 82M parameters under Apache 2.0 become much more interesting than its position on any Elo board suggests. Check the licence before you fall in love with a demo.

Every TTS deployment needs one more thing that most postpone. Voice cloning from a few seconds of reference audio is now routine, which makes consent and provenance an engineering requirement as well as a policy matter. Keep a record of which reference clip authorised which voice.

06Isolation, separation and noise suppression

These three get used interchangeably and they are different problems with different inputs, different outputs and different metrics. Naming yours correctly narrows the model list immediately.

Noise suppression

One speaker plus non-speech noise in, clean speech out. Traffic, keyboards, fans, room tone. Must be causal and cheap enough to run per frame on a laptop or a phone.

RNNoise, DeepFilterNet 2 and 3, DPDFNet, FRCRN, MossFormer2-SE.

Source separation

A mixture of several sources in, one stream per source out. Two people talking over each other, or a song split into vocals, drums, bass and other.

BS-RoFormer and Mel-Band RoFormer for music, SepFormer and MossFormer2 for speech.

Target speaker extraction

A mixture plus a short enrolment clip of one person in, that person's voice alone out. The right framing when you know who you want and do not care about the rest.

ClearerVoice-Studio's target-speaker models, including versions cued by lip video.

The metrics

  • SI-SDR, scale-invariant signal-to-distortion ratio in decibels. It is the standard separation metric and is usually reported as an improvement over the mixture, written SI-SDRi.
  • SDR. This is the older, scale-dependent version. It is still the reported metric for music separation, where a system at 9 to 12 dB average SDR is state of the art.
  • PESQ and STOI. Intrusive quality and intelligibility scores that need the clean reference. Fine for benchmarks, useless in production.
  • DNSMOS P.835. is a model that predicts human MOS ratings without a reference, reported as three numbers: signal quality, background intrusiveness and overall. The Deep Noise Suppression Challenge standardised on it, and you can run it on live audio.

On music separation the RoFormer family is the current answer. BS-RoFormer won the Sound Demixing Challenge separation track, and a version trained only on the standard MUSDB18HQ data reaches roughly 9.8 dB average SDR, rising to about 12 dB with extra training data. Mel-Band RoFormer improves on it further on vocals and drums.

The trap: denoising can make transcription worse

Enhancement models are trained to maximise perceptual quality. Removing noise aggressively also removes speech detail that a human ear does not miss and an ASR model relies on. The audio sounds cleaner and the word error rate goes up.

Modern ASR models were trained on noisy audio and are already robust to a great deal of it. If the output of your pipeline is a transcript and no human listens to the audio, A/B the denoiser against no denoiser before you assume it helps, because it frequently does not.

07Diarization and speaker attribution

Diarization works out who spoke when. It uses a separate model from the transcriber, and it is where multi-speaker transcripts usually fall apart.

Diarization error rate (DER) sums three failures: speech the system missed, non-speech it labelled as speech, and speech attributed to the wrong speaker. It is scored against a reference with a small tolerance around each boundary.

The open default is pyannote. Its community-1 pipeline under CC-BY-4.0 improves on the earlier 3.1 release across the standard sets, for example 11.7% against 12.2% DER on AISHELL-4 and 20.3% against 24.5% on AliMeeting.

Read those numbers next to the ASR table and something jumps out. A good transcriber is at a few percent error while a good diarizer is at eleven to twenty. Knowing who spoke is much harder than knowing what was said.

Two structural reasons for that gap.

  • Overlapped speech. When two people talk at once, a single-label-per-frame formulation cannot be right. Meeting corpora are full of overlap, which is why meeting DER is roughly double telephone DER.
  • Speaker counting. The system usually has to infer how many people are present. Getting that count wrong scrambles every label downstream, and it is a discrete decision with no graceful degradation.

The alternative to running two models is a speaker-attributed ASR model that emits words and speaker labels together, which is what VibeVoice-ASR-Streaming does. Joint modelling avoids the alignment problem between two separate outputs, at the cost of a much smaller field of models to choose from.

08Voice agents: cascaded or speech-to-speech

The architectural fork in every conversational product. Both are viable in 2026 and they fail in opposite ways.

Cascaded

ASR, then a text LLM, then TTS. Each part is swappable and every intermediate is a readable string. You can log the transcript, apply text-level guardrails, and use any language model you like. The costs are latency, because three models run in series, and information loss, because tone, emotion and emphasis do not survive the transcript.

→
Speech-to-speech

One model takes audio in and emits audio out, with no text bottleneck. It hears hesitation and laughter and can respond in kind. Kyutai's Moshi demonstrated full-duplex conversation at roughly 200 ms, and it runs on a single GPU. LiquidAI's LFM2.5-Audio-1.5B does the same at 1.5B parameters, pairing a 1.2B backbone with a 115M FastConformer encoder.

→
In practice

Enterprise deployments still lean cascaded, because auditability and control usually outrank naturalness. Speech-to-speech wins where the interaction is the product and interruption handling matters more than a transcript. Many teams run both: speech-to-speech for the conversation, a parallel ASR pass purely for logging and analytics.

Whichever you pick, turn-taking is where voice products fail. Barge-in, endpointing and backchannel handling are engineering problems in the transport layer and do not come with a model card, so budget for them.

09Build this

Published WER on read audiobooks tells you almost nothing about your telephone audio. Two hours of measurement replaces a month of guessing.

Project Your own WER, and the denoiser that made it worse ~3 hours · 30 minutes of your own audio

Measure three ASR models on audio that looks like yours, then test the assumption that cleaning the audio first helps. The second half is the part that surprises people.

  1. Collect 30 minutes of real audio and transcribe it by hand, carefully. This is the only tedious step, and it makes everything after it meaningful.
  2. Run three models from different architecture families: Whisper large-v3-turbo, Parakeet TDT 0.6B v3, and one streaming model. Record wall-clock time as well as output.
  3. Score WER twice, once with a standard normaliser and once with no normalisation at all. The gap between the two numbers is what normalisation is worth.
  4. Score an entity error rate over just the names, numbers and product terms. Compare the ranking to the WER ranking.
  5. Now run a denoiser such as DeepFilterNet over the same audio and re-run all three models on the cleaned version.
You'll know it worked when your three models rank differently on entity error rate than they do on WER. That inversion is the whole argument for measuring the thing your product cares about, and it shows up on most real corpora.
What the denoiser step teaches. On reasonably clean input, enhancement usually leaves WER flat or slightly worse, because the model was already robust to that noise and the denoiser removed speech detail along with it. On bad input it helps a lot. Finding the SNR where the crossover happens for your audio tells you whether to run the extra stage at all, and that is a real infrastructure saving.

10What breaks

Audio systems fail in ways that never appear in a benchmark table.

  • Benchmark audio is not your audio. Published WER comes from clean, close-miked, mostly read speech. Telephone bandwidth, accents, cross-talk and background music each cost you more than the gap between any two models on the board.
  • Normalisation hid the errors you care about. Aggressive normalisation forgives number formatting and casing, so a system that writes dates and amounts badly can still post an excellent score.
  • Chunk boundaries eat words. Any long-form pipeline built from a fixed-window model loses words at the seams, and no model change fixes what is a chunking bug.
  • The denoiser hurt the transcript. Enhancement optimises perceptual quality and does not optimise downstream recognition, so measure both.
  • Diarization collapsed on overlap. Two people talking at once breaks the one-speaker-per-frame assumption, and the resulting speaker count error corrupts every label after it.
  • The licence forbids the product. Several of the best open TTS and separation models are research-only. It is the most common late-stage surprise in audio work.
  • Latency was measured on an idle GPU. A TTS model at RTF 0.2 alone can exceed 1.0 under concurrency, at which point the audio stutters. Load-test the whole chain and avoid testing each stage separately.

11Pick by use case

Six shapes, and the stack each one implies.

Meeting notes and call analytics

This case is batch, long-form and multi-speaker. Accuracy and speaker labels matter and latency does not, and diarization quality will be your ceiling before transcription quality is.

Parakeet or Whisper for the words, pyannote community-1 for the speakers, or a speaker-attributed model for both.

Live captions

Human reading in real time. A native streaming model with a sub-second delay, and a display strategy that tolerates the model revising a word it already showed.

Voxtral Mini Realtime at 480 ms, or the quantised Nemotron streaming model if it has to run on CPU.

Chunked Whisper is the wrong choice.

Phone voice agent

The full chain under a 300 ms conversational budget, over 8 kHz telephone audio that every benchmark ignores. Endpointing and barge-in decide whether it feels human.

Streaming ASR plus a fast LLM plus streaming TTS, or a speech-to-speech model if auditability permits.

Bulk archive transcription

Ten thousand hours, nobody waiting. RTFx is the only metric that affects the invoice, and a 40-fold throughput difference dwarfs a one-point WER difference.

A CTC or TDT model such as Parakeet, batched hard.

Cleaning up recordings

This is podcast or video post-production. A human is the consumer, so perceptual metrics are the right ones and aggressive enhancement is fine.

DeepFilterNet for noise, a RoFormer separator to pull voice away from music.

On-device, no network

Fixed memory, no round trip, privacy as the requirement. This tier became usable in 2026, having been a compromise before it.

Quantised streaming ASR under a gigabyte, Kokoro-82M for synthesis, DeepFilterNet for noise.

12Interview questions

BeginnerWhat is word error rate and what are its limitations?

Word error rate is the sum of substitutions, deletions and insertions divided by the number of words in the reference transcript, so it is an edit distance normalised by length. Because insertions are counted, it can exceed 100%. Its limitations are that it treats every word as equally important, when in practice names, numbers and product terms carry nearly all the value, and that it is highly sensitive to text normalisation choices such as casing, punctuation and number formatting. Two published figures for the same model can differ by several points purely from the datasets averaged and the normaliser applied.

BeginnerWhat is the difference between noise suppression, source separation and target speaker extraction?

Noise suppression takes one speaker plus non-speech noise and returns clean speech, so the output is a single stream and the model is usually small and causal for real-time use. Source separation takes a mixture of several sources and returns one stream per source, which is the framing used for splitting two overlapping speakers or splitting a song into instrument stems. Target speaker extraction takes a mixture plus a short enrolment clip of one person and returns only that person's voice, which sidesteps the problem of deciding how many sources are present. They use different metrics too: DNSMOS and PESQ for enhancement, SI-SDR improvement for separation.

IntermediateWhy does a streaming ASR model have higher error rate than a batch one, and what is the dial?

Because it has less context. A batch model conditions on the entire utterance, including audio that comes after the word it is deciding, while a streaming model must commit using only past audio plus a short window of future audio. That window is the emission delay or right context, and it is the dial: widening it raises accuracy and raises latency in direct proportion. Systems that expose it, such as Voxtral Mini Realtime with its 240 millisecond to 2.4 second range, let you sit anywhere on that curve, and around half a second is where several models report parity with offline transcription.

IntermediateWhy do CTC and TDT decoders have such enormous throughput advantages over LLM decoders?

Because they are not autoregressive over text. A CTC decoder produces a frame-level output in a single pass and decodes greedily, so the cost is one encoder forward pass over the audio. An LLM decoder generates the transcript token by token, each step requiring a forward pass through a language model, which is the same memory-bandwidth-bound decode loop that limits text generation. The result is a throughput gap measured in one to two orders of magnitude, with Parakeet-class models reaching inverse real-time factors in the thousands against double digits for Whisper large-v3. The trade is accuracy, since the language model decoder repairs ambiguous audio using linguistic context that CTC has no mechanism to apply.

IntermediateYour voice agent feels sluggish but the LLM's time to first token is 200 ms. Where would you look?

At the endpointer first. Deciding that the user has finished speaking commonly costs 200 to 500 milliseconds on its own and is usually the largest single item in the budget, yet it is invisible in any model-level metric. After that, check time to first audio from the TTS rather than total synthesis time, since a streaming synthesiser should start speaking long before the sentence is planned. Then check whether the ASR is genuinely streaming or is chunked, because a chunked pipeline has a hard latency floor equal to the chunk length. The transport and jitter buffer are worth ruling out last.

DeepAdding a state-of-the-art denoiser in front of your ASR made accuracy worse. Explain.

Because the denoiser and the recogniser are optimising different objectives. Enhancement models are trained against perceptual targets such as DNSMOS or PESQ, which reward suppressing anything that sounds like interference. That suppression removes low-energy speech detail, particularly consonant onsets and fricatives, which a human listener does not miss but an acoustic model uses. Meanwhile modern ASR models were trained on large volumes of noisy audio and are already robust to moderate noise, so the enhancement gives them nothing to gain and something to lose. The practical rule is to A/B the denoiser at several signal-to-noise ratios and find the crossover point, below which it helps and above which it hurts.

DeepCompare cascaded and speech-to-speech voice architectures.

A cascaded system runs ASR, then a text model, then TTS. Every intermediate is a readable string, so you can log it, filter it, apply text-level guardrails, and swap any component independently. The costs are latency from three serial models and information loss, since prosody, emotion and hesitation do not survive the transcript. A speech-to-speech model consumes and emits audio directly, preserving that information and handling interruption naturally, with full-duplex open models such as Moshi demonstrating roughly 200 millisecond turnaround. Its costs are controllability and auditability: there is no transcript to inspect unless you run a separate ASR pass, and the field of available models is much smaller. Regulated deployments generally choose cascaded for that reason, and some run both, using speech-to-speech for the interaction and a parallel ASR pass purely for logs.

DeepWhy is diarization error rate so much higher than word error rate, and what drives it?

Because the task has a harder structure. Transcription maps audio to a word sequence and modern systems reach a few percent error, while diarization must additionally decide how many distinct speakers exist and assign every frame to one of them. Two failure modes dominate. Overlapped speech breaks the assumption of one speaker per frame, which is why meeting corpora score roughly double what telephone corpora do. And speaker counting is a discrete decision with no graceful degradation, since getting the count wrong permutes or merges labels across the entire recording. Strong open pipelines sit in the eleven to twenty percent range on standard sets, against a few percent for transcription on comparable audio, and the gap is structural rather than a matter of model quality.

13Go deeper

The live boards and the papers behind the metrics.

📊
Leaderboard
Open ASR Leaderboard
WER and RTFx across a dozen datasets, with separate English, long-form and multilingual tracks. The board to shortlist a transcriber from.
📄
Paper
Open ASR Leaderboard: Reproducible and Transparent Evaluation
86 systems across 12 datasets, and the finding that LLM decoders win on WER while CTC and TDT win on throughput — arXiv:2510.06961
📊
Leaderboard
Artificial Analysis — Speech to Text
Open and closed systems side by side, with price per thousand minutes. Note that its WER definition differs from the Open ASR board's.
📊
Leaderboard
Artificial Analysis — Speech Arena
Blind listening votes scored as Elo, filterable to open-weight models. The only honest way to rank synthesis quality.
📄
Paper
Robust Speech Recognition via Large-Scale Weak Supervision
Radford et al. — the Whisper paper, and the argument that scale of weak supervision beats architecture tuning — arXiv:2212.04356
📄
Paper
Music Source Separation with Band-Split RoPE Transformer
Lu et al. — BS-RoFormer, winner of the SDX23 separation track — arXiv:2309.02612
📄
Paper
Mel-Band RoFormer for Music Source Separation
Wang et al. — mel-scaled sub-bands, and better vocals and drums than BS-RoFormer — arXiv:2310.01809
🔧
Tool
ClearerVoice-Studio
Enhancement, separation, super-resolution and target speaker extraction in one toolkit, with pretrained FRCRN and MossFormer models — arXiv:2506.19398
🔧
Tool
DeepFilterNet
Real-time full-band noise suppression small enough for embedded devices. The default open denoiser.
🔧
Tool
pyannote.audio
The open diarization default: segmentation, embedding and clustering as one pipeline.
🔧
Tool
Moshi
Kyutai's full-duplex speech-to-speech model and the Mimi codec underneath it. The reference open implementation of real-time conversation.
📄
Paper
The INTERSPEECH 2020 Deep Noise Suppression Challenge
Where the DNS datasets and the subjective testing framework behind DNSMOS come from — arXiv:2005.13981

My Notes — 34 Model Atlas: Speech & Audio

Free notes

Highlights on this page