Descript Research

We develop our own AI models for audio and video generation, editing, and understanding. They power Descript's video editing workflows.

What we work on

Audio editing by latent inpainting

A two-stage system — a continuous neural codec plus a flow-matching transformer — that regenerates a masked span of speech conditioned on surrounding audio and target text, zero-shot, with inaudible boundaries.

Read more
Audio editing by latent inpainting

Anchored tree sampling beats autoregressive drift

A training-free, inference-time scheduler that replaces left-to-right rollout with anchor-bounded tree imputation — converting horizon-compounding drift into bounded drift, and running 5.3× faster than autoregressive baselines.

Read more
Anchored tree sampling beats autoregressive drift

Audio-driven lip-sync by masked frame regeneration

A flow-matching transformer that regenerates the lower face from new audio. Whisper-encoded speech and reference frames condition the model; masked velocity prediction during training keeps identity, lighting, and the boundary to untouched video in place.

Read more
Audio-driven lip-sync by masked frame regeneration

Invisible jumpcuts by regenerating the bridge

Trim a talking-head clip and the head jumps to a new pose. Jumpcut Smoothing masks the frames across the cut and generates a short bridge, then runs Video Regenerate over it to re-sync the lips — so the join plays like a continuous take.

Read more
Invisible jumpcuts by regenerating the bridge

Our ethics

Descript believes a person's likeness is part of their identity. We pledge to do our part to help individuals retain control of their voice.

Publications

Goodbye Drift: Anchored Tree Sampling for Long-Horizon Video-to-Video Generation

Goodbye Drift: Anchored Tree Sampling for Long-Horizon Video-to-Video Generation

May 19, 2026
PoDAR: Power-Disentangled Audio Representation for Generative Modeling

PoDAR: Power-Disentangled Audio Representation for Generative Modeling

May 11, 2026
High-Fidelity Audio Compression with Improved RVQGAN

High-Fidelity Audio Compression with Improved RVQGAN

October 26, 2023
Wav2CLIP: Learning Robust Audio Representations From CLIP

Wav2CLIP: Learning Robust Audio Representations From CLIP

October 21, 2021
Chunked Autoregressive GAN for Conditional Waveform Synthesis

Chunked Autoregressive GAN for Conditional Waveform Synthesis

October 19, 2021
NU-GAN: High resolution neural upsampling with GAN

NU-GAN: High resolution neural upsampling with GAN

October 22, 2020
MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis

MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis

December 2019

About Descript Research

Alexandre de Brébisson, Kundan Kumar, and Jose Sotelo founded Lyrebird in 2017, while studying under Yoshua Bengio at Mila. In 2019, after pioneering research in AI-generated speech and media synthesis, Lyrebird was acquired by Descript. While the team has evolved, its focus remains locked on developing technologies that power creative tools and advance the state of generative media.

About Descript Research

Most of this is still hard. That’s the job.

See Open roles