Work Hard Everywhere logo Work Hard Everywhere

Applied Research Scientist, AI Research

Descript
📍 Anywhere in the World 💰 🕑 Any timezone
Full-time Mid-level Engineering All Other Remote

Job Description

Headquarters: San Francisco, CA or Remote, US

Descript's Research team builds the models behind the product's most distinctive features: Video Regenerate and lipsync, video translation, zero-shot voice and roomtone cloning, and Studio Sound. We don't build general-purpose generative models. We pick specific problems in the editing workflow and build specialized models for them. This isn't research for its own sake. Everything we build is meant to ship, and most of it has, going from prototype to a production feature used by millions of creators within months.

This role is focused on multimodal understanding: training models to perceive edited media the way a human video editor does. Underlord, our AI editing agent, reasons about a project largely through a textual representation of it. Giving it direct perception of the media it's working on is what will let it judge its own output and reason about the creative choices in an edit, not just the structure of a project. It's also an open research problem, since there's no settled way to represent or evaluate editorial craft, whether a cut lands or whether the pacing works. We have a unique dataset to work with.

Some recent work from the team:

- Audio editing by latent inpainting : regenerating a masked span of speech

- Video Regenerate : regenerating a speaker's lower face to match new or translated audio

- Jumpcut Smoothing : generating a bridge across a cut so the join plays like a continuous take

- Anchored Tree Sampling : tree-based imputation that bounds drift in long video generation

- PoDAR : disentangling power from semantics in audio latents to make them easier to model

More at descript.com/research .

What you'll do

- Multimodal understanding: build vision-language systems that let Descript's agentic editing features reason over the visual and audio content of a project.

- Evaluation: design the benchmarks and evals that make editorial quality measurable, and that balance quality against cost and...