ByteDance keeps shipping, and the newest “seed” model isn’t a dance or a dream — it’s for your ears. I think a lot of people missed the mark on this one, because it looks like a strong video feature, and it’s cheap, which is not a word we get to use often anymore. There’s also a bonus at the end: a way to get Blender-style camera control without touching Blender.
Full breakdown’s in the video; here’s the part you can use.
What SeedAudio actually is
ByteDance released SeedAudio 1.0, and I think it got slept on because everyone saw “audio model” and immediately went to music. It can do music — it’s just not in the same league as a Suno or a Udio or any of the flagship music generators. That’s fine, because that’s not where it’s setting up home base. Where it lives is audio scenes: more like text-to-speech, but with the entire audio landscape available around the voices. In typical modern ByteDance fashion, this is basically Seedance 2.0, but audio. And it gets really good when you combine the two.
First look on Fal
I ran it over on Fal, for no better reason than I had about twenty bucks of credits sitting in there. It’s also fairly widely available via API. My prompt was a bit long — essentially the opening chapter to a high-fantasy audio drama — and the output was, for the most part, audio-drama worthy. (I’d make a Game of Thrones joke here, but I decided about six years ago that George can’t hurt me anymore.)
One aside worth keeping: I originally wrote the prompt as “the opening chapter to a high-fantasy audio novel, complete with voices and sound effects.” The model read the words “sound effects” out loud, in the drama. So watch your phrasing.
The real trick: audio as a Seedance reference
Here’s where it gets interesting for video. You generate the audio in SeedAudio first, then plug it into Seedance as an audio reference.
First test: sniper woman. She was generated up in Midjourney 8.2, which is in preview right now — you can access it by tacking --preview onto the end of your prompt. Quick mini-review while I’m here: 8.2 is looking good. It’s still Midjourney, so it’s not controllable, but it’s the first time in a while I’ve felt like that Midjourney aesthetic spark is back.
So I generated the audio for that image first, plugged it in, and got a one-liner: “Got you, clanker.” (My little ode to the Fat Man from Fallout, or I guess the BFG from Doom.) Then I ran it straight vanilla — no audio driving it — for comparison.
Honest caveat: the two generations sound pretty similar, but on the video side I actually preferred the vanilla version here. The catch is that vanilla fumbled the word “clanker.” Since you have both, though, you can cut them together into the shot you actually want. Granted, that one was a bit of a gimme example.
Where it shines: back-and-forth dialogue
Where things start to really work is with back-and-forth dialogue. Credit where it’s due — this whole rabbit hole kicked off for me with an experiment from Tom Likes Robots, running SeedAudio as the reference. It’s a genuinely cool output.
It’s hard to quantify how good that sounds without a point of reference, so: another creator ran the same lines through Seedance 2.0 fast. That’s fast mode, so of course it takes a quality hit — but it does illustrate that the audio reference is making a real difference.
The detective office
Back to our recurring Midjourney detective office, with a new Tues Sinclair generated in Midjourney 8.2, a new Malloy, and some generated audio. From a performance standpoint, this is some of the best acting we’ve gotten in this scene.
There are bugs, obviously. Where did the Venetian blinds come from? Where did Malloy get his magically appearing whiskey, and why can’t I get magically appearing whiskey? And this version of Malloy comes out like the love child of Pedro Pascal and Orlando Bloom while legally being neither of them.
The context-overload fix
I did hit problems on that generation. If you’ve ever messed with audio referencing in Seedance, you know it’s usually pretty disrespectful of your source material — I ran all three references at once (both characters, the location, and the audio) in a single prompt and got the typical Seedance mush: music in the background, repeating lines, mush-mouth at the end.
I figured at least part of that was context overload. So I pulled the image references into Nano Banana Pro — I did this over on Flora — to build one clean first frame. Nano Banana Pro seemed to do a really good job. (GPT Image 2 went off in a whole other direction and basically recast the office, so in this case NBP is the way to go.)
Timestamp your audio
The other thing that helps a lot: timestamp your audio. After you generate it in SeedAudio, you want a Seedance prompt that carries the dialogue, the action, and the timing of when each line lands.
Most LLMs can do this. I just dropped the MP3 into Claude — I’m in Code — and asked it to write a Seedance prompt with timestamps for that dialogue. It gave me back both the timestamps for each line and the Seedance prompt itself. You don’t have to use Claude; Gemini or GPT will handle it too. And you don’t need the big thinking model for this — leave it on medium thinking and save some tokens.
Plan on long takes
One note: SeedAudio likes long takes. That timestamped generation came out at 35 seconds; I tried prompting a bunch of times to make it 15 and it just wouldn’t listen. Most of mine landed around 35 seconds up to a minute, even up to a minute and a half. So plan on some light audio editing — for me these were really just cuts, nothing fancy.
It’s an extra step for now, but it’s worth noting Seedance 2.5 is not far off, and that has 30-second outputs. Clearly that’s where this is heading.
Voices and settings
A few of the other options. There are templated voices — I didn’t use any of them for these generations, and I don’t know if they’re a Fal thing or carry across the APIs, so I didn’t dig in. You can upload your own audio sample to lock a consistent voice, then tag @character1 / @character2 in the prompt, the same way you do in Seedance. You can also upload an image to have the model derive audio from it — I tried that here and there and it didn’t do much for me, so take that as you will.
Outputs come in wav and MP3, plus PCM and OGG Opus if you need something more esoteric. Sample rate runs from 8k up to 48k — 44.1 or 48 is probably where you’ll want to live. There are speed, volume, and pitch controls too; I left them all on default.
The animated scene, and one key tip
On to the animated example, with a tip that sounds silly but matters: explicitly call out the audio and dialogue as coming “from audio one” in your prompt. When I just threw the audio into the reference material without saying that, the generation was more prone to going off the rails. So spell it out.
Video outputs aside, the SeedAudio voices sound a little less stock — a little less video-game-character-y — than the vanilla versions. And the ambient sound in the generations is noticeably better, which makes sense given SeedAudio is only doing audio.
Why this matters
Here’s the whole point. Seedance is a pricey model, and I can pretty much guarantee 2.5 won’t come in cheaper than 2.0. So whatever we can do to improve our odds of nailing a solid generation before we pay for the expensive video pass, the better off we are. This isn’t foreign to filmmaking, either — it’s basically the animation workflow: audio first, then the visual.
Take one of our channel benchmarks: the quirky FBI agent drinking coffee in a Pacific Northwest diner. Running the audio first, we get the full exchange — “damn fine cup of Joe,” and the pie recipe “from the log lady who got it from the owls.” That “owls” line normally gets chopped or lopped off when you generate cold. Working audio-first, you’re guaranteed to keep it.
What it costs on Fal
And it’s cheap. Over on Fal right now it’s 9.4 cents for 30 seconds, or about 18 cents a minute — cheap enough that you’ve got real room to iterate and make sure you’re landing the generation you actually want. I’m curious whether this gets folded into Seedance 2.5 or stays separate, but either way, via this workflow it’s worth thinking about now, mostly to save yourself money.
Camera control without Blender
You’ve probably seen the Blender-to-Seedance workflows going around — simplified geometry in Blender to capture camera and motion control. I also know a lot of you will never install Blender. (Technically you don’t have to know Blender — you can drive it through the Claude-to-Blender MCP, which I covered in my “Claude vs. the Blender Donut” video, and I’m the perfect use case for that because I’m a Blender idiot.)
If you don’t feel like going through all that, the gang at Martini has a new feature that gets into a similar ballpark, and it’s very user-friendly. Quick note up front: I’ve done sponsored segments with Martini in the past. This is not one of them — it’s just something they shipped that I thought was cool and wanted to play with.
Their step-into-set feature takes an image and builds a Gaussian splat of the location, characters included. From there you can move and reshape the characters (mine came off the ground, so I nudged them back down), change camera position, and set the lens focal length. Hit capture and Nano Banana renders out the shot you need.
The new part is camera motion. On top of step-into-set, you now get a timeline: move the camera, change the angle, drop a keyframe. You can swap the lens length — go wider if you want — and the WASD keys move the camera around. Add another keyframe, and there are motion sliders for how much sweep you want to allow. When you’ve got a move you like, you save the camera move and render it out. I’d already rendered one — you can see the original point-based splat, then the final render with the keyframes in place. Genuinely interesting, and I’ll be playing with it more. It’s live on Martini now.
Between a feature like this and Kling OmniDirector, the next wave of camera-control updates is clearly coming, and those of us who are obsessive-compulsive about camera control are about to be very happy — and I don’t say that dismissively, I’m one of you.
Sound and vision, both in one video.