Meta has a long history of announcing video models that never ship — Make-A-Video in 2022, MovieGen in 2024 (the one that promised baked-in audio and stayed a research concept), then Vibes, which was Midjourney and Flux under the hood. This time is different: Muse Image is live right now at meta.ai, it’s free, and Muse Video is coming. Full breakdown’s in the video; here’s where things landed.
Muse Image: the gauntlet
Muse is a thinking (autoregressive) image model in the same family as GPT Image 2 and Nano Banana — and it has a “show thinking” toggle, so you can watch it reason through the image. Which is occasionally comedy: on one generation it got a storefront’s OPEN sign facing the correct direction mid-iteration, then talked itself out of the right answer for the final. AI — it can talk itself out of the correct decision, just like people.
Through the standard tests:
- Man in a blue business suit: solid. He’s on the sidewalk (progress), background characters hold up, decent text on the signage. The skin has that waxy AI texture and the suit could use one more pass — upscaler territory. Also a one-way sign pointing both ways.
- Wine glass at 3:30: the analog clock actually reads 3:30 instead of the usual 10:10 smiley face, and the glass is filled to a reasonable top.
- Pelican on a bike with a glass of wine: good, nice satchel detail, but the wine glass is held by a human hand. Unsettling.
- The Elder Scrolls 6 imagination test: genuinely impressive. Level 24 Breton Battlemage, mini-map in Hammerfell — both accurate to the lore — running 4K at 78fps. Nano Banana 2 gives you what amounts to an Elden Ring screen grab on the same prompt; Muse dug into the source material.
- Text (the Sharknado AI poster): clean title treatment, and it wrote its own tagline — “the bite is bite-sized.” Well played.
Where it falls down: laziness. Give it a long storyboard job in the chat interface and it generates a few shots, quietly stops, then asks if you’d like it to do the rest. Character consistency also drifts by the third shot. And editing across aspect ratios gets stretchy.
The ranking: one of the arenas has Muse Image at #2. In the video I put GPT Image 2 and Nano Banana 2 in a tie at the top, with Muse below them — and for cinematic outputs specifically, nothing beats Nano Banana Pro at 2K. The Banana survives.
One mid-edit curveball: ByteDance dropped Seedream 5.0 Pro while the video was being cut — selectable interactive editing and apparently layers. That could reshuffle this whole leaderboard; circling back on it.
Muse Video: real this time
Meta says Muse Video offers competitive prompt adherence, visual fidelity, and temporal consistency, and that they’re investing in audio-to-video sync and physically accurate fast motion — which is a description of Seedance with the serial numbers filed off. The samples (all 10 seconds, multi-shot with cuts around the 5-second mark) look decent, and one arena already slots it just below Seedance 2.0 at #3 — though that same arena has Gemini Omni Flash at #1, so take it as you will. Too early for a verdict from samples alone. The real story is that, unlike every prior Meta video paper, this one is actually going to release.
The free tool: pose + depth motion control
The second half of the video is a tool I built and am giving away — TheoreticallyMotion Control: drop in any footage and it bakes out an OpenPose skeleton plus a depth map — using Video Depth Anything for the consistent mode — that you can feed as a video reference into Seedance, Runway, or anything that accepts one. Real motion control over your generations, running locally, for free.
What’s in it:
- Pose + depth baking with fast and consistent (VDA) modes — VDA needs a Chromium browser and downloads the model from Hugging Face on first run (~30 seconds, faster after)
- Multi-character tracking, up to five, color-coded so you can tell the model “character in green, character in red”
- In/out points so you only render the section you need — fewer cuts means fewer chances for the model to screw up
- Smoothing, tracking, and confidence controls
The proof run was the hardest video-to-video test I’ve put on the channel — heavy movement, occlusion, characters leaving and re-entering frame. Kling turns it into interpretive dance and hands one fighter a spontaneous mustache. With the pose + depth reference, Seedance held on. And you’re not limited to swapping the character — the whole location can go too: same fight, now on a derelict spaceship.
One honest caveat: Seedance sometimes just ignores the video input entirely on an identical prompt. Might be a context issue; hoping 2.5 solves it.
The whole thing is on Gumroad — put $0 in the box and it’s yours. You get the pose UI, a user guide, and all the source code, so you can modify it, extend it, or gut it. Donations welcome, never obligated. It cost roughly 2.5 million tokens to build; consider those saved.
Google Omni and the new AI Apps are available over on Artlist, who sponsored the video.