Today we ran a small experiment. We opened a blank Arsaze project, uploaded two images, one solid black and one solid white, and handed the project to Claude Sonnet 5 with one instruction: make the best short video you can.
The setup
There was no footage, no audio and nothing else in the project. The video had to run between 30 and 45 seconds. The concept, story, pacing, style and mood were all the model's call. It wasn't allowed to ask us anything.
We added two constraints that shaped everything. First, it couldn't generate video. It could generate music and sound effects, and it could write animated clips, but it couldn't put the whole piece into a single animated clip and render that. Everything had to live in layers on the timeline, built from Arsaze's own text, shapes, transitions and keyframes. We told it this was a benchmark and to prove what it could do.
We also gave it a project name, claude-nothing-to-something. We didn't think much of it at the time.
What it decided
Before building anything, it pitched a concept in a couple of sentences. The video would be called Nothing to Something, and it would play as five acts, going from a black void with a single dot to a black dot on white paper as one blank image becomes a whole world.
We want to be accurate about where the title came from, because it's the part that surprised us. The name was in our prompt, as the name of the project. We never asked for a video about it. The model treated the only line in the prompt that wasn't a rule as the brief, and it built the whole piece around that idea.
What's on the timeline
The finished video is 40 seconds long. The base track is the two images themselves, cut on the act boundaries. From 30 to 34 seconds they strobe every half second, locked to the 120 BPM music.
On top of that are seven animated scene layers. Four story layers and the finale sit back to back. A HUD frame runs the full 40 seconds and switches its ink colour to contrast with whichever image is underneath it. A slice wipe is reused three times, at 7.45s, 15.45s and 27.45s, to hide the cuts between acts.
Above those are 12 overlay lanes holding about 65 clips of native timeline text and shapes. The text includes typed captions running from "THERE WAS NOTHING." to "SOMETHING.", two word-by-word cascades that land on the beat, and a closing subtitle. The shapes include orbiting dots, label bars that grow in, flying squares, sixteen twinkling stars, a shooting star, lines that swap colour with the strobe, and a burst of particles at the finale. It also cropped the black image into four circle-masked tiles that spin up on the drop.
The sound is all generated: one music track plus 14 sound effects across four lanes, each placed on a cue.
Why the layers matter
The rule we cared about most was the one against a single baked clip. One generated clip is a demo of a generator. A timeline where every star, caption, wipe and sound effect is its own clip is an edit. You can move any piece, retime it, recolour it or delete it, and the rest of the video still holds together.
The layers are also where you can see the model's most careful work, which was timing. The strobe is on the beat, the cascades are on the beat, the sound effects sit on their cues, and the wipes cover the cuts. That kind of alignment doesn't come from luck. It comes from an agent that can read the timeline and place things at exact times.
Why it needed Arsaze
Claude Sonnet 5 can't render a video by itself. What it can do is reason about structure, rhythm and taste, and it needs somewhere to apply that.
In Arsaze, every capability is exposed as an API and MCP tool. The agent reads the project, the assets and the timeline as files. It makes changes through tools, and it can apply a multi-step change as one batch that can be undone in a single step. So when it decided the HUD should change colour at 15 seconds, or that a wipe belonged at 7.45s, it could make that happen directly and then check the result.
Without an editor it can read and act on, the model has ideas and nowhere to put them. Without the model, Arsaze is an empty timeline waiting for someone. Together they turned two flat images into a 40-second film.
What this does and doesn't show
It doesn't show that an AI replaces an editor. The input was deliberately thin, and real projects come with footage, a client and taste that comes from people.
It does show how high the floor is. With a black frame, a white frame and the editor's native building blocks, the model made something with structure, rhythm and a point. Give it real footage and that floor only goes up.
Try it
If you want to run the same test, connect Claude to Arsaze over MCP, give it an empty project and two images, and see what comes back. Arsaze is live at arsaze.com.

