We wanted to answer a simple question. If you put a frontier model inside a real editing environment, strip away every shortcut, and give it a genuinely messy project the way a human editor would receive one, can it actually edit a video, and how does it compare to a person doing the same job with their normal workflow.
So we ran an experiment inside Arsaze, our own agent native video editing platform, using Claude Opus 5.5.
Setting up the test
We started with raw footage and assets from a real content creator, the same material they were planning to edit themselves. Inside Arsaze we organized everything into four folders.
One folder held the raw footage. One held images, some from the creator and some unrelated filler mixed in on purpose. One held sound effects and music, again a mix of what the editor intended to use and random extra files. The last folder was the important one. It contained animated clips relevant to this specific piece of content, along with unrelated animated clips, all saved as mp4.
Then we did something deliberate. We renamed every single file across all four folders to generic names like IMG_0001.mov, IMG_0002.png, IMG_0003.mp3, and so on, except the numbers had no order or relationship to each other. It wasn't a clean sequence like 0001 to 0002 to 0003. It was scattered, something like 0973, then 6173, then 9726, with no pattern connecting one file to the next. No filenames carried any meaning at all. The model could not skim a folder and guess what anything was from its name or its position in the list. It had to actually look inside every single file and understand the content for itself.
This creates an obvious problem. If a model tries to brute force its way through dozens of files by dumping everything into its context at once, it runs out of room fast. Arsaze is built specifically to prevent that. The platform manages how content gets loaded and processed so a model can work through a real project without blowing its context window, no matter how much raw material is sitting in the folders.
We told the model plainly that this was a test. The same content creator would be editing this exact video using their normal workflow, and we would compare the total time each of them took.
The model had full creative authority. It could generate one video clip using the Seedance model, generate one image using any model it chose, pull in sound effects or music from what Arsaze has built in, and decide which of the provided assets actually belonged in the final cut. It could remove anything it judged wasn't a good fit. From there, we stepped back and watched.
Finding the story before touching a single clip
The first thing the model did was read through the transcripts of every raw footage file, one at a time. You would think matching clips to a storyline is the easy part, but the creator had written the hook deliberately to be tricky.
Clip one was the creator asking, do you think bodybuilding is only for men. Clip two was the actual series intro, welcome to bodybuilding bible series part fifteen.
Before running this test with Claude, we tried the same setup once with ChatGPT. It swapped the order, treating clip two as the natural starting point and reasoning its way into the wrong sequence. Gemini handled it correctly. So did Opus 5.5. Both understood that the hook question was meant to come first, with the series intro following right after, exactly how the creator intended the video to open.
The raw folder also had a number of topically related clips that had nothing to do with this particular video. The model correctly separated those out and locked onto the real storyline before moving forward.
Reading everything before acting
From there the model moved into the second folder, the one mixing real assets with junk. It read through the contents and spent a noticeably long stretch thinking before taking any action at all. No guessing, no jumping straight to edits.
Then it opened the last folder, read through everything inside it, opened our animation page, and quickly started creating animated scenes with green screen backgrounds. At this point we genuinely had no idea what it was doing. We let it keep going because we had given it the right to pause and ask a clarifying question at any point during the test, and we wanted to see where that would happen naturally.
Partway through, it stopped and asked us for a screenshot of one of the raw clips showing the content creator on camera. We hadn't offered that, and at first we had no idea why it wanted it. We sent over a screenshot of one of the clips. Looking back, it makes complete sense. The model was planning a series of custom motion designs and animations that needed to sit around the creator in the frame, and it needed a reference point for where the person actually stood in the shot, their height, their positioning, how much space surrounded them, so it could place b-roll and animated elements without anything covering his face or cutting off awkwardly at his chest. That one image was enough for it to work out the rest of the framing on its own.
It stopped again once it had decided which assets it planned to use for b-roll. It had even generated custom sound effects through ElevenLabs to match specific keyframe motion designs it was building. Then it asked a direct question, should everything be placed on the timeline now.
We said yes, and also asked why it had rendered those clips with green screen backgrounds when Arsaze supports direct alpha channel transparency for animated video. It answered honestly. Every file in the animated clips folder was an mp4, so it had assumed they were all chroma keyed rather than alpha channel video. Within about thirty seconds of us pointing that out, it went back and converted every animated scene to a proper transparent background.
From planning to a finished edit
After that correction, the model went into another long thinking pass. When it came back, it moved fast, completing roughly seventy five percent of the overall edit in a single pass across the timeline. Then it shifted to clip level work, building custom keyframe animations for entrances and exits on individual clips.
The entire video was finished in twenty four minutes. Actual working time was closer to twenty two minutes once you separate out the parts that were not really the model thinking or editing, things like rendering time for both the AI generated clips and the animated scenes, us typing out clarifying questions about keyframes, delays from the music generation server, and ordinary network latency, which barely factored in.
The real timeline placement and clip level keyframing work took somewhere between fifty and eighty seconds. At the end, the model gave us its own closing note, something along the lines of, that's done, now let's see how the real editor does.
The human editor, working without any of this tooling, completed the same video in three hours and twelve minutes. His final cut was slightly better in overall polish. But given the difference in time, the speed the model delivered at deserves real credit.
One screenshot versus a stream of them
That single reference image is worth sitting with for a moment. A computer use agent that operates a screen the way a person would generally has to keep taking screenshots throughout a task just to understand what state things are in and what has changed. It is a constant loop of look, act, look again.
Here, one image was enough. The model asked for exactly what it needed, understood the spatial relationship between the creator and the frame from that single reference, and used it to plan complex keyframe based clip placement and motion design across the entire edit without needing to check back in visually again. Arsaze could support a continuous screenshot style workflow if a task called for it, but this experiment showed that with the right system around the model, a single well chosen image can do the same job.
The part that actually matters
The detail we think is most important here isn't the render quality or even the time difference. It's that the model never broke its context window despite working through a large volume of unlabeled, mixed, and partially irrelevant assets across four folders. That's not an accident. Arsaze is engineered specifically to manage context so a model can operate on a real editing project without hitting that wall.
Arsaze also runs a dual rendering engine, one built for standard multitrack timeline editing, and a separate one, closed source, that renders directly from code to video. That second engine is what let the model generate its own custom animated scenes from scratch and drop them straight into the timeline alongside the rest of the footage, rather than being limited to only arranging pre-made clips.
Alongside this post, we're attaching a sample video showing the raw clips, the actual AI edited output, and the timeline itself laid out at the top so you can see how the pieces came together.
In the next post, we'll be sharing a precise, clip by clip and minute by minute breakdown of the human editor versus the AI editor, which we think can serve as an early benchmark for this kind of work.

