When you ask Arsaze's agent to do something — "cut the dead air," "add captions," "find the best clip of the dog" — it isn't guessing from your prompt alone. It's operating on structured timeline state: tracks, clips, timestamps, and (when you've asked for it) an analysis of what's actually in your footage.
That analysis step is opt-in by design. Uploading a file never triggers automatic scanning of its video or audio. Only when you explicitly run analysis on a clip does the agent gain the ability to reason about its content — what's said, what's shown, where the good takes are. Skip that step and Arsaze still edits, renders, and exports your project; you just lose the features that require understanding what's inside the clip.
Every action the agent takes lands on the same timeline you see and can edit by hand. There's no separate "AI version" of your project living somewhere else — one timeline, driven by you, the agent, or both in the same session.
We'll keep writing about the pipeline as it evolves — next up is a look at how b-roll generation picks shots.

