Company News

Invideo Agent One ranks #1 in Physion-Arc 1.0 Benchmark

July 27, 2026

physion-benchmark

The agent led all three quality dimensions and placed first on 12 of 16 metrics in a seven-agent evaluation conducted by Physion Labs.

San Francisco, July 27, 2026

Invideo Agent One has ranked first overall in Physion-Arc 1.0, an independent benchmark from Physion Labs that evaluates AI video agents on minute-long, multi-scene video generation. Agent One led all three of the benchmark's quality dimensions and placed first on 12 of its 16 individual metrics, never falling below second on any metric.

invideo physion labs New.jpg

Image Caption: Overall quality, all 16 metrics weighted equally. Source: Physion Labs, Physion-Arc 1.0.

Physion Labs builds evaluation systems for generative video, work the company describes as building world critics for world models: systems that judge whether generated footage obeys the rules people expect reality to follow. The lab is led by co-founder and CEO Evelyn Qin Zhang, previously at AWS AI Labs working on Amazon Bedrock and Amazon Rekognition, and holder of an MIT PhD in physics-based simulation and machine learning, and by co-founder and CTO Bing Shuai, formerly a principal scientist at AWS AI. Its technical staff is drawn from AWS AI

Labs, Meta, Amazon's AGI group and Tencent Hunyuan. Physion-Arc 1.0 is the third benchmark the lab has released this year, following Galileo-0 and Physion-Atlas 1.0, which diagnoses prompt misalignment and visual defects in individual clips. Arc 1.0 extends that scope from single clips to complete minute-long films.

Physion-Arc 1.0 tests whether a generated video holds together as a piece of filmmaking rather than as a sequence of clips. Each agent received a complete screenplay and planned, generated and assembled the video in a single end to end run. Seven agents were evaluated across 100 screenplay prompts, producing 700 videos scored on 16 metrics grouped into Narrative Coherence, Cinematic Language and Production Quality. Scoring was done by human annotators, with every objective question answered independently by two annotators and disagreements resolved by a third-party adjudicator. Human taste was captured separately through pairwise comparisons and aggregated with a Bradley-Terry model.

invideo physion labs 2 New.jpg

Image Caption: Rankings within each of the three quality dimensions. Source: Physion Labs, Physion-Arc 1.0.

Measured Correctness

Physion-Arc separates its metrics into two families. Eight are objective, scored against fixed rubrics that measure correctness and consistency. Agent One placed first on this family with an aggregate of 82.3.

Its strongest result was Narrative Alignment at 97.8, the single highest metric score recorded by any agent in the evaluation. The measure asks whether each scripted beat lands in a chapter, in order, without padding or omission, and Agent One hit 97.5% recall and 98.4% precision. It also led Technical Coherence at 97.3, Spatial Consistency at 88.4, Identity Consistency at 79.6 and Causal Logic at 74.2.

Human Preference

The remaining eight metrics capture human judgement rather than measurement, scored by annotators comparing videos head to head. Agent One placed first on this family too, with an aggregate of 62.5, leading Rhythm and Pacing at 66.4, Audio Integration at 65.8, Emotional Intent at 65.4, Director Style Adherence at 63.9 and Camera Grammar at 63.8.

Leading both families matters more than leading either. Objective metrics establish that the output is correct. Preference metrics establish that people want to watch it.

invideo physion labs 3 New.jpg

Image Caption: Mean of the eight objective metrics and mean of the eight subjective metrics. Source: Physion Labs, Physion-Arc 1.0.

The split matters because the two families are not the agent's to win equally. Holding a character, a location and a story together across a full runtime is work no person can track by hand, and it is precisely what an agent should carry. The choices that sit on the other side, where the camera goes, where the cut lands, how a scene is paced, are the ones a serious creative keeps for themselves. Agent One is built to own the first job so the director is free for the second.

Consistency Across Prompts

Agent One posted the tightest spread in the field, with a standard deviation of 8.2 across the 100 prompts and only three prompts scoring below 50. Its worst single prompt scored 46.4, and removing its three weakest prompts moves its overall mean by 2.0 points. For studios and teams putting an agent into a production schedule, the absence of a long tail of failures is often the number that matters.

The Margin

Agent One's overall score of 72.4 places it first, 2.8 points ahead of the second-placed agent. Physion reports that gap as statistically significant, with a 95% confidence interval of +0.1 to +5.9 and a p-value of 0.046.

The Full Metric Breakdown

invideo physion labs 4 New.jpg

Image Caption: All 16 metrics, scored 0 to 100, with per-metric rank. Source: Physion Labs, Physion-Arc 1.0.

How Agent One Came to be Evaluated

Physion Labs published Physion-Arc 1.0 on July 15, 2026 with six agents. Invideo Agent One was not among them. The benchmark, its 16 metrics and its prompt set were designed, built and published by Physion Labs before invideo's involvement.

Invideo subsequently requested that Physion Labs evaluate Agent One against that same published benchmark. Physion executed the run end to end. Invideo supplied no generated outputs, and had no involvement in annotation, adjudication or scoring.

Invideo founder and CEO Sanket Shah framed the result as a marker of where the category is heading rather than a finish line.

"Agentic systems are the future of generative media, and of how creative work gets made at all. Ranking first on an evaluation we did not design is a testament to what this team has built, and it settles something worth saying plainly: this category will not be won on funding size. It will be won on technical depth and on strategy. Invideo builds for serious creatives, and a third of our own team are working creatives, which is why we build Agent One to hold the film together while the director makes the film. Agent Two arrives next week, four times faster and twice as cost efficient as Agent One. Mainstream film, television, advertising and serious media are moving to agentic production, and we intend to lead that shift."

About Physion Labs

Physion Labs builds evaluation systems for generative video. Physion-Arc 1.0 draws its screenplay prompts from T2F-Bench, an open collection of scripts derived from well-known cinematic scenes.

About invideo

Invideo is the agentic filmmaking platform for serious creatives. Agent One, the first generation of the platform, is a creative collaborator for filmmakers and creative professionals: they direct in plain language, the way they would on a real set, and the agent executes the production at scale. Agent One orchestrates more than 200 frontier image, video, audio and music models with long-term project memory, multi-shot editing and real-time collaboration. Invideo is backed by investors behind Stripe, Spotify, Flipkart and ByteDance. Learn more at invideo.io.

Share