Episode 88
AI VIDEO WARS: COMPARING OPENAI'S SORA 2 AND GOOGLE'S VEO
Keywords AI video generation, Sora 2, OpenAI, Google, video simulation, physics engine, audio integration, copyright issues, social media, creative industries Summary In this episode of the…
Keywords
AI video generation, Sora 2, OpenAI, Google, video simulation, physics engine, audio integration, copyright issues, social media, creative industries
Summary
In this episode of the Midget Step Podcast, host Dalton Anderson explores the advancements in AI video generation, focusing on the differences between OpenAI's Sora 2 and Google's Vio. He discusses the innovative features of Sora 2, including its built-in physics engine, audio integration, and the ability to create consistent narratives. The conversation also touches on the social implications of AI-generated content, copyright issues, and the contrasting approaches of Sora and Vio in the creative industry landscape. Dalton emphasizes the need for individuals to adapt to the evolving landscape of AI in creative fields.
Takeaways
Sora 2 represents a significant advancement in AI video generation. The built-in physics engine enhances realism in video simulations. Audio integration allows for more immersive storytelling. Consistent character narratives improve the quality of AI-generated stories. Users can opt into how their likeness is used in AI videos. OpenAI's approach may lead to copyright challenges in the future. The Sora and Vio models represent different philosophies in AI video creation. The future of creative industries will be shaped by AI advancements. Individuals must adapt to remain relevant in an AI-driven landscape. There is no right or wrong approach, only different perspectives.
Episode content
Explore every layer of this episode.
Each article, guide, analysis, and field note has its own focused page and stays linked to this source conversation.
Articles & stories
Narrative and editorial pieces that carry the conversation forward.
Google Veo 3.1 Product Profile and Channels
A sourced profile of Google DeepMind's Veo 3.1, including video and audio capability, current channels, Flow workflow, model-card evidence, and comparison limits.
Venture Step E088: Sora 2 and Veo Product Bets
A historical reconstruction of Dalton Anderson's Sora 2 and Veo comparison, with corrected physics, availability, product-channel, and legal language.
Sora 2 Product Profile and Historical Status
A sourced profile of OpenAI's Sora 2 launch, video and audio capabilities, social product, likeness and provenance controls, evidence limits, and 2026 discontinuation.
Research & analysis
Evidence-led work that tests and expands the claims in the conversation.
World Models and AI Video Research Note
This note supports a technical explainer that separates observed output behavior, vendor language, research definitions, interactive simulation, and hidden architecture.
AI Video Product and Likeness Research Note
This note supports the product-strategy and consent-based-likeness pages. Product controls are described as designed safeguards, not guarantees and not substitutes for pu
AI Video Evaluation and Continuity Research Note
This note supports the model-comparison and narrative-consistency guides. The protocol is designed to be reusable across vendors and versions. It does not contain a Ventu
AI Video Copyright and Character Research Note
This note supports an information-only issue map. It does not give legal advice, decide fair use or infringement, or replace review by qualified counsel in the relevant j
Field notes
Focused observations and durable ideas worth carrying into other work.
What Does World Model Mean in AI Video?
Distinguish visual plausibility, temporal coherence, learned dynamics, counterfactual control, interactive simulation, vendor aspiration, and architecture.
Test AI Video Character and Story Consistency
Use a five-shot story, continuity bible, controlled changes, blinded scoring, retry and edit-burden records, and archived outputs to test AI video.
How to Compare AI Video Models Fairly
Define the production job, run matched and native-capability tests, preserve versions and settings, score workflow value and failures, then retest.
AI Video Product Strategy: Feed or Workflow?
Choose the product loop before the model surface by comparing identity, creation, collaboration, distribution, monetization, safety, retention, and control.
AI Video Copyright, Characters, and Style
Separate output authorship, source use, derivative works, fair use, character protection, trademark, contracts, platform rules, style, and likeness.
How AI Video Likeness Consent Should Work
Design likeness consent across identity proofing, purpose, permitted creators, context, discovery, revocation, existing outputs, reporting, audit, and deletion.
Full episode
Read the complete record.
The show notes, transcript, and source trail remain on this canonical episode page.
TranscriptRead the full conversation.
E88 AI VIDEO WARS_ COMPARING OPENAI'S SORA 2 AND GOOGLE'S VEO
Transcript
Dalton Anderson (00:01.442) Welcome to Midget Step Podcast, where it's going to entrepreneurship, industry trends, and the occasional book review. Think about it, big picture. Where are we going with AI video generation? There's two core players, one core player, recently released, Soar 2, which is by OpenAI. Today we're going to be talking about some of the features and the difference in direction. Sorry, I accidentally hit my mic. The difference in direction between
Google and OpenAI. That's what we're going to be discussing. But before we do this, me tell you that I'm your host, Dalton Anderson. I do various things like this podcast, like to program, work in insurance and run, read, et cetera, et cetera. And if you do like this show at the end of the episode or in the middle, or you find something that I said insightful, or you just enjoy listening to my voice,
whatever it may be, please like, or end or follow slash subscribe to the podcast. It's very helpful. And Spotify is on my back. Two episodes ago, the Sora files, I discussed how Spotify is like, Hey, you have compared to your peers, a good amount of listens, but you're underperforming in the amount of followers that you have per your listens. And so, Hey,
Spotify is telling me I gotta be the bad guy. I'm just doing what Spotify told me to do. Anyways, but of course, let's jump right into it. I think the biggest thing between Sora 1 and Sora 2 is going from estimation of the world to legitimate world simulator. And what the biggest difference between Sora 1 and Sora 2 is Sora 2 has a built-in physics engine.
And that built-in physics engine allows the model to fundamentally understand how objects would interact with other objects in the simulated world and or the video that it's creating, which is simulated. And so that gives you an opportunity to do cool things. And not only is it able to do cool things, it's able to understand how things would interact with others. So a good example that I am thinking of off the of my head is if you're paddle boarding,
Dalton Anderson (02:29.398) and your buddy's like, Hey, I want to do a back flip. And the buddy does a back flip. If that was a Sora video and Sora one, the buddy might be able to do a back flip and then land on the paddle board, but it would be an odd way of the mechanics of the body moving and doing the back flip might be a little off the water from the person shifting their weight might not look realistic, but now with
Sora 2 with the physics model built in, the shifting of the weight right before the jump would make the water ripple a certain way. Once the person lands, the paddleboard will shift with the water. Then once the person lands, there's the shifting and then there's the ripples. And all of that stuff wouldn't have been done properly with the old model. And that's the biggest difference is
the world simulator is probably the chat GPT 3.5 moment where now it legitimately understands what's going on externally besides the context of the prompt, right? Where it takes the context of the prompt, knows how the world should interact with that prompt, and then it can create the story, a vibrant story for you, which is quite interesting.
So besides the physics engine that they created that's integrated into this video model, the next thing they did was integrated audio. So this is something that Google had for their rollout of VO, the new VO model that, I know the exact date when it came out, but it's been some time. They had integrated audio as well. Now SOAR 2 has integrated audio. And then another thing that
that Sora came out with was consistent narrative. So you could tell stories. One thing that these models were really limited until the most recent releases of Vio and Sora was you could tell a story. You could use it to do things. But it was really difficult to get a consistent character in that story, which makes it hard to tell stories if you've gotta figure out a way to hide the person's face or you can only show the person's face one time and then...
Dalton Anderson (04:58.286) that's the main character and then the sub characters are this, I guess the support characters in this instance would be able to jump in, jump out. Cause it wouldn't have to be that consistent person. But for the main character, it's a little different, right? Like you can't have, you can't have the main character be a black man and then later on Asian and then later on white, like it's gotta be the same person.
or like the hairstyles can't be drastically different all the time in the same scene on the same day. Now you're able to do that, which you wouldn't have been able to achieve before, which is great. Similar to the VO model. they're basically upping the capabilities of the model. With the consistent narrative, audio generation, and the physics engine built in, it just has
a better understanding of not only the prompt, how the prompt should interact, but also the whole story that you're trying to tell. You can inject your prompt into the story and change the narrative of what's being created, all the while just typing it out or using your voice, which is pretty great. Very straightforward.
But the next thing that was quite interesting was this cameo toolkit. I think they call it remix or recut, but basically you have this process now that you can, you can like insert yourself or your friends into these stories. If it's not you, then, and if it's someone else, they have to opt in. I haven't done it before, but basically,
you can opt into what type of videos you would be okay with your friends creating of you. So you could say, okay, like I don't want to be any political things or no nothing raunchy or something like that. And it won't allow that person to create those types of videos with you. But otherwise they have full permission to create whatever they want, which is interesting. It removes the liability from open AI, but also
Dalton Anderson (07:24.596) opens you up to your friends just making something crazy and being like, look what Jimmy put together or something. I don't know. But quite interesting this opt in deepfake where you're opting in to the service of your friends to able to create memes or stories with you. And then what type of stories would you be okay with them telling with yourself?
Interesting concept. I've seen it work out quite well and.
Dalton Anderson (07:58.544) Yeah, it's just, it's a, it's a different approach. And we'll talk about the two schools of thought, similar to the last episode where I talked about Microsoft Bing and then Google search engine, how they have two different schools of thought. OpenAI and Google have two different schools of thought on generation of video. So there's this social, seems like there's this social little narrative here going around. And then there's the next thing is,
There's the memes, the mashups, and then the copyright chaos where OpenEye was just, okay, let's throw it out there. It's gonna go viral. It's gotta go viral. So people are creating, there was an example of a Pikachu rescue meme or story using Pikachu, which is definitely against copyright rules and is an infringement. there's just fundamental things. They just threw it out there.
and people are just doing all absurd things with the model. But at the same time, it's going to cause a lot of copyright claims because people are using art, like art styles or proprietary characters to create stories and do things that might not be in the brand guidelines of that corporation. they're not thrilled about it. And nor are they able to monetize it because people are just doing whatever
they want with the characters that they've created. So it causes a lot of issues with their intellectual property in that regard. I think, I think R would count as like, I guess it wouldn't. Hmm. So intellectual poverty, but.
Dalton Anderson (09:47.702) It is an issue when you're the corporation on the receiving end, so okay, like Pokemon, for example, like multi-generational game, people play it, kids play it, it's a hugely popular thing. And for people to create stories with their Pokemon without their permission, it kind of opens you up to...
more abuse, right? If you don't do anything about it and it becomes okay, then when it gets to a certain point, it's not enforceable anymore. Because if you go to the court and then they have hundreds of examples of people using your characters and you didn't do anything about it, then it's kind of like, okay, well, what are you really defending? Because all these other people you didn't pursue, so why are you pursuing this person? And so it's like, okay, if you let a couple of people through, now it's a mosh pit.
which is what you'd want to avoid for sure.
So yeah, it's interesting. So the example was saving private Pikachu and then SpongeBob and Breaking Bad were the two examples that I was able to pull from the internet.
Dalton Anderson (11:03.949) Okay. So the next thing is I think the Sora versus Vio approach. So as I talked about earlier, Sora is a separate app. It's got this, and I skipped over it, but it's got this TikToki like feature. And I talked about Meta's feed, the AI feed where it just gives you AI slop the whole time. And you just scroll, scroll, scroll, scroll. It's infinite video, AI videos. Similar.
to open-ended approach where it's a separate app and you can just scroll the feed, this AI feed over and over and over and over again. Not my thing, nor do I support it. Yeah, I don't like those type of algorithms, not a fan. Because it fundamentally doesn't maximize your time. It's supposed to take your time away. It's not.
maximize to give you knowledge or to provide you the resources. It's just there to take up space. And there's a lot of things in life that are taking up space from you, your mind and your body. And there's enough of it, right? Like you shouldn't add things to take away from what it means to live, right? Without getting too philosophical.
But regardless, now you know my stance. Like I don't have TikTok. I have reals blocked on my phone. I have YouTube shorts blocked on my phone. I'm locked out. I'm locked out of all that stuff. I have it. I've had it and it wasn't an issue, but it's just like, I don't support it. And there's no way for you to take it off your phone. But anyways, before we get too off topic, the duopoly, Sora and Vio.
So I talked about how there's this social aspect with the remix and the cameo features where you can insert your friends if they opt into this deep fake capability. And then there's the separate app that has the infinite scrolling feed. is a completely different approach where VO is more of an enterprise solution to create ads. It's in the workspace account.
Dalton Anderson (13:30.049) It's not similar to the Nano Banana that I talked about like five episodes ago.
NanoBanana went viral with its code name, but this VO thing is not viral because it's used at an enterprise level. It's highly protected. You don't just throw it out there. They're being very careful on how they utilize the model and who utilizes it and how they want to proceed. And it's a very careful and thoughtful initiative where
It's less about the social aspect and more about the future relationship with that corporation. Whereas Sora is trying to create mainstream viral videos and
Dalton Anderson (14:24.826) provide a different perspective on how to create and how to interact in society. So I think it's more impactful, but it's less of a enterprise approach. So they just have two schools of thought, right? There's no fundamental wrong or right answer. There's just difference. And you can feel how you feel about the different approaches.
There's that always that Delta, right? But there is no fundamental wrong way of doing it. It's just a different way of going about it. And that's the beauty of it, right? That's the beauty of life is there's just, there's a lot of answers and there's not necessarily a lot of support to them. There's just people, people have their way of doing things and that may be the wrong way. Like I've got a buddy that I recently met and he hates fully clothes. And so he hangs everything up.
on the hanger except his socks, I think. So everything is hung up. He doesn't like folding clothes. So he hangs up everything except his socks. So the only thing that's not on a hanger is his socks, which I find crazy. Like that doesn't make any sense. It's not that big of a deal. But to him, it is. And so it's honestly wrong. It's just odd. And I have some oddities as well. So that being said,
there's not really a wrong approach. There's just different perspectives and we'll see how it all plays out, right? Like the outcomes might be different and the results might be different. And then looking back, you can say, well, you you could have seen this from the start and, and you could have, you could have saw that this is the wrong approach. This is a new era. You don't know what's going to happen. Like everything that we've got
going on right now is the first time it's been happening at this capability at this scale ever in human history. So you just don't know. You just have to roll with it and make your best decision at that moment and see how it plays out. That's it. There is no right or wrong answer. There's gonna be this future copyright battlefield and
Dalton Anderson (16:47.268) just complete restructuring of some of these marketing agencies and video videography. I think fundamentally that there's still gonna be different tiers, right? Like the best artists in the world, like the top 30%, they're not gonna be affected because they make the best stuff. They make the art that the AI models train on. But if you're just creating normal stuff,
On a day to day.
you might have to up your skills because there's just not gonna be that much of a differential between an ordinary, and this is for everybody. This isn't just for art or videos. This is legitimately for everybody. There's not gonna be that much of a difference between ordinary folk and what they do day to day and the AI model. It's gonna be able to do similar stuff.
And unless you're exceptional, it's not going to work out. I'm not trying to freak out or anything, but hey, I mean, you either rise with the boats or you drown. It's what it is.
I hope you found this interesting, this interesting perspective about...
Dalton Anderson (18:09.349) the new Sora video and
I don't know, let me know how you feel about it. How would you feel if your friends put you in a meme and it was just unhinged? I would be worried about giving my friends access to make unprovoked memes of me and insert me into stories. I just don't think it's gonna be good for me. I love my friends, right, but.
You know how it is the group chats like they would destroy you with these these videos and I don't know if I'm willing to give somebody that leverage on me. yeah, darling. You don't want to make another video of you, do you? Because I'll do it. I've got 45 minutes to spare. I can make a nice five minute. I could do a five minute video on you. No, please don't do that. I'll do it.
or whatever.
Dalton Anderson (19:17.519) No. So, just don't think I would allow my friends to make videos of me, and nor would I give that out to anybody as well. I guess it depends, but.
Not for me, not for me. But of course, wherever you are in this world, good morning, good afternoon, good evening, right? Thank you for listening in. Hope you listen in next week. Until then, doodaloo, goodbye, and see you next time.
SourcesFollow the source trail.
E088 Sources
Preserved episode evidence
[[E88 - Transcript - dalton (Dropbox copy 1)]] is the canonical raw monologue. It preserves Dalton's launch-era comparison of Sora 2 and Veo, demonstrations, product-strategy thesis, likeness discussion, feed critique, and copyright concerns.
The transcript supports Dalton's dated tests and interpretation. It does not establish hidden model architecture, a permanent capability ranking, the legal status of every generation, or current product availability.
OpenAI Sora 2
OpenAI's September 30, 2025 launch page described Sora 2 as a video-and-audio model with improved physical accuracy, synchronized sound, remixing, and consent-controlled likeness characters. It supports launch-era product identity and first-party claims.
openai.com/index/sora-2-system-card
The system-card page described more accurate physics, realism, synchronized audio, and progress toward simulating real-world complexity. It did not document a discrete built-in physics engine. The page now states that the Sora product was no longer available as of April 26, 2026.
deploymentsafety.openai.com/sora-2/overview-of-sora-2
The safety record documents launch-era restrictions, explicit opt-in likeness controls, provenance measures, teen safeguards, and content policies. These controls should be described as designed safeguards, not guarantees against misuse.
openai.com/index/creating-with-sora-safely
This later safety update described visible and invisible provenance signals, C2PA metadata, and stricter rules for real-person image-to-video. It also carries the current product-unavailability notice.
Google Veo
Google's current Veo page describes Veo 3.1 across Gemini, Flow, and developer channels, with native audio, creative controls, prompt adherence, consistency, and physics-related first-party claims. It shows that Veo is not solely an enterprise product.
deepmind.google/models/model-cards/veo-3-1-lite
The April 2026 model card documents inputs, outputs, distribution through API, AI Studio, Google Cloud, Flow, and Workspace, evaluation methods, and links to limitations and safety information.
deepmind.google/models/model-cards
The model-card index provides current version and update dates. Check it at draft time rather than treating the episode's Veo version as current.
Copyright and digital replicas
The U.S. Copyright Office AI initiative links the governing reports on digital replicas, copyrightability, and generative-AI training. These are separate issue areas.
copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-2-Copyrightability-Report.pdf
The Copyright Office says purely AI-generated material is not protected by copyright, while human-authored expression, selection, arrangement, or modification may be protected case by case. This is about copyrightability of outputs, not permission to use existing characters.
uspto.gov/sites/default/files/documents/copyright-and-ai-digital-replicas-report-part-one.pdf
The digital-replicas report explains the uneven state-law landscape for privacy and publicity rights and recommends federal legislation. It supports caution against universal likeness claims.
uspto.gov/trademarks/name-image-and-likeness
The USPTO's current guidance distinguishes trademarks, publicity and privacy rights, contracts, and AI-generated depictions of identity.
Provenance
Media provenance, C2PA, watermarks, and detection are covered canonically in [[E093 Sources]]. E088 should link there rather than duplicate the full standards explanation.
Evidence boundaries
Vendor benchmarks are first-party and often use different prompts, durations, access tiers, or evaluator conditions. A useful comparison needs matched tasks and preserved outputs.
Model capability and product strategy are related but not interchangeable. A model can appear in multiple products, and a product can change or disappear.
Draft-time checks
Recover the episode's prompts, outputs, settings, and dates. Identify the exact Veo and Sora access tiers used. Verify current availability. Obtain legal review for the copyright page and technical expert review for world-model claims.
Recovered episode media
The recorded Dropbox episode folder was searched on July 27, 2026. The full 392,751,551-byte episode video and six themed MP4 clips were found in 02_CONTENT_VAULT, including clips about the Sora comparison, physics language, opt-in likeness, and the two product directions.
The original prompts, settings, model identifiers, raw Sora outputs, and raw Veo outputs were not recovered. The dated Episode Story can reconstruct Dalton's argument from the transcript and recorded episode, but it must not publish a clip-level benchmark or recreate missing evidence.
Executed research
[[AI Video Evaluation and Continuity Research Note]] documents the task-specific comparison protocol, current model-card evidence, NIST evaluation controls, and the boundary against reporting unrun scores.
[[AI Video Product and Likeness Research Note]] documents current distribution, social-versus-workflow strategy, consent lifecycle, identity proofing, withdrawal, and the distinction between product controls and legal rights.
[[AI Video Copyright and Character Research Note]] separates output copyrightability, source and derivative-work questions, fair use, characters, trademark, contracts, platform rules, and likeness. Qualified legal review remains required before publication.
[[World Models and AI Video Research Note]] distinguishes learned dynamics, output plausibility, interactive simulation, vendor aspiration, and disclosed architecture. Qualified technical review remains required before publication.