back
146 comments
In my day job I program rigid body behaviour in real time amongst other simulations. I think rigid body contact is hard to learn as it is inherently discontinuous.. something you discover when trying to code a solver.

As such I always use this prompt as a test: "A video of a jenga brick tower falling over as a brick is removed. The physics of each brick must be realistic."

It gave me a video of where bricks suddenly disapper or morph into others[1]. The linked video is after 2-3 iterations of me insisting on realistic physics. If you are just glancing at this, you would believe it is realistic.

That said this is still very impressive and one more step towards .. IDK what. But I am a bit reasurred that at least my job won't be fully replaced with AI :)

[1] https://streamable.com/2em1r3

> But I am a bit reasurred that at least my job won't be fully replaced with AI :)

I honestly can't comment with certainty that training from videos alone and whatever tokenization scheme they're using will ever get perfect dynamics.

However it is worth noting that transformers can do a pretty good job at learning dynamics with the right pipeline (not video): https://arxiv.org/pdf/2605.15305 https://arxiv.org/pdf/2605.09196

My point here being that representationally, it might be possible to learn good dynamics without a radically different approach/arch. There are already models that extract 3D tracking points from videos, so they could possibly be leveraged for learning dynamics (which on its own gives precedent for end-to-end approaches also possibly working).

I'm not sure why, especially because you're a developer... But damn, the amount of people that expect AI to just one shot stuff is hilarious. Half of the time I make a typo or something, should I be laughed out of the room?
Such videos are essentially dreams: how it feels that the planks should move, not what equations of rigid body physics would compute. And the feeling is realistic (even if overly dramatic in the end). If "stylistic transfer" works for static pictures spread out in space, why won't it work for the character of motion spread out in time?
I wonder what's the training data that makes it generate the final "explosion"...
Classic 3d simulation artifact with boundary conditions. I remember for an assignment where I had to model liquid with rigid bodies, they would suddenly gain infinite force at the corner and just disappear. It's clear that they must have used a lot of these kinds of synthetic data. But what's impressive to me, every release of these models, I am feeling less and less uncanny valley.
Totally unrelated, but what would you say the feasibility of writing simulation software for simulation of/replicating body movements during/in a martial arts technique would be?

I’ve often thought it would be very handy to have a proper simulator for being able to simulate and identify inefficiencies in one’s technique, but no idea whether it would be feasible to do.

That[1] video looks very Twin towers. Falls in on itself and then explodes.
thanks for intro to streamable
Some serious clipping
While at a cursory glance it looks as impressive as always, subtle spatial errors, and geometry that changes as it goes out of sight and comes back again hints at the fact that Google has still yet to solve the problem of deep spatial understanding.

Which considering just how pretty and detailed this whole thing looks, imo points at a fundamental issue at how these things are trained - it's as if there's no structure to its knowledge and training, like how an artist trained to draw would first try to understand simple 2d composition, then perspective, then light and shadow, mastering each concept and gradually building up a hierarchical understanding - it seems like its trying to learn everything at once.

I would rather see an AI model that I could give a floorplan of a building and it would generate an accurate flythrough on any path, even if it looked like butt.

Im not just talking out of my arse, I did work for a while in data science/engineering, and one of the big lessons people needed to be reminded of is to clean/downsample the data - a dataset consisting of a million samples could very well take 1000x as long to process as if we downsampled the whole thing to just a couple of thousand samples and we could learn the same conclusions with the fraction of expended time/effort.

I'm sure there's a similar logic in RL, that if you dump a trillion samples into the datacenter that consumes the same power as a city, what the model learns is what it could've learned with a much more curated training set and directed approaches.

At first usage I'm not impressed. I've probably spent a couple grand on Seedance 2 to date, and I can't find anything google omni flash does better than Seedance from running a handful of samples through the system. You can find some of the videos I've made in my HN bio link.
Just curious - are you at all concerned about the legal implications of ai-generating property listing videos?
Seedance 2 is amazing, compared with anything else American tech is producing. It does struggle with consistency like all other models.

The other problem is Seedance is heavily censored because of copyright concerns.

I have exactly the same thought. Anyone who had used seedance 2.0 a bit can tell Gemini is a bit behind, and seedance 2.1 is on the horizontal already.
> Prompt: Make it look like the weird shape of my hand hole super zooms and magnifies the ground it's looking at in sharper quality.

There's got to be a reason this is phrased so insanely, right?

Even weirder:

> Prompt: A skeuomorphism stop motion explainer about how the brain hippocampus works with a compelling voiceover. Don’t add seahorses. No voice cuts at the end. Don’t add text

Seahorses???

Yes, if you watch the video closely you can see that the "lensing" effect only really covers a circular area—this prompt probably went through multiple iterations where the author was trying to improve it so that the shape of the hand was reflected more closely.
Image-search for “hand hole” at your own peril.
At the bottom there is a "Try in Youtube Shorts" button.

Oh god...

We could be solving fusion power and instead we’re generating videos of birds in space or something. The market is a harsh mistress sometimes.
I'm an AI optimist. But AI video is probably the one thing that does depress me. Seeing that we can make anything visually, there's nothing that impresses me visually. I watch a video that two years ago I would've thought was really cool, and now my first thought is, "Yawn, is this AI?".

Video, more than anything else, is the place where I really care if something is AI or not. If I could get a TikTok that had no AI usage -- I'd be in. Which is weird for me, because I'm typically the guy who is all-in on AI.

It ruined the whole category of "cute animals acting goofy" content for sure.
I think the opposite. It allows more people to be creative. Similar to how the DAW allowed more people to become musicians. You can produce a hit song with just a laptop now.

Now you can have people producing videos without needing a crew of people.

I think it is like around 2010 or so I use to upload just this god awful music to early Sound Cloud because it was easy to make music with a DAW.

I even remember being on a psytrance production music mailing list 25 years ago and 95% of the tracks people posted were absolutely terrible, including myself.

I have seen a few incredible pieces from AI video but most has just not been that interesting. Then even the incredible pieces are 5 second one offs. No narrative, no continuity. I think of a random, real 5 second clip from Clockwork Orange with no backstory or context in the movie, who cares? Even the most visually interesting scenes wouldn't make sense and would be boring.

Right now it seems like we are at the stage of sampling random 5 second clips from early sound cloud and concluding this is the artistic utility of an entire new technology like DAW software and VST synths. That is obviously absurd.

For a few weeks, YouTube thought I wanted to see videos of package thieves being surprised by a booby-trapped box that was actually a glitter bomb. Video after video were these AI created shorts of supposed doorbell camera footage showing a thief running away with a box that explodes into a giant pink cloud.

I eventually picked one and opened the comments and the top comment was something like "This is obviously an AI video. Who watches this?" and the reply was along the lines of "me because I like seeing thieves get what's coming to them".

So you, like me, aren't interested in AI videos but I think there's a lot of people who don't care if it's real or not.

Thankfully, YouTube eventually stopped showing those to me. Now it thinks I'm interested in road rage videos. My YouTube feed outside of the three of four channels I've subscribed to is terrible.

> I can create more videos as soon as your limit resets. Check your usage in Settings

I did not create any videos yet.

Google, building great AI that nobody can try out.

But thx for the press release.

Browser crashes while scrolling because of all the auto playing videos. Please use IntersectionObserver to pause the video when not in display.
It's funny how they specifically use the phrase "output that follows real-world physics" to describe the marble rolling video. At the end of the zigzag track, the marble jumps up for no reason. In a couple of other places it speeds up with no apparent energy source. It's still an amazing result, but they could have picked a better example for this claim!
I think Hollywood is in for a rough era. The disruption is happening at break neck speeds.
What I'm hoping/waiting for is IMDB users creating alternative endings of movies.

It could make the comments section even more fun.

> I can create more videos as soon as your limit resets. Check your usage in Settings.

I have not used Gemini in a month.

This model is very misunderstood. It's not good at raw generation, but it has really deep world knowledge. Not just knowledge of physics, but just knowledge in general.

There are examples like providing a basic google maps view and then asking it to simulate driving from point A to point B, and it will generate landmarks from that location.

It's also great at consistency and editing. Probably SotA in editing.

Not sure what use cases will be unlocked but I sense something is there.

Does anyone else feel like Google is just always a dollar short and a day late here? Maybe not a dollar short, but it's like they've consistently been focused on the wrong thing. First they missed chatbots, now they're missing coding agents while they double down on chatbots and video gen (which OpenAI has already basically abandoned). Maybe this strategy is actually genius and I'm too stupid to grasp it.
Who is creative enough to drive this in any meaningful way?

Certainly not me - you have to be a great artist /designer to even imagine what to do with it.

To be honest, I think the performance of Gemini Omni Flash is still not as good as Seedance 2.0. You can try using both models on this platform. https://omnivideoai.co
What's the end goal of video generation? It feels unnecessary. Text generation leads to AI that can replace workers. Video generation is bad and only for video content generation, like movie and tv show production?
So it's really good, and we have reason to believe, never again, anything that happens in a video. Unless there's a super-product somewhere to authenticate footage?
Even though I don't have words to express how impressive this capability looks. I am genuinely scared at the harmful use cases of this.
Interestingly the `o` in GPT-o4 stood for Omni too (which I never realized until yesterday when reading random 3rd party documentation)
The people that think this output looks good are the same people that "don't get" art.

From a technical perspective, it's very impressive, no doubt. But from an artistic perspective I thought all of these examples on the site look bad.

When I click the link, the website crashes on my iPhone 13 iOS Chrome lol