When it comes to AI though, humans are using biological neural net much more capable than any today's AI you can cram into a car. So, even if one accepts your premise of targeting human performance as a design guideline, more sensors is still logical at this point as way to compensate for the weaker AI.
Also, if you read how Tesla does vision it is very different from, and i think inferior to, how your eyes and brain build the 3d map of the surroundings. If one is limiting oneself to only vision, the first thing would be to try to get as good as possible that 3d mapping, and the vision seems to be among the simplest and most researched brain functions, ie. easiest to reproduce. As Tesla doesn't seem to be doing it - only may be couple years ago they only started to elicit the 3d model - i think they aren't on the shortest path to success when it comes to FSD.
Rotation is very common in nature.
Planetary rotation, inner-core rotation, spinning galaxies, dung beetle rolling, Keratinocyte migration, Rotifers, spirals, rotational symmetry, etc.
What isn’t common (but not non-existent) is using rotation for locomotion in biology.
Apples and oranges fall on the ground and can roll far and wide. Walnuts too.
Partial rotation is still rotation, of course: see animal joints in walk, trot and gallop.
And then there’s the belly-up pig drunk on brewery grain rolling down the hill. That mash packs a wallop!
For instance, when we see a ball rolling onto the street, we know that there is probably a young person nearby who wants that ball back. We don't have to be trained on the visual patterns of what might happen next.
Of course AI can be trained on the visuals of high probability events like this. But the number of things that can potentially happen is far greater than the number of training examples we could ever produce.
Models don't need to have been trained on every single possibility - it's possible for them to generalize and interpolate/extrapolate.
But, even knowing that it's theoretically possible to drive at human-level with only the senses humans have, it does seem like it makes it unnecessarily difficult to limit the vehicle to just that. Forces solving hard tasks at/near 100% human-level, opposed to reaching 70% then making up for the shortcoming with extra information that humans don't have.
They do have some in-distribution generalisation capabilities, but human intentions are not a generalisation of visual information.
Clearly that's possible to some extent, and in theory it should be possible for some system receiving the same inputs to reach human-level performance on the task, but it seems very challenging given the imposed constraints.
Also, for clarity, note that the limitations don't require the model be trained only on driver-view data. It may be that reasoning capability is better learned through text pretraining for instance.
Cars can't do this.
And not surprisingly the biggest problem with FSD is the accuracy of its bounding boxes.
But we do move our heads around pretty frequently. Enough to build mental records of what the bounding boxes are going to be for a range of objects.
If we're talking purely about going off memory, there's no reason why machines couldn't build up a similar catalog (which could be used by every self driving AI once learned). And human ability to judge distances varies significantly between drivers.
There may even be an AI built into your photo library app.
The fact your phone can identify an object doesn't inform you on the capabilities of self-driving car's vision stack. It's complete non-sequitur.
You know how big your own team is, and that your team is itself an abstraction from the outside world. You know you get the shortcuts of being able to look at what nature does and engineer it rather than simply copy without understanding. You know your own evolutionary algorithms, assuming you're using them at all, run as fast as you can evaluate the fitness function, and that that is much faster than the same cycle with human, or even mammalian, generational gaps.
> CLIP is proof of what AI can and can't do
CLIP says nothing about what AI can't do, but it definitely says what AI can do. It's a minimum, not a maximum.
Maybe this is an old post and your understanding has dramatically improved to the point where you're able to offer useful insight on ML/AI/self-driving?
2. Most ML is basic calculus and basic linear algebra — to the extent that people who don't follow it, use that fact itself as a shallow argument.
3. I'm not asserting how fast it can advance, I'm asserting that the comparison with "6 million years of evolution" is a as much a shallow hand-wave as saying it's trivial, as evidenced by what we've done so far.
- Over the speed limit (it's called a limit for a reason)
- Too fast for the conditions (speed limit != speed target)
- Too close to the vehicle in front of them
There are very few situations that can't be prevented by driving properly in the first place.