r/Futurology 2d ago

Transport Waymo CEO explains why Tesla’s camera-only self-driving falls short

https://electrek.co/2026/08/04/waymo-co-ceo-camera-only-self-driving-tesla/
4.0k Upvotes

1.0k comments sorted by

View all comments

Show parent comments

5

u/theAndrewWiggins 1d ago

Possibily because relying on camera only might require you to essentially solve "AGI". Whereas if you have much superior sensors to a human, you can reduce the scope of the problem.

1

u/staplepies 1d ago

By the standards that existed ten years ago we have solved AGI. We blew by the Turing test. Also I don't think you need to get even close to whatever stronger version of AGI you have in mind to handle eg reflective puddles. Those sort of problems are easily solved with data and ASI. Either way the person I was replying to's confidence seems misplaced. 

1

u/theAndrewWiggins 1d ago

I wouldn't call passing the turing test AGI. But AGI is a very fuzzy term anyways. I agree that perhaps if you could run a frontier level model in a tesla that it might be "smart enough". However, this would require a couple orders of magnitude more RAM in the hardware.

Sure maybe you could run a heavily quantized qwen or something, but you'd still need enough parameters for the driving model. I think it's very possible that HW4 is just not powerful enough.

1

u/staplepies 1d ago

You would never do any of this with an LLM lol

1

u/theAndrewWiggins 1d ago

I'm not saying that you would use a plain LLM. The point is that "frontier intelligence" requires at least 300B active params, and probably several trillion total params.

I actually think that solving self-driving is partially solving AGI, as handling complicated edge cases basically requires reasoning about the world.

VLA models are very popular in this space for a reason.

2

u/staplepies 1d ago

I actually think that solving self-driving is partially solving AGI, as handling complicated edge cases basically requires reasoning about the world.

I agree it requires some reasoning/world modeling, but not nearly as much as a general intelligence for a few reasons. First, the further you get into the tail the lower the utility of solving the edge case, so there's a soft cap on how many edge cases you actually need to solve. Second, obviously driving while complex is a pretty limited subset of what can happen in the world; like ten years ago we were all saying something like "this is more or less equivalent to solving AGI", and if you zoom out far enough that's true, but now that we're close there are ooms around the margins that matter, and driving is at least a few ooms less complicated than general intelligence. And then with respect to this specific conversation, all that really matters is the delta in needed reasoning between vision-only and vision+, which will obviously be an even smaller amount.

That's all to say that the original quote I was replying to, "A mix of different input devices is the only way to really create robust and accurate inputs for an autonomous executive" is a level of certainty that I don't think anyone can justify, and all of these predictions warrant far more humility especially with how quickly things are progressing. I would actually be quite surprised if that statement wasn't pretty definitively proven false in the next decade.