
One of the hardest things to prove about autonomous driving is that it actually works. It is easy to show a video of a car negotiating traffic, making a U-turn, or finding a parking space. It is tougher to put a passenger in the car, take away the obvious clues, and ask them a simple question – who’s driving?
That was exactly the idea behind XPENG’s “Guess Who’s Driving” activity when we went to Guangzhou. During our recent visit to XPENG headquarters, we were given the opportunity to experience the company’s VLA 2.0 intelligent driving system in the XPENG L03 on public roads around the company’s headquarters.

Before we got into the cars, XPENG gave us a brief explanation of what makes VLA 2.0 different. VLA stands for Vision, Language and Action. Instead of relying heavily on high-definition maps or sophisticated sensors such as LiDAR, the system uses the vehicle’s cameras and its proprietary hardware and software to understand what is going on around it, then makes decisions and translates those decisions into driving action.

But the real demonstration came when I got into the L03. The instructor in the car with me was a professional driver, so I assumed he would do most of the driving. There was a barrier between me and the instructor, and the challenge was to figure out when he was actually driving and when the car had taken over. We headed out onto the public roads surrounding XPENG’s headquarters. We accelerated, slowed down, followed traffic, and negotiated the road as you would expect from a competent human driver. Nothing felt particularly unusual. That was the point.
At one point, he asked me to guess who was driving. My first instinct was to look for the subtle clues. How was the steering? How was the braking? Did the car hesitate before crossing an intersection? Honestly, there was very little to give anything away.

Eventually, I was told that VLA 2.0 had been doing most of the driving since we left XPENG’s HQ. And that was probably the most impressive part of the demonstration. There was no sudden change in driving style that announced that artificial intelligence had taken over or was in control. No exaggerated steering corrections. No awkward braking. It simply drove. Like a human.
XPENG says the system can process what it sees around it and turn that information into driving actions with inference latency of less than 80 milliseconds. The objective is to make the car’s reactions quick, decisive, and, importantly, natural.
After experiencing the system as a passenger, I was given the opportunity to sit in the driver’s seat and experience the same technology firsthand. This changed my perspective completely. From the driver’s seat, I was aware of what the vehicle was doing. I watched the steering wheel move, felt the accelerator and brake inputs, and saw how the car responded to traffic around it.

And again, the biggest surprise was how normal and uneventful it felt. That may not sound like a compliment for an advanced autonomous-driving system, but it is probably the best one you can give it. The goal of autonomous driving shouldn’t be to constantly remind you that a computer is driving. It should simply make the journey feel natural or even ordinary.
XPENG’s “Guess Who’s Driving” exercise was more than a clever demonstration. It was a fun way of showing what XPENG believes VLA 2.0 can bring to everyday driving. After experiencing it from the passenger seat and then from behind the wheel, I came away with a slightly different view of autonomous driving. The real breakthrough may not be when the car drives itself. It may be when you don’t notice that it is.




