Imagine your car slams on the brakes for no obvious reason. There’s no car ahead, no pedestrian, nothing.
Was it a glitch, or did it see something you didn’t?
And will it happen again at the next junction?
That’s the problem with AI driving systems. They can drive well, but they can’t tell you why they do what they do.
A new study in Nature tested a fix: a self-driving car that shows, in plain words, what it thinks is happening. It helped drivers understand the car, and predict it.
- The car’s explanations revealed hidden mistakes, including braking for a “stopped vehicle” that wasn’t there.
- People who saw the explanations got better at predicting what the car would do next.
- It didn’t make the car drive any worse, but it wasn’t a crash-safety test.
Are self-driving cars safe?
There’s no simple yes or no, because “self-driving” covers very different things.
Most cars people can buy today have driver assistance, sometimes called partial automation. The car can steer and brake, but you must watch the road and be ready to take over.
The Insurance Institute for Highway Safety (IIHS), a US safety research group, said in 2024 that there’s little evidence these systems prevent crashes.
It also rated 14 such systems on how well they keep drivers paying attention. Only one earned an “acceptable” rating, two were “marginal” and 11 were “poor”.
Fully driverless robotaxis are a separate category, operating in a few cities. Either way, the weak point is often the handover between human and machine.
That’s why this study matters. It’s about helping the human understand what the machine is thinking.
What is explainable AI?
Explainable AI means systems that can show why they made a decision, in terms people can understand.
It matters most where mistakes are costly, such as driving, medicine or lending. If you know why a system acted, you can judge when to trust it.
Many explanation tools work backwards. They look at a finished decision and guess which inputs mattered most.
The problem is that such a guess can sound convincing and still be wrong. This study took a different route, as you’ll see below.
The study: a car that shows its reasoning
Who did it
The research was led by scientists at MIT, working with Motional, a self-driving company formed by Hyundai and the car-parts maker Aptiv.
Motional funded the study, and several authors worked there. We’ll come back to what that means.
The black box problem
Modern driving systems learn from examples. This one was trained on 80 hours of expert human driving.
That kind of system can’t explain itself. Its decisions come out of millions of numbers that no human can read.
Adding plain-language concepts
The team added a new layer, called CW-Net, on top of the existing driving system. It translates the car’s situation into simple concepts, such as “Close to another vehicle” or “Approaching stopped vehicle”.
Crucially, the car’s final decision is based only on those concepts.
So the explanation isn’t a guess made afterwards. It shows what actually drove the decision.
On the dashboard, each concept appeared as a percentage. That showed how strongly the car believed each one applied.
In computer tests, adding the layer changed driving performance by less than 1%.
What happened on the test track
The team put the system in a real self-driving car, with a trained safety driver, on a private track. Three surprising moments stood out.
| Situation | What the driver thought | What the explanation revealed |
|---|---|---|
| Car kept stopping near a drop-off zone | The drop-off zone caused it | It sensed it was “close to another vehicle”, meaning parked cars |
| Car braked beside a traffic cone | The cone caused it | It believed there was a stopped vehicle ahead that wasn’t there |
| Car stopped for a cyclist | The system saw the cyclist | Its cyclist concept stayed below 1%; a backup brake did the stopping |
Testing the explanation
The team didn’t just take the explanations on trust. They tested them.
In the first case, the driver moved the car further from the parked cars. The “close to another vehicle” signal dropped, and the car started moving again.
That’s exactly what the explanation predicted, and the opposite of the driver’s first guess.
The phantom car
The second case is striking. The team removed the cone, and the car still braked at the same spot.
The explanation was right. The car was “seeing” a stopped vehicle that didn’t exist, a phantom braking problem.
The invisible cyclist
The third case is the most sobering. The main driving system had never been set up to use the cyclist data it received.
Left to itself, it would have chosen paths that hit the cyclist. The car only stopped because a separate emergency brake stepped in.
Once the safety driver noticed the explanation, they became more cautious. Later analysis showed that caution was justified.
Plain-language explanations showed drivers when a self-driving car was confused. They also helped people predict what it would do next.
Testing it on more people
Experts and ordinary users
To check it wasn’t a one-off, the team showed videos of the track events to 9 Motional experts and 30 members of the public.
People answered questions before and after seeing the explanations. Their understanding improved for 8 of the 9 experts and 27 of the 30 members of the public.
Those whose understanding improved also got better at predicting what the car would do.
Public roads in Las Vegas
Next, the system ran on public roads in Las Vegas, while a safety driver drove manually. For safety, the experimental system only ran in the background there.
That still gave real, messy scenarios to test.
In a large online study, 100 people watched replays, with 99 counted after attention checks. Half saw the car’s explanations and half saw neutral information, such as speed.
| What people had to do | Effect of explanations in surprising moments |
|---|---|
| Notice what the car was reacting to | Large improvement |
| Understand why it acted | Large improvement |
| Predict what it would do next | Medium improvement |
In unsurprising moments, the explanations didn’t make things worse. That matters, because extra information can distract.
So the explanations helped most exactly when a driver needs help most: when the car does something odd.
What earlier research found
Explanations help, if done well
A 2025 systematic review in Accident Analysis and Prevention looked at 59 studies on explanations in automated vehicles.
It found that explaining why a car acts works better than just describing what it does.
More detail didn’t always increase trust. Most studies were in simulators.
Trust depends on the machine
A 2016 meta-analysis in Human Factors pooled 30 studies on trust in automation.
Features of the automation itself had a medium-sized effect on how much people trusted it. That includes how reliable and understandable it is.
How it fits together
Past work suggested explanations could help, mostly in simulators. This study is one of the first to show useful explanations from a real self-driving car.
How much should you trust this?
Early. It’s a careful, peer-reviewed study in a top journal, but the real-world evidence is limited and the funder has a stake.
What makes it convincing
- The explanations are built into the decision itself, not guessed afterwards.
- It was tested in a real car, not only in a simulator.
- The findings held across experts, the public and a larger online study.
- The larger study used a control group that saw neutral information.
- It found and flagged a real safety gap, which is not the easy result for a funder to publish.
What makes me cautious
- The track tests involved only three observed situations with one safety driver.
- Most human testing used video replays, not live driving.
- The concept labels were often wrong: accuracy averaged 54%, and the cyclist concept barely worked.
- Motional funded the study, several authors worked there, and it has a pending patent on the method.
- Better understanding isn’t the same as fewer crashes, which the study didn’t measure.
| This study shows | This study does not show |
|---|---|
| Explanations can reveal a self-driving car’s hidden mistakes | That the car is safe |
| People predicted the car better with explanations | That drivers would react better in live traffic |
| Adding explanations didn’t hurt driving performance | That this works for every car or system |
| A cyclist gap was caught | How common such gaps are in cars on sale |
What this means for you
You can’t buy this system. But the lessons apply to any car with driver assistance.
- Driver assistance isn’t self-driving. Keep your eyes on the road and your hands ready.
- Learn what your car can’t do. Read how its lane keeping and cruise control handle junctions, cyclists and road works.
- Treat surprises as warnings. If the car brakes or steers oddly, take over and stay alert on that stretch.
- Don’t be fooled by names. Some systems sound more capable than they are.
- Check safety ratings. IIHS publishes ratings for how well these systems keep drivers alert.
In this TED-Ed lesson, Sajan Saini explains how self-driving cars sense the world around them:
What we still don’t know
- Do explanations prevent crashes? No study has shown that yet.
- Do they help in live traffic? Most testing used replays.
- Could they overload drivers? Longer use might bring fatigue or overtrust.
- Will they work for newer AI systems? The authors think so, but it needs testing.
- Would independent teams get the same result? The work was funded by an interested company.
My take: I want my car to show its working
What I like about this study is its honesty. It shows a car that was confused, and a cyclist it didn’t register, in a paper funded by the carmaker’s partner.
I’m less convinced about how far it goes. A few track moments and video replays aren’t the same as busy real roads.
But the idea feels right. If a machine can take control of a two-tonne vehicle, it should be able to tell you what it thinks it sees.
Understanding a system is the first step to knowing when not to trust it.
Paper: Explainable deep learning improves human mental models of self-driving cars
Published: Nature, 2026-09-02
Study: Engineering and human-factors study: real-car tests on a private track and public roads, plus online experiments
Who: A self-driving research car; 9 experts, 30 members of the public and 100 online participants
Funding: Motional, which employed several authors and has a pending patent on the method
Evidence: Early — promising real-car evidence, but few on-road observations, mostly video replays and a funder with a stake

Comments
No comments yet. What did you think?