UAE Researchers Find Leading AI Models Fail Most Autonomous Drone Decisions

2 Min Read

Researchers in the UAE have introduced PhysAI-Bench, a benchmark designed to test whether AI models can make reliable real-time decisions as autonomous drone pilots.

Developed by United Arab Emirates University with Khalifa University’s Digital Future Institute and Abu Dhabi’s Technology Innovation Institute, the benchmark evaluated 29 models from 14 AI organisations. It used 10,178 decision points drawn from real UAV mission traces, covering mission objectives, physical constraints, sensor readings, tool calls and simulated 6G network conditions.

OpenAI’s GPT-5.3 Chat achieved the highest score, correctly selecting the next action in 52% of 500 previously unseen test questions. GPT-5.2 Chat followed at 49.40%, while xAI’s Grok 4.5 scored 49.07%.

All 29 models performed worse on unseen questions than on the development data used to tune them, with performance gaps ranging from about 2% to 26%. The findings also suggest that model size alone does not determine performance. Some smaller open-weight models matched or outperformed larger systems while operating more efficiently.

The researchers said the benchmark could eventually be extended to robotics, autonomous vehicles and industrial automation. The work builds on UAVBench, a 2025 dataset of 50,000 validated drone flight scenarios developed by UAE University and Khalifa University.

Source: Middle East AI News

Share This Article