New research maps the security and ethical risks created when vision-language models guide these embodied systems, where a mistaken description or manipulated command can become a physical action.
Research connects failures across perception, planning, instruction following and human-robot interaction, covering hallucinations, synthetic forgeries, adversarial attacks, privacy leakage and unsafe execution.
It also shows how the same multimodal capabilities can support defense through contextual checking, forgery detection, privacy protection and risk-aware reasoning, offering a roadmap toward embodied intelligence that is not only capable, but dependable in real-world settings.
Vision-language models (VLMs) connect images with text, while vision-language-action models (VLAs) extend that connection to robot plans and control signals. This makes natural-language instruction, scene understanding and flexible task execution possible, but it also creates a chain of dependency: Flawed data can distort perception, weak visual-language alignment can produce hallucinations, and malicious inputs can redirect decisions.
Results of Errors
In a chatbot, such errors may generate misinformation; in an autonomous vehicle or industrial robot, they may lead to collisions, damaged equipment or failed missions.
Existing safeguards are often benchmark-specific, fragmented across system layers or too computationally costly for real-time use. Because of these challenges, needs to occur to gain a deeper understanding into unified, adaptive safeguards for multimodal agents operating under uncertain physical conditions.
Researchers from the Institute of Automation, Chinese Academy of Sciences; University College London (UCL); Minzu University of China; and the China Academy of Electronics and Information Technology conducted the research.
The team examined how VLMs and VLAs end up used in embodied intelligence (EI), organized the major security threats and defensive approaches, and connected technical safety with accountability, fairness, privacy, environmental sustainability and human oversight.
The survey first tracks VLM and VLA use across four linked functions: Perception, planning, instruction following and human-robot interaction (HRI).
It then shows how failures can cascade. Biased training data, weak visual encoders or poor cross-modal alignment can make a model describe objects that are not present.
Additionally, forged traffic signs, altered labels, cloned voices or deceptive captions can misguide perception and planning. Tiny adversarial perturbations, hidden backdoor triggers and multimodal jailbreak prompts may bypass safety controls, while persistent sensing can expose identity, location, possessions and social behavior.
The authors organize countermeasures into equally connected layers. These include hallucination filtering and vision-grounded alignment; cross-modal forgery detection, watermarking and provenance tracing; defenses against perturbations, backdoors and jailbreaks; differential privacy (DP), secure multi-party computation (SMPC) and homomorphic encryption (HE); and safeguards for navigation, communications and physical control.
Not One Filter Reigns
A further strand uses causal explanations, intent alignment and risk assessment so robots can interpret ambiguous instructions, anticipate hazards and correct actions. The review’s central insight is that no single filter can secure an embodied agent: Protection must follow the entire path from sensor input to model reasoning, system architecture and physical execution.
The authors said the central challenge is not simply making models more accurate, but ensuring a system remains safe when its sensors, language inputs and operating conditions are imperfect.
They said defenses should end up combined rather than deployed as isolated patches, with transparent risk metrics, continuous monitoring and human oversight for critical decisions. A trustworthy robot must also explain what it is doing, recognize when it is uncertain and fall back safely instead of acting with false confidence. The authors added that technical progress must move alongside privacy protection, fairness, accountability and responsible governance.
For developers and regulators, the survey provides a practical checklist for evaluating embodied systems before large-scale deployment.
Future platforms could combine interpretable reasoning, attack detection, privacy-preserving computation and dynamic safety controls under reproducible, open evaluation protocols.
To that end, the authors call for designs that address four dimensions together: Technical robustness, regulatory alignment, social equity and environmental sustainability. Such an approach could support safer autonomous transport, healthcare assistance, warehouse automation, industrial inspection and collaborative robotics, while making responsibility easier to trace when failures occur.
Furthermore, the review also warns that strong laboratory results may not transfer cleanly to noisy, culturally diverse and resource-constrained environments. Progress will therefore depend on cross-disciplinary cooperation and testing that measures not only task success, but safe behavior under stress.
Click here to view the paper.

