News
Date:2026-07-17
Views:522
An international team led by Professor Bin Sheng argues that medical AI should be assessed not only by predictive performance, but also by whether it can justify recommendations, respect clinical constraints and remain auditable in real-world use.
16 JULY 2026 | SCHOOL OF COMPUTER SCIENCE, SHANGHAI JIAO TONG UNIVERSITY
A multidisciplinary team led by Professor Bin Sheng of the Ministry of Education Key Laboratory of Artificial Intelligence, School of Computer Science, Shanghai Jiao Tong University (SJTU), has published a Comment titled “Neuro-symbolic artificial intelligence in medicine” in Nature Biomedical Engineering. Published online on 10 July 2026, the article examines how neuro-symbolic artificial intelligence (NeSy AI) could bridge the gap between strong predictive performance and clinically accountable decision-making.
The Comment is a conceptual synthesis and roadmap rather than a report of a new clinical model. It compares neural, symbolic and hybrid AI paradigms; identifies the settings in which NeSy AI may offer the greatest value; and sets out the technical, governance, human-factors and evaluation priorities required for responsible clinical deployment.

The Comment “Neuro-symbolic artificial intelligence in medicine” was published online in Nature Biomedical Engineering on 10 July 2026. Image source: Nature Biomedical Engineering.
Deep learning and large language models have achieved human-level performance in selected biomedical tasks, but their internal reasoning can remain opaque and their errors difficult to anticipate or audit. In medicine, these limitations are not merely technical inconveniences: a confident but unsupported recommendation can affect clinical decisions, patient safety and public trust.
“Performance alone is insufficient for clinical trust, accountability and safe deployment.”
The authors therefore frame the central question differently. The issue is no longer only whether AI can recognize a pattern at expert level, but whether it can explain the basis for a recommendation, comply with explicit safety boundaries and behave robustly when patient populations, clinical settings or available data change. The paper uses a simple clinical analogy: purely neural AI can resemble a brilliant but erratic intern, symbolic AI a rigid textbook, and a well-designed hybrid an experienced clinician who combines perceptual skill with explicit knowledge and constraints.
NeSy AI brings together two complementary capabilities. Neural models extract signals from unstructured clinical notes, medical images and physiological waveforms. Symbolic components encode clinical guidelines, causal or physiological relationships, contraindications and other formal constraints. This is not a retreat from data-driven learning; it is an attempt to make the decision process more structured, inspectable and accountable where those properties matter.
The Comment distinguishes two main designs. In composite systems, neural and symbolic components remain separate and communicate through defined interfaces, allowing a symbolic back end to check candidate outputs. In monolithic systems, logical structure is integrated directly into training or inference. Composite designs may offer modularity and clearer audit trails, while tighter integration may reduce latency; the right choice is therefore a deployment decision as much as a modelling preference.

Figure 1. Two ways to build neuro-symbolic medical AI. A composite design passes neural outputs to a separate symbolic reasoner, whereas a monolithic design integrates neural perception and symbolic reasoning within one system. Source: Wong et al., Nature Biomedical Engineering (2026), Fig. 1. DOI: 10.1038/s41551-026-01728-1.
Two potential clinical benefits follow. First, a NeSy system can expose an audit trail linking an output to the evidence and rules used during inference. For example, it could show that tachycardia and immobilization extracted from a note triggered criteria for further assessment of pulmonary embolism. Second, a symbolic knowledge base could encode a maximum safe dose or an absolute contraindication and intercept an out-of-bounds recommendation. These examples illustrate design possibilities rather than a guarantee of correctness or a deployed system reported in the Comment.
Two potential clinical benefits follow. First, a NeSy system can expose an audit trail linking an output to the evidence and rules used during inference. For example, it could show that tachycardia and immobilization extracted from a note triggered criteria for further assessment of pulmonary embolism. Second, a symbolic knowledge base could encode a maximum safe dose or an absolute contraindication and intercept an out-of-bounds recommendation. These examples illustrate design possibilities rather than a guarantee of correctness or a deployed system reported in the Comment.
The authors explicitly caution against applying NeSy AI everywhere. Its strongest case is in safety-critical, context-sensitive tasks that depend heavily on statistical inference, including treatment planning, sepsis management, multimodal decision support and personalized medicine. By contrast, the added complexity may contribute little to low-risk tasks such as routine documentation or note summarization, while straightforward rule-based systems may already be sufficient for functions such as immunization reminders.

Figure 2. NeSy AI is most warranted for tasks that are both high-risk and heavily dependent on statistical inference. Simpler neural or rule-based approaches may be more appropriate elsewhere. Source: Wong et al., Nature Biomedical Engineering (2026), Fig. 2. DOI: 10.1038/s41551-026-01728-1.
The Comment also makes clear that symbolic rules do not automatically make AI safe. A rule set can be incomplete, outdated or biased, and may silently validate a recommendation that no longer reflects current practice. Medical knowledge should therefore be treated as governed clinical infrastructure: machine-readable, versioned, traceable to its source, and supported by clear clinical ownership over how rules are created, validated, updated and retired.
Constraint design must also preserve clinical judgement. Hard constraints should be reserved for non-negotiable safety invariants, such as absolute contraindications, while guideline preferences can be represented as soft constraints that clinicians may override with appropriate justification. Explanations should adapt to clinical role and context so that transparency does not become cognitive overload, alert fatigue or a new source of automation bias.
Evaluation must consequently move beyond predictive accuracy. The authors call for studies that measure the frequency and severity of safety-rule violations, whether symbolic overrides improve outcomes, performance after rule updates or distribution shift, latency, clinician workload, workflow integration and subgroup disparities. Because most medical NeSy systems remain at proof-of-concept stage, prospective real-world clinical validation is a central unmet need.
The proposed roadmap begins with near-term priorities such as rule-based safety checks, auditable decision pipelines and regulatory alignment. The mid-term agenda includes lightweight reasoning, context-aware constraints and neuro-symbolic digital twins that can support personalized disease trajectories and resource-constrained settings. Over the longer term, the authors envision systems with meta-cognitive capabilities that monitor uncertainty, detect distribution shift and defer or escalate when their own reasoning becomes unreliable.

Figure 3. The roadmap progresses from safety checks and auditable decision pipelines, through context-aware systems, toward AI that can monitor uncertainty and escalate when unreliable. Governance, human factors and evaluation span every stage. Source: Wong et al., Nature Biomedical Engineering (2026), Fig. 3. DOI: 10.1038/s41551-026-01728-1.
NeSy AI is not presented as a universal solution. Its promise depends on rigorous validation, interoperable knowledge governance and careful integration into clinical workflows. Nevertheless, by connecting scalable learning with explicit constraints and governed deployment, the Comment shifts the conversation from whether medical AI can produce an answer to whether that answer can be examined, maintained and responsibly acted upon.
The article was authored by Matthew Yu Heng Wong, Jasmine Chiat Ling Ong, Nigam Shah, David C. Klonoff and Bin Sheng. The collaboration brought together researchers from Shanghai Jiao Tong University, the University of Cambridge, Singapore General Hospital, Duke-NUS Medical School, Stanford University and the Diabetes Research Institute at Mills-Peninsula Medical Center (Sutter Health). Wong and Sheng co-conceived the project; Wong drafted the paper, and Sheng supervised the work and served as corresponding author.
The work reflects the School of Computer Science’s continuing commitment to interdisciplinary research at the intersection of artificial intelligence, clinical medicine and responsible technology. It was supported by the National Natural Science Foundation of China (grant T2525004). The authors declared no competing interests.
ARTICLE INFORMATION
Citation: Wong, M. Y. H., Ong, J. C. L., Shah, N., Klonoff, D. C. & Sheng, B. Neuro-symbolic artificial intelligence in medicine. Nature Biomedical Engineering (2026).
Original article: View the published article on Nature Biomedical Engineering
DOI: 10.1038/s41551-026-01728-1
Corresponding author: Professor Bin Sheng, shengbin@sjtu.edu.cn