Introduction
Artificial intelligence (AI) has moved rapidly from a largely technical and clinical topic to a health system and policy concern. Machine learning, natural language processing, predictive analytics, and related methods can process large and heterogeneous datasets, identify patterns that may be difficult to detect manually, and generate forecasts or classifications that can inform planning and resource allocation [1, 2]. For policymakers, the attraction of AI lies less in replacing judgment than in augmenting the speed, scale, and timeliness with which evidence can be assembled and interrogated.
The policy relevance of AI has grown alongside the wider expansion of AI in healthcare. A recent peer-reviewed review reported that Web of Science publications on AI in healthcare increased from 158 in 2014 to 731 by October 2024 and cited commercial estimates placing the 2023 global healthcare-AI market at approximately USD 11.2 billion, with rapid projected growth [3]. These figures reflect technology and research activity rather than policy adoption, and comparable statistics on government investment specifically for AI-enabled health policymaking remain limited. This distinction is important: High investment in AI does not necessarily translate into accountable, equitable, or evidence-informed policy use.
Potential policy applications include identifying emerging population needs, forecasting disease burden, targeting scarce resources, monitoring system performance, analyzing public feedback, and supporting scenario modelling [4]. At the same time, policy-level use can amplify harms if training data are incomplete or unrepresentative, if algorithms are difficult to explain, or if responsibility for AI-supported decisions is unclear [5]. Digital infrastructure, interoperability, cybersecurity, technical capacity, and sustained financing are therefore policy prerequisites rather than secondary implementation details.
Existing reviews have examined AI in healthcare broadly and AI in health policymaking more specifically. Secinaro et al. described a rapidly expanding multidisciplinary healthcare-AI literature centered on health-service management, patient data, predictive medicine, and clinical decision-making [1]. Ramezani et al. mapped 90 studies published between 2000 and 2023 using a policy-triangle framework and concluded that AI may support multiple policymaking functions, with particular emphasis on evaluation [2]. Public health reviews have also highlighted surveillance, risk prediction, and epidemic forecasting, while noting infrastructure, data, and ethical constraints [6].
The present review was designed to add a different lens: A structured evidence map organized by policy-cycle stage and AI application area, with explicit attention to where evidence is concentrated and where important gaps remain. These issues are particularly relevant to health systems undergoing digital transformation while managing fragmented information sources, variable infrastructure, and workforce constraints. In Iran, for example, big-data approaches have been discussed as a means of supporting universal health coverage and health system monitoring [7]. More broadly, Panch et al. highlighted that algorithmic bias can arise from limitations in underlying data and model development and may reproduce or amplify existing inequalities when AI is introduced into health systems [8]. Together, these concerns emphasize that the value of AI for health policymaking depends not only on technical capability, but also on data quality, representativeness, governance, and the capacity of health systems to use AI responsibly.
This review therefore addressed three questions: 1) How is AI applied across stages of the health policy cycle? 2) What benefits, risks, and implementation challenges are reported? 3) Where are the main evidence gaps, particularly for equity, governance, primary care, and resource-constrained health systems? The objectives were to map the distribution of available evidence, summarize reported opportunities and challenges, and identify priorities for policy and future research.
Materials and Methods
Study design and reporting approach
An evidence-mapping review was undertaken because the available literature spans heterogeneous study designs, AI applications, and policy functions. The purpose was to systematically identify, categorize, and display where evidence is concentrated and where gaps remain, rather than to estimate pooled effects or determine intervention effectiveness. The review was conducted using a scoping-style approach and reported with reference to preferred reporting items for systematic reviews and meta-analyses extension for scoping reviews (PRISMA-ScR) principles to improve transparency in searching, selection, data charting, and synthesis [9]. A completed PRISMA-ScR checklist is provided as
Appendix 1.
The evidence map was designed to identify concentrations and gaps in evidence across AI application areas and policy dimensions.
No protocol was prospectively registered or publicly deposited for this evidence-mapping review.
Operational definitions and review questions
For this review, AI included machine learning, deep learning, natural language processing, predictive analytics, conversational AI, and other computational approaches explicitly identified by the source as AI. Digital health technologies without an identifiable AI component (such as conventional electronic health records, telemedicine platforms, mobile-health applications, or electronic referral systems operating without AI-enabled analysis or decision support) were excluded. Clinical decision-support applications were eligible only when the source explicitly linked them to a health system or policy function.
Policy relevance required more than a general mention of policy. A source was considered policy-relevant only when the AI application was explicitly connected to at least one health-policy or health-system function: Agenda setting or priority identification; policy or programme formulation/design; resource allocation; implementation; monitoring/evaluation; regulation or governance; or equity-related decision-making. Studies describing AI solely for individual-level diagnosis, prognosis, or treatment without an explicit health-system or policy connection were excluded.
Each included source was coded against the policy-cycle stage or stages explicitly addressed in its objectives, methods, results, or policy implications. Sources could receive more than one policy-stage code when they substantively addressed multiple stages. Codes were therefore non-mutually exclusive and were assigned on the basis of the reported function of AI rather than a simple mention of a policy-stage term.
Search strategy
PubMed, Scopus, and Web of Science were systematically searched for English-language publications published between January 2015 and August 2026. The search was conducted on 30 August 2026 using the same core search concepts and eligibility criteria. Search terms combined concepts relating to AI with terms relating to health policy, governance, decision-making, agenda setting, policy formulation, implementation, monitoring, and evaluation. Google Scholar was used as a supplementary source for citation chasing and identification of potentially relevant publications and was not included in the database-specific record counts reported in the PRISMA flow diagram. Search Strategy is presented in
Appendix 2.
No dedicated systematic search of organizational grey-literature repositories, government or ministry websites, international agency repositories, or policy think-tank databases was undertaken. Institutional reports identified through database searching, Google Scholar, or citation chasing were considered when they met the eligibility criteria. Google Scholar was used as a supplementary source to identify potentially relevant publications and citation leads; it was not treated as a separate database for the database-specific record counts reported in
Figure 1.
Because individual sources could contribute to more than one benefit theme, the supporting-source counts are descriptive and non-mutually exclusive.
Table 3 shows across the included literature, the potential benefits of AI were closely linked to the health system context in which it was implemented.

Reported opportunities included applications in primary health care, public health decision-making, organizational implementation, system monitoring, prediction, service delivery, and workflow improvement. However, these benefits were generally dependent on adequate data infrastructure, appropriate governance, workforce capacity, validation procedures, and ongoing monitoring to support safe and sustainable implementation [36, 41, 43, 44, 49].
Reported challenges and risks
Five recurring areas of concern were identified across the included literature: bias and inequity, data privacy and security, infrastructure and workforce limitations, ethical and accountability gaps, and regulatory uncertainty. These themes were reinforced by more recent studies, particularly in relation to governance, organizational readiness, equity, lifecycle oversight, and regulatory requirements. Because individual sources could contribute to more than one challenge theme, the supporting-source counts are non-mutually exclusive (
Table 4).

Across the included literature, governance-related challenges were prominent, particularly those concerning algorithmic bias, data privacy and cybersecurity, workforce and infrastructure readiness, accountability, and regulatory oversight. These challenges were often interrelated, as limitations in data quality and infrastructure could affect both technical performance and equity, while weak governance arrangements could increase concerns related to transparency, liability, privacy, and public trust. The evidence also indicates that responsible AI implementation requires continuous monitoring and oversight throughout the AI lifecycle, particularly as models, data environments, and regulatory requirements evolve over time [13, 14, 26, 33, 35, 36, 45-47].
Evidence gaps
Several gaps remained apparent across the updated evidence map. First, the literature continued to be concentrated in the formulation/design and implementation stages of the policy cycle. Formulation/design accounted for 41 coded sources, and implementation accounted for 42, compared with 18 sources addressing monitoring/evaluation and 11 addressing agenda setting. Although recent studies have expanded attention to governance, organizational readiness, regulation, and lifecycle oversight, comparatively less evidence examines how AI influences the initial setting of health policy priorities or how AI-supported policies perform after sustained implementation.
Second, equity was explicitly represented in 19 sources, but much of this evidence addressed equity through concerns about algorithmic bias, fairness, access, representation, and governance. Empirical evidence demonstrating whether AI-supported policies reduce, maintain, or widen health inequalities in routine health-system settings remains limited. This distinction is important because identifying equity risks does not demonstrate the distributional effects of AI after implementation.
Third, evidence from prospective and sustained real-world implementation remained relatively limited. The updated literature includes several organizational governance frameworks, implementation case studies, policy analyses, and systematic reviews, but comparatively few studies evaluate long-term adoption, sustainability, costs, unintended consequences, or measurable policy outcomes. Recent implementation-focused studies have begun to address these issues, particularly in relation to organizational governance and post-deployment monitoring, but the evidence base remains at an early stage [25, 26, 35, 36, 43, 44].
Fourth, representation of resource-constrained and underserved settings has improved but remains limited relative to the evidence from higher-resource health systems. The included literature now covers resource-poor and rural settings, Association of Southeast Asian Nations (ASEAN) countries, Iran, and selected Middle Eastern contexts [14, 15, 17, 19, 28, 47]. However, empirical evidence on sustained implementation, scale-up, financing, workforce requirements, interoperability, and governance in low-resource settings remains comparatively sparse.
Evidence in some settings and policy areas should also be interpreted cautiously because the review did not include a dedicated systematic search of grey-literature repositories, government websites, or policy think-tank databases. Government evaluations, implementation reports, and other policy documents that are not indexed in bibliographic databases may therefore be underrepresented.
Finally, evidence related to primary health care and community-based applications has increased in the updated literature, including recent work on AI implementation and future health system preparedness in primary care [14, 15]. Nevertheless, empirical evidence on community-level implementation, social determinants of health, patient and public participation, and the longer-term effects of AI-supported policies on access and service delivery remains limited. Future research should therefore prioritize prospective implementation studies that examine feasibility, adoption, sustainability, cost, equity, governance, and policy outcomes across diverse health-system contexts.
Discussion
This evidence-mapping review provides a descriptive overview of how AI has been examined across health policy functions and should not be interpreted as an assessment of comparative effectiveness. The updated evidence base indicates a field that is expanding rapidly but remains uneven in its maturity. Earlier literature largely positioned AI as an analytic tool for processing large datasets, supporting prediction, surveillance, and decision-making [1, 2, 4]. More recent studies have broadened this focus toward governance, organizational readiness, regulatory oversight, lifecycle monitoring, equity, and implementation within health systems [12-15, 25, 26, 28, 33-38, 43-47]. Across the mapped evidence, formulation/design and implementation remained the most frequently represented policy-cycle stages, while agenda setting and monitoring/evaluation were less prominent. This pattern suggests that considerable attention is being given to how AI should be designed, governed, and introduced into health systems, but comparatively less empirical evidence is available on how AI shapes policy priorities or performs over sustained periods after implementation.
The smaller evidence base identified in this review compared with the 90 studies reported by Ramezani et al. [2] reflects important differences in scope, time period, eligibility criteria, and evidence classification. Ramezani et al. examined studies published from 2000 to 2023 and considered a broad range of AI applications using a policy-triangle framework. The present review focused on literature published from 2015 through August 2026 and required an explicit connection between the AI application and a defined health policy or health system function. Purely clinical studies concerned only with diagnostic or treatment performance were excluded when no wider policy or system relevance was evident. The present evidence maps also included a heterogeneous range of source types, including empirical studies, reviews, policy and ethical analyses, governance frameworks, guidance documents, and selected conference papers. The resulting 46-source evidence base should therefore be interpreted as a focused map of policy-relevant evidence rather than a comprehensive inventory of all healthcare AI applications.
The apparent underrepresentation of some policy areas and settings should also be interpreted cautiously. Evidence from resource-constrained and underserved settings has increased, with studies addressing rural healthcare, resource-poor settings, ASEAN countries, Iran, primary health care, and selected Middle Eastern contexts [14, 15, 17, 19, 28, 47]. Nevertheless, this literature remains less extensive than evidence from higher-resource health systems. The observed gap may reflect both a genuine lack of sustained implementation research and limitations in evidence retrieval. Restriction to English-language publications, reliance primarily on indexed bibliographic databases, and the absence of a dedicated systematic search of grey literature may have reduced the visibility of government reports, programme evaluations, and locally produced policy documents from resource-constrained settings.
The practical value of AI applications depends on balancing potential benefits with corresponding risks. Data-driven decision support can make service gaps, population-health patterns, and resource needs more visible, while predictive modelling may support earlier action and planning. Digital and conversational tools may extend access, and automation may improve selected workflows. However, these benefits depend on the quality, representativeness, and interoperability of underlying data and on the ability of organizations to integrate AI into existing systems and decision processes. Incomplete data, model drift, limited external validity, and poorly communicated uncertainty can undermine apparently precise outputs. Likewise, digital tools may widen rather than narrow inequalities when access differs by language, literacy, connectivity, disability, socioeconomic position, or trust [12, 14, 16, 19, 29, 30, 32, 34].
The updated literature places particularly strong emphasis on governance throughout the AI lifecycle. Recent frameworks and implementation studies highlight the need for clear accountability, transparent decision rules, appropriate validation, human oversight, fairness assessment, privacy and cybersecurity safeguards, and post-deployment monitoring [25, 26, 33-36, 43-47]. This represents an important development from earlier conceptual discussion toward more operational approaches to AI governance. Nevertheless, much of this work still consists of frameworks, organizational case studies, policy analyses, or early implementation experience. Evidence showing whether specific governance arrangements produce better policy outcomes, reduce harms, or improve equity over time remains limited.
The growing focus on organizational readiness also highlights that AI implementation is not primarily a technical issue. Recent studies consistently identify workforce capability, leadership, infrastructure, interoperability, financing, governance arrangements, and implementation processes as important conditions for sustainable adoption [14, 15, 25, 26, 28, 33, 43, 44]. This is particularly relevant in health systems where electronic information systems remain fragmented or where technical and financial capacity is limited. Introducing AI without strengthening underlying data systems and organizational capacity may reproduce existing weaknesses rather than overcome them.
Primary health care is beginning to receive greater attention in the recent literature. Studies addressing AI implementation and future preparedness in PHC identify opportunities for decision support, screening, service coordination, monitoring, and resource planning, while also emphasizing stewardship, infrastructure, financing, workforce readiness, and governance [14, 15]. However, evidence on community-based implementation, public participation, social determinants of health, and longer-term effects on access and service delivery remains limited. This represents an important area for future implementation research, particularly because policy decisions in PHC directly influence population-level access and equity.
Evidence maps are useful in this context because a larger concentration of publications should not be interpreted as demonstrating effectiveness or stronger evidence. Similarly, a sparsely populated area may represent a genuine research gap, limitations of the search strategy, or differences in how policy functions are described and indexed. The present findings therefore characterize where research attention has been concentrated and where further empirical evaluation is needed rather than ranking AI applications according to effectiveness or methodological quality.
Implications for health systems and policy
The findings suggest that responsible adoption of AI in health policymaking requires attention to the wider health system environment in which these technologies are introduced. Reliable and interoperable health information systems are important foundations for policy-facing AI, particularly where applications depend on routinely collected population or service-delivery data. In settings with fragmented information systems, variable data quality, or limited technical capacity, introducing AI without first addressing underlying information-system weaknesses may reduce rather than improve the reliability of policy decisions.
Initial implementation may therefore be most appropriate for clearly defined and auditable functions in which AI supports rather than replaces human decision-making. Potential applications include identifying geographic or socioeconomic gaps in service coverage, forecasting service demand, strengthening public-health surveillance, supporting resource planning, and prioritizing information for human review. Such applications require appropriate validation, transparent decision rules, defined accountability, and mechanisms through which AI-supported decisions can be reviewed, challenged, or modified.
Governance should extend throughout the AI lifecycle, from problem definition and data acquisition to implementation, monitoring, modification, and eventual withdrawal. Particular attention is required for data quality and representativeness, privacy and cybersecurity, subgroup performance, transparency, human oversight, accountability, and mechanisms for responding to performance deterioration or unintended consequences. Equity should be incorporated throughout this process because differences in representation, digital access, language, socioeconomic position, geography, and disability may influence who benefits from AI-supported policies.
Implementation research should accompany real-world deployment. Evaluations should move beyond technical accuracy to examine feasibility, acceptability, adoption, cost, workflow effects, sustainability, unintended consequences, equity, and policy outcomes. Longer-term monitoring is particularly important because AI models, underlying datasets, organizational environments, and regulatory expectations may change after deployment. Greater attention is also needed in public health, primary care, community-based, and resource-constrained settings, where evidence remains comparatively limited despite their importance for population health and health equity.
Implications for future research
Future research should move beyond descriptive, conceptual, and early implementation studies toward prospective and longitudinal evaluations of AI-supported policy and health system applications under routine conditions. Priority questions include whether AI improves policy decisions, resource allocation, service delivery, and population health outcomes compared with existing approaches; whether these benefits are sustained over time; and whether implementation reduces or widens inequalities between population groups. Further research should evaluate which governance arrangements are most effective in supporting transparency, accountability, human oversight, privacy, and appropriate responses to model or system failures.
Greater attention is also needed to the organizational conditions required for sustainable implementation, including interoperable data infrastructure, workforce capability, financing, leadership, procurement, and integration with existing workflows. Research remains particularly important in primary care, community-based services, and resource-constrained settings, where evidence on feasibility, scale-up, sustainability, cost, public participation, and equity outcomes remain comparatively limited. Future studies should also examine AI systems across their full lifecycle, including post-deployment performance, unintended consequences, model drift, modification, and withdrawal.
Conclusion
The updated evidence map indicates that AI-related health policy research is increasingly concentrated in formulation/design, implementation, governance, and organizational readiness, with growing attention to regulation, equity, and lifecycle oversight. However, agenda setting, sustained monitoring and evaluation, empirical equity outcomes, and evidence from resource-constrained and community-based settings remain less developed.
The findings suggest that the policy challenge is no longer simply whether AI can support health system decision-making, but how it can be introduced, governed, evaluated, and sustained responsibly. Realizing potential benefits will require reliable and interoperable data systems, appropriate workforce and organizational capacity, transparent governance, human oversight, equity-sensitive implementation, and continued evaluation after deployment. Because this review maps the distribution and characteristics of the literature rather than its comparative effectiveness, stronger prospective evidence is still needed before broad policy conclusions about AI’s impact can be drawn.
Policy Implications
Implications directly supported by the mapped evidence
AI initiatives should be accompanied by investment in data quality, interoperability, cybersecurity, technical infrastructure, and workforce capability, as these conditions repeatedly influence implementation readiness and sustainability.
Equity and representativeness should be considered throughout the AI lifecycle, including data selection, validation, implementation, and post-deployment monitoring, because bias and unequal digital access may affect who benefits from AI-supported policies.
Policy-facing AI requires clear governance and accountability arrangements, including transparent decision processes, appropriate validation, privacy safeguards, human oversight, auditability, and mechanisms for reviewing or challenging consequential decisions.
Implementation should include ongoing monitoring rather than one-time assessment, because model performance, underlying data, organizational conditions, and regulatory requirements may change after deployment.
Decisions about wider scale-up should be informed by real-world evaluation of feasibility, adoption, cost, sustainability, unintended consequences, equity, and policy outcomes, rather than technical performance alone.
Priorities for health systems
Developing governance arrangements that address the full AI lifecycle, including procurement, data governance, validation, implementation, equity assessment, monitoring, incident management, modification, and withdrawal.
Prioritizing AI applications that address clearly defined health-system needs, such as public health surveillance, service demand forecasting, resource planning, primary care support, and identification of inequities in service coverage.
Integrating AI initiatives with broader health information-system strengthening so that AI does not reproduce existing problems arising from fragmented, incomplete, or poorly interoperable data.
Building interdisciplinary capacity among policymakers, public health practitioners, clinicians, health system managers, data scientists, legal and ethics specialists, and community representatives, with opportunities for public communication and stakeholder participation.
Generating more implementation evidence from primary care, community-based, underserved, and resource-constrained settings to ensure that future AI policy is informed by a wider range of health system contexts.
Limitations
This review has several limitations. First, no formal methodological quality or risk-of-bias appraisal was undertaken. The evidence map therefore describes the distribution, characteristics, and thematic content of the literature rather than the strength of evidence or effectiveness of individual AI applications. Areas with a larger number of mapped sources should not be interpreted as having stronger or more reliable evidence than areas with fewer sources. Second, the search was restricted to English-language publications, which may have reduced the representation of evidence from non-English-speaking and resource-constrained settings. Third, no dedicated systematic search of organizational grey-literature repositories, government websites, ministries of health, or policy think tanks was undertaken. Although some institutional guidance and policy reports were identified through database searching and supplementary citation searching, relevant policy documents and implementation reports may have been missed. Fourth, the included evidence was heterogeneous and comprised empirical studies, reviews, governance frameworks, policy and legal analyses, guidance documents, perspectives, and implementation case studies. This diversity limits direct comparison across sources and restricts conclusions about the effectiveness or long-term impact of AI-supported policy interventions. Fifth, policy-cycle stages were coded non-mutually exclusively because individual sources could address more than one policy function. For the quantitative evidence map, each source was assigned to a dominant AI application area to improve interpretability; however, this classification may simplify studies that addressed multiple AI functions simultaneously. Finally, areas that appear sparsely populated in the evidence map may reflect not only genuine evidence gaps but also the search strategy, language restriction, eligibility criteria, indexing practices, and coding framework. The findings should therefore be interpreted as a map of the available and retrievable evidence rather than a complete representation of all AI-related health policy activity.
Ethical Considerations
Compliance with ethical guidelines
This article is a review paper with no human or animal sample.
Funding
This research received no specific grant from any funding agency in the public, commercial or not-for-profit sectors.
Authors contributions
Conceptualization, data analysis, and writing the original draft: Marziyeh Najafi and Roya Rajaee; Methodology: Marziyeh Najafi and Anahita Behzadi; Investigation: Roya Rajaee, Tejas Wakde, and Mohammad Barzegar Rahatlou; Review & editing: Roya Rajaee, Anahita Behzadi, Leila Vali, Tejas Wakde, and Mohammad Barzegar Rahatlou; Supervision: Roya Rajaee and Leila Vali; Final approval: All authors.
Conflict of interest
The authors declared no conflict of interest.
Acknowledgements
The authors would like to thank everyone who contributed their time, support, and valuable input during the development and preparation of this manuscript.