Alexandre Tran, David Granton, Eddy Fan, Bram Rochwerg
Clinical practice guidelines (CPGs) are used by critical care clinicians to guide practice and inform best care. According to the GRADE framework, evidence synthesis should preferentially rely on randomized controlled trials (RCTs) because they minimize bias and establish causality.1 Despite challenges, critical care is well-suited to randomized studies given its (1) high incidence of acute conditions, (2) protocolized interventions, (3) standardized outcomes, and (4) strong data infrastructure and trial networks.2,3 Despite the advantages, RCTs are often unavailable for CPGs or leave knowledge gaps, particularly for subgroup effects or patient-important outcomes like long-term quality of life related to heterogeneous populations, urgent interventions, and recruitment constraints.4,5 Physicians may hesitate to apply RCT results because (1) enrolled patients differ from real-world populations, (2) key outcomes may be unmeasured, (3) effect estimates may be imprecise, and (4) subgroup analyses may be lacking.6 When RCT evidence is insufficient, high-quality non-randomized studies of interventions (NRSI) can complement trials by approximating causal inferenceâestimating exposure effects while separating systematic bias from random error.7 High-quality NRSI Ârequire large, well-validated datasets with minimal missingness and adequate temporal resolution. Without these, even advanced analytics cannot yield credible estimates. NRSI often emulate target trials, aligning eligibility, time zero, and predefined interventions and outcomes.8,9 Design must reflect strong knowledge of confounders and time-varying biases, addressed through advanced data and statistical methods. When based on explicit and credible assumptions (eg, exchangeability, no residual confounding), NRSI can yield valid and generalizable estimates, though such assumptions cannot be proven and still require caution in interpretation.10 Most NRSI are retrospective and lack safeguards standard in RCTs such as trial registration or prespecified outcomes. In target-trial emulation (Table 1), preregistration before data access is critical to prevent selective reporting and analytic flexibility, mirroring RCT practice. These limitations are especially relevant in critical care, given dynamic physiology, urgent decisions, and substantial clinical heterogeneity. These factors complicate exposure timing, increase time-varying confounding, and challenge stability assumptions in target-trial designs. Rigorous cohort definition and analytic strategy are essential when applying NRSI in this context. As causal-inference methods such as target-trial emulation spread, cautious application with methodological rigor and transparency is essential to avoid poorly executed, misleading, or irreproducible NRSI. High-quality NRSI depend not only on analytical sophistication but also on careful data acquisition, explicit protocolization, and transparency in prespecifying exposures, outcomes, and analytic plansâprinciples that mirror RCT standards. Target trial (ideal RCT) versus emulation. 1. Treatment with ECMO therapy if PaO2/FiO2 < 80 mmHg 2. Treatment with conventional mechanical ventilation without the use of ECMO therapy Adapted from: National Academies of Sciences, Engineering, and Medicine; Health and Medicine Division; Board on Health Care Services; Committee on Developing a Protocol to Evaluate the Concomitant Prescribing of Opioids and Benzodiazepine Medications and Veteran Deaths and Suicides. An Approach to Evaluate the Effects of Concomitant Prescribing of Opioids and Benzodiazepines on Veteran Deaths and Suicides. Washington (DC): National Academies Press (U.S.); 2019 Sep 24. 2, Specifying the Target Trial. Available from: https://www.ncbi.nlm.nih.gov/books/NBK547516/. Case example: Venovenous extracorporeal membrane oxygenation in patients with acute covid-19 associated respiratory failure: comparative effectiveness study.22 This commentary examines the evolving role of NRSI in developing critical care CPGs. We outline key challenges in conducting and synthesizing critical care research, then describe how high-quality NRSI can complement randomized evidence by (1) aligning effect estimates with RCTs, (2) informing certainty of evidence (CoE), and (3) guiding clinical practice recommendations. We propose practical strategies for CPG panels and domain experts to maximize the utility of NRSI while maintaining methodological rigor. Our goal is to support CPG panelists, researchers, and clinicians in interpreting recommendations that integrate NRSI. These recommendations align with evolving GRADE guidance, operationalizing its principles for critical care applications. GRADE provides a structured approach for rating CoE, the confidence that an estimated effect is close to the truth.11 When Âdeveloping guidelines, the GRADE Evidence-to-Decision (EtD) framework translates synthesized evidence into recommendations by weighing intervention effects, CoE, patient-valued outcomes, and contextual factors such as resource use, equity, acceptability, and feasibility.12 These contextual judgments ensure that evidence is interpreted through a patient- and system-centered lens, recognizing that even high-certainty data require value-based consideration before adoption into practice. A review of critical care CPGs showed reasonable uptake of GRADE, with recommendation strength generally aligned with CoE.13 However, strong recommendations are still often made from low or very low-certainty evidence, often related to evidence gaps in RCTs. This highlights the need to integrate high-quality NRSI into CPG development to strengthen evidence synthesis and uptake. Critical care CPG panels should consistently apply GRADE principles, incorporating all high-quality evidence, including NRSI to augment situations where RCT data may be limited or absent. RCTs are resource-intensive and difficult to conduct in critical care.1 To maintain feasibility, investigators often overestimate effect sizes, leading to underpowered studies that may miss true effects.14,15 Reviews of critical care RCTs show that predicted treatment effects are often exaggeratedânearly 10-fold higher than observed, and that few trials sufficiently justify their sample-size targets.16 Similar overestimation has been reported in sepsis, stroke, and trauma trials.17â19 Among high-profile publications, fewer than half of trials had reproducible results.20 Moreover, a meta-epidemiologic review of more than 600 critical care trials found that only 1 in 16 was at low risk of bias, with little improvement over 4 decades.21 These findings suggest that RCTs alone may not provide sufficient high-quality evidence to inform strong guideline recommendations. Critical care populations are highly heterogeneous, encompassing subgroups with different baseline risks and treatment Âresponses. RCTs often target broad syndromes such as sepsis or acute respiratory distress syndrome (ARDS), which likely contributes to many ânegativeâ trials unable to detect differences in outcome.22 Because these studies estimate average treatment effects (ATEs) across diverse patients, potential subgroup benefits can be obscured when other subgroups experience harm.23 This variability, termed heterogeneity of treatment effect (HTE), reflects non-random differences in benefit or harm linked to patient characteristics.24 Understanding HTE (Table 2) is central to precision medicine: treatments that appear neutral on average may conceal offsetting benefit and harm across biologically or contextually distinct subgroups. Explicit exploration of these differences can refine trial design, improve interpretation, and guide targeted recommendations. Methods for assessing heterogeneity of treatment effect. Case example: Heterogeneous treatment effects of therapeutic-dose heparin in patients hospitalized for COVID-19.19 Causal forest and other machine-learning approaches allow for non-linear and interactive modeling of treatment effect heterogeneity but are more susceptible to overfitting and typically require larger sample sizes and external validation. In contrast, regression-based risk modeling approaches are generally more interpretable but may oversimplify interaction effects. RCTs typically assess HTE using pairwise subgroup analyses, but these are often underpowered, rely on arbitrary subgroup thresholds (eg, age <65 vs â„65), and cannot capture complex interactions.25 The American Thoracic Society (ATS) and European Respiratory Society (ERS) guideline on non-invasive ventilation illustrates these limitations: subgroup evidence for conditions such as acute hypoxemic respiratory failure or ARDS came mostly from small or secondary analyses, yielding sparse data and very low certainty.26 These challenges highlight the need for improved data science approaches to identify and characterize HTE: a priority emphasized in the recent ATS research agenda for sepsis and ARDS.27 Data-driven subgroups (subphenotypes) can integrate multiple patient characteristics to assess effect modification and estimate individualized treatment effects.28,29 These models require rigorous derivation and validation to avoid overfitting, yet no consensus framework currently guides their validation or clinical use. Critical care trialists should adopt realistic effect size and recruitment targets and predefine strategies to evaluate clinically relevant HTE. When RCT evidence is insufficient, we propose strategies for CPG panels to integrate NRSI within the GRADE framework to complement RCTs and strengthen recommendations. In accordance with GRADE guidance, if the CoE from RCTs is judged to be high then the role for NRSI is minimal for the specific comparison and outcome of interest.7 However, RCTs often do not report certain patient-important outcomes such as adverse events, quality of life, or longer-term morbidity or mortality. Even if a particular question and outcome of interest have RCT evidence, the estimates of treatment effect are often limited by imprecision due to aforementioned recruitment and sample size concerns. Treatment effects are often assessed in highly selected populations; trial participants typically represent a small fraction of those screened and even meta-analyses may yield low certainty due to imprecision or inconsistency.30,31 In these situations, guideline panels should consider high-quality NRSI, defined by adherence to TARGET (Transparent Reporting of Observational Studies Emulating a Target Trial) standards, acceptable risk of bias, and robust sensitivity analyses, to supplement RCT evidence.7 Target-trial emulation exemplifies this approach: investigators first design a hypothetical randomized trial addressing the question of interest, then emulate it using observational data.8,32 For instance, an international study using the COVID-19 Critical Care Consortium dataset estimated the effect of VV-ECMO versus conventional ventilation in patients with severe COVID-19, providing real-world evidence where an RCT was impractical due to complexity and cost.33 Similar emulations have evaluated intubation,34 ventilation,35 and corticosteroid strategies36 in critical careâdemonstrating how NRSI can inform practice when trials are unfeasible. Consider the example of drotrecogin alfa (activated protein C, rhAPC). Following the PROWESS RCT,37 which demonstrated benefit of rhAPC in patient with septic shock, the large open-label ENHANCE observational study38 reported a similar reduction in mortality with rhAPC but was the first to raise important concerns about serious bleeding, including intracranial hemorrhage. These observational findings influenced early guideline discussions, tempering enthusiasm for the drug, and subsequent RCTs39,40 confirmed this harm and rhAPC was ultimately withdrawn. This highlights that replication across larger datasets remains essential to confirm findings and ensure generalizability beyond selected RCT populations. This sequence illustrates an iterative process: observational signals can generate early warnings or hypotheses that subsequent RCTs confirm or refute. When results diverge, these contrasts can highlight methodological limitations or context-specific factors that warrant further investigation. The TARGET statement outlines 21 reporting items to standardize eligibility, interventions, outcomes, and analyses, improving transparency and reproducibility of emulated trials.41 Adherence to TARGET helps guideline panels assess NRSI rigor and determine when such evidence can complement or upgrade certainty around RCT findings. Similarly, the RCT-DUPLICATE initiative evaluated whether database-derived emulations can reproduce findings from RCTs across 32 cardiovascular studies, including interventions for anticoagulation, antiplatelet therapy, and chronic disease management. The authors found that effect estimates from well-designed emulations closely mirrored their RCT counterparts in both direction and magnitude, demonstrating that real-world data can yield valid causal inference when study design and analytic methods are rigorous.10 Whether successes from other fields will translate to critical care remains uncertain, given its confounding, physiologic complexity, and HTE. A blinded target-trial emulation in this setting reproduced findings of the PreVent RCT examining bag-mask ventilation and hypoxemia,42,43 providing proof-of-principle that short-term physiologic effects can be predicted from observational data, though its value for longer-term or patient-centered outcomes remains untested. Valid causal inference in NRSI requires adherence to key assumptions: exchangeability (no unmeasured confounding), positivity (each patient could receive any treatment), and consistency (observed outcomes reflect potential outcomes under that treatment).8,9 Meeting these assumptions demands careful cohort design, proper time alignment, and analytic techniques that address confounding, such as target-trial emulation, inverse-probability weighting, or doubly robust estimators.44,45 Studies must also handle time-varying confounding and competing risks (eg, death precluding extubation), which can otherwise bias effect estimates.46 To address these concerns, marginal structural models may be used to estimate the causal effect of a time-varying treatment and address the challenge of estimating treatment effects when confounders are influenced by prior treatmentâa situation conventional regression models struggle with. CPG panels should systematically appraise NRSI by verifying TARGET adherence, assessing bias with validated tools such as ROBINS-I, and judging how results affect GRADE domains such as imprecision, inconsistency, and indirectness.41,47 Robust sensitivity analyses, testing alternative models, handling missing data, and probing unmeasured confounding, are essential to confirm result stability and should be clearly reported.48,49 Transparent presentation of assumptions and their plausibility further strengthen credibility. When high-certainty RCT evidence already exists for all relevant target populations, additional NRSI are seldom needed (Figure 1). More often, however, critical care trials involve highly selected populations, making complementary NRSI useful for confirming Âtreatment effects in broader or under-represented groups.50,51 When RCT and NRSI results are consistent, guideline panels may consider upgrading certainty and recommendation strength in line with GRADE guidance.7 GRADE also allows rating up observational evidence when large effects, dose-response relationships, or confounding that would only diminish an observed benefit are present.52 Conversely, inconsistent or methodologically weak NRSI such as those with implausible assumptions, poor reporting, or critical bias, should be excluded, with the rationale documented. Expanding use of target-trial emulation is promising but must be paired with training and standards to prevent low-quality proliferation that could erode confidence in observational evidence.48 Framework for incorporating NRSI into critical care CPGs. CPG panels should incorporate well-conducted NRSI to strengthen CoE and adopt structured workflows: (1) verifying TARGET adherence, (2) considering potential risk of bias, and (3) linking NRSI results to GRADE domains to ensure transparent, reproducible use of observational evidence. Critical care RCTs often study heterogeneous syndromes using strict eligibility criteria that limit generalizability and obscure subgroup effects. A multicenter simulation of 15 landmark trials found that over half of real-world ICU patients would have been ineligible,53 and a review of 75 high-impact trials showed that 60% used at least one poorly justified exclusion such as language barriers or lack of insuranceâfurther restricting applicability.54 Most RCTs originate from high-income countries, leaving major evidence gaps for critically ill patients in the Global South.55 For example, a Zambian sepsis RCT found higher mortality with early fluid resuscitationâcontradicting prior goal-directed therapy trials.56,57 This discordance may be explained by the fact that these trials enrolled predominantly young, malnourished individuals predisposed to pulmonary edema and respiratory failure in a setting with limited ventilatory support. Beyond generating estimates of effectiveness in underrepresented populations, NRSIs also offer a pathway to address structural inequities in evidence generation and utilization. Conducting RCTs in the Global South is often hindered by logistical, regulatory, and infrastructural challengesâincluding limited research infrastructure, ethical oversight, or funding mechanisms, which systematically exclude these populations from RCTs.55 Well-designed NRSI can help bridge such gaps by leveraging local data to assess effectiveness, feasibility, and contextual factors in resource-limited settings. They can also identify structural and contextual modifiers such as malnutrition, health-system capacity, and disease epidemiology; thereby supporting more equitable, context-specific guideline recommendations.58 Embedding such evidence from the Global South not only broadens external validity but also enhances the global relevance of CPGsâthereby promoting more equitable and relevant evidence-based decision-making for clinicians practicing in resource-limited settings. NRSI can also inform feasibility, acceptability, and which are key factors in CPG For instance, the ATS guideline on ARDS a recommendation for VV-ECMO based on NRSI substantial in and across and NRSI can RCT findings to real-world which patients benefit or are based on risk or A key is which to assess how RCT results to external populations and to identify contextual effect improving both evidence relevance and trial Causal inference using real-world data can evaluate HTE across broader populations, including and patients typically underrepresented in a systematic review found major in methodological rigor for HTE analyses, particularly in testing and for confounding, the need for standardized methods and In critical care, HTE from secondary analyses of RCT In the modeling showed that patient characteristics predicted benefit from specific oxygenation targets for patients with and higher for those with The subsequent Care Medicine a recommendation higher oxygenation targets based on very low-certainty an of the trial found that even when are machine-learning models can identify clinically subgroups with benefit or the value of HTE modeling in acute respiratory These secondary analyses are and but should be by observational studies to evaluate HTE beyond RCTs. The ARDS cohort illustrates the value of non-randomized showed that patients had mortality with higher while no benefit in the example of HTE using real-world ICU Beyond also a global of guideline adherence, and ARDS outcomes. not its and rigor how observational studies can yield at a RCTs informing international ARDS When developing panels should consider how best to incorporate NRSI in HTE. this requires systematically HTE analyses, particularly for subgroups in the and assessing how these findings complement subgroup no GRADE yet panels should still evaluate whether HTE evidence recommendations or can guide research for or in RCTs. CPG panels should apply well-conducted causal-inference analyses to confirm the generalizability of RCT findings and identify clinically important HTE. RCTs the standard for and but well-designed NRSI can augment both the certainty and of evidence. Critical care CPG panels should integrate observational evidence when while recognizing methodological standardized (1) TARGET for reporting, (2) validated risk of bias and (3) explicit GRADE will ensure use of NRSI across guideline High-quality NRSI can CoE and generalizability beyond selective RCT populations, providing a to evaluate HTE. incorporating such studies into CPG development may improve both the generalizability and of recommendations. such as the dataset highlight how NRSI can HTE not in trials As analytic methods and target-trial NRSI will an important role in addressing evidence gaps in critical care. will rely on close across and to ensure that NRSI are and with the rigor of randomized authors the the authors to the and of the is at American of and Critical Care Medicine the which have been as tools used in this