Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “probabilistic machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

234 records · Page 13

Digital polycrystalline microstructure generation using diffusion probabilistic models

Accurate micromechanical simulation of polycrystalline materials requires a realistic digital representation of the grain scale microstructure. Here, this work demonstrates the use of a generative diffusion probabilistic model for synthesizing single phase polycrystalline realizations. The model performs well and is capable of producing realistic microstructures consisting of not just simple equiaxed structures but also structures exhibiting more complex spatial arrangements. Masked microstructure generation reveals that the model is context aware of morphological descriptors which may be encoded in the latent space. Training on more diverse data sets, with scaled up architectures, may enable development of future models capable of synthesizing even more complex microstructural features.

36 MATERIALS SCIENCE↗

Understanding the Scalability of Bayesian Network Inference using Clique Tree Growth Curves

Bayesian networks (BNs) are used to represent and efficiently compute with multi-variate probability distributions in a wide range of disciplines. One of the main approaches to perform computation in BNs is clique tree clustering and propagation. In this approach, BN computation consists of propagation in a clique tree compiled from a Bayesian network. There is a lack of understanding of how clique tree computation time, and BN computation time in more general, depends on variations in BN size and structure. On the one hand, complexity results tell us that many interesting BN queries are NP-hard or worse to answer, and it is not hard to find application BNs where the clique tree approach in practice cannot be used. On the other hand, it is well-known that tree-structured BNs can be used to answer probabilistic queries in polynomial time. In this article, we develop an approach to characterizing clique tree growth as a function of parameters that can be computed in polynomial time from BNs, specifically: (i) the ratio of the number of a BN's non-root nodes to the number of root nodes, or (ii) the expected number of moral edges in their moral graphs. Our approach is based on combining analytical and experimental results. Analytically, we partition the set of cliques in a clique tree into different sets, and introduce a growth curve for each set. For the special case of bipartite BNs, we consequently have two growth curves, a mixed clique growth curve and a root clique growth curve. In experiments, we systematically increase the degree of the root nodes in bipartite Bayesian networks, and find that root clique growth is well-approximated by Gompertz growth curves. It is believed that this research improves the understanding of the scaling behavior of clique tree clustering, provides a foundation for benchmarking and developing improved BN inference and machine learning algorithms, and presents an aid for analytical trade-off studies of clique tree clustering using growth curves.

Mengshoel, Ole Jakob↗

NaroNet: Discovery of tumor microenvironment elements from highly multiplexed images

Understanding the spatial interactions between the elements of the tumor microenvironment -i.e. tumor cells. fibroblasts, immune cells- and how these interactions relate to the diagnosis or prognosis of a tumor is one of the goals of computational pathology. We present NaroNet, a deep learning framework that models the multi-scale tumor microenvironment from multiplex-stained cancer tissue images and provides patient-level interpretable predictions using a seamless end-to-end learning pipeline. Trained only with multiplex-stained tissue images and their corresponding patient-level clinical labels, NaroNet unsupervisedly learns which cell phenotypes, cell neighborhoods, and neighborhood interactions have the highest influence to predict the correct label. To this end, NaroNet incorporates several novel and state-of-the-art deep learning techniques, such as patch-level contrastive learning, multi-level graph embeddings, a novel max-sum pooling operation, or a metric that quantifies the relevance that each microenvironment element has in the individual predictions. We validate NaroNet using synthetic data simulating multiplex-immunostained images where a patient label is artificially associated to the -adjustable- probabilistic incidence of different microenvironment elements. We then apply our model to two sets of images of human cancer tissues: 336 seven-color multiplex-immunostained images from 12 high-grade endometrial cancer patients; and 382 35-plex mass cytometry images from 215 breast cancer patients. In both synthetic and real datasets, NaroNet provides outstanding predictions of relevant clinical information while associating those predictions to the presence of specific microenvironment elements.

60 APPLIED LIFE SCIENCES↗

Fuzzy logic, neural networks, and soft computing

The past few years have witnessed a rapid growth of interest in a cluster of modes of modeling and computation which may be described collectively as soft computing. The distinguishing characteristic of soft computing is that its primary aims are to achieve tractability, robustness, low cost, and high MIQ (machine intelligence quotient) through an exploitation of the tolerance for imprecision and uncertainty. Thus, in soft computing what is usually sought is an approximate solution to a precisely formulated problem or, more typically, an approximate solution to an imprecisely formulated problem. A simple case in point is the problem of parking a car. Generally, humans can park a car rather easily because the final position of the car is not specified exactly. If it were specified to within, say, a few millimeters and a fraction of a degree, it would take hours or days of maneuvering and precise measurements of distance and angular position to solve the problem. What this simple example points to is the fact that, in general, high precision carries a high cost. The challenge, then, is to exploit the tolerance for imprecision by devising methods of computation which lead to an acceptable solution at low cost. By its nature, soft computing is much closer to human reasoning than the traditional modes of computation. At this juncture, the major components of soft computing are fuzzy logic (FL), neural network theory (NN), and probabilistic reasoning techniques (PR), including genetic algorithms, chaos theory, and part of learning theory. Increasingly, these techniques are used in combination to achieve significant improvement in performance and adaptability. Among the important application areas for soft computing are control systems, expert systems, data compression techniques, image processing, and decision support systems. It may be argued that it is soft computing, rather than the traditional hard computing, that should be viewed as the foundation for artificial intelligence. In the years ahead, this may well become a widely held position.

Zadeh, Lofti A.↗

Using the optimal combined index weight ratio to improve the probability of anomaly detection in big area additive manufacturing

Big Area Additive Manufacturing (BAAM) of composites requires significant time, energy, and material, so it is critical to reduce production inefficiencies to make functional parts without multiple iterations. Statistical process control coupled with Principal Component Analysis (PCA) is a powerful technique that provides a quick, computationally inexpensive, and intuitive way for operators to detect defects that form in a manufacturing process without massive datasets. Recently, a combined index that is a weighted sum of the Hotelling's T 2 and squared residual error statistics has been proposed that can be monitored in one chart, improving interpretation accuracy and simplicity. However, the literature does not offer a formal method to optimise the weights. Here, we introduce two new approaches to the traditional weight selection approach using simulated and BAAM image data. Approach 1 uses a theoretically motivated optimum inspired by probabilistic principal component analysis. Approach 2 systematically varies the ratio of the weights to find the optimum. We show that approach 1 delivers optimal anomaly detection performance in select cases while approach 2 fares better in practice. Surprisingly, we also show that choosing a more complex PCA model has a minimal negative impact on anomaly detection performance compared to a more simplistic model.

3-dimensional printing↗

Exploring Requirements for Software that Learns: A Research Preview

Context & motivation: The development of software that learns has revolutionized how many systems perform. For the most part, these systems are neither safety- nor mission-critical. However, as technology and aspirations advance, there is an increased desire and need for Machine Learning (ML) software in safety- and mission-critical systems, e.g., driverless cars or autonomous space robotics. Problem: In these domains, reliability is crucial and systems have to undergo much scrutiny in terms of both the developed artefacts and the adopted development process. Central to the development of such systems is the elicitation and definition of software requirements that are used to guide the design and verification process. The addition of software components that learn, and the associated capability for unforeseen behavior, makes defining detailed software requirements especially difficult. Principal ideas/results: In this paper, we identify unique characteristics of software requirements that are specific to ML components. To this end, we collect and examine requirements from both academic and industrial sources. Contribution: To the best of our knowledge, this is the first work that presents real-life, industrial patterns of requirements for ML components. Furthermore, this paper identifies key characteristics and provides a foundation for developing a taxonomy of requirements for software that learns.

Probabilistic requirements↗

BMINN: Learning chemical potentials and parameters from voltage data for multi-phase battery modeling

Free-energy landscapes and chemical potentials govern the dynamics of phase transitions, transport, and stability in functional materials, yet they remain experimentally inaccessible under realistic operating conditions. Here we introduce a Bayesian model-integrated neural network (BMINN) that embeds physics-based formulations of non-autonomous partial differential-algebraic equations into probabilistic learning. This approach reconstructs hidden thermodynamics directly from macroscopic current-voltage data, providing quantitative access to metastable states, staging transitions, and energy barriers without synchrotron probes. Demonstrated on lithium-graphite electrodes, BMINN recovers full Gibbs free-energy landscapes with fidelity validated against operando X-ray diffraction. The framework generalizes across dynamical regimes, enabling accurate voltage prediction, internal state estimation, and inference of governing parameters. Beyond batteries, BMINN exemplifies a broadly applicable strategy for learning missing physics in multiphase, non-equilibrium systems, offering a new pathway to uncover hidden thermodynamic functions across condensed matter and materials physics.

25 ENERGY STORAGE↗

Mixture Model Framework for Traumatic Brain Injury Prognosis Using Heterogeneous Clinical and Outcome Data

Prognoses of Traumatic Brain Injury (TBI) outcomes are neither easily nor accurately determined from clinical indicators. This is due in part to the heterogeneity of damage inflicted to the brain, ultimately resulting in diverse and complex outcomes. Using a data-driven approach on many distinct data elements may be necessary to describe this large set of outcomes and thereby robustly depict the nuanced differences among TBI patients’ recovery. In this work, we develop a method for modeling large heterogeneous data types relevant to TBI. Our approach is geared toward the probabilistic representation of mixed continuous and discrete variables with missing values. The model is trained on a dataset encompassing a variety of data types, including demographics, blood-based biomarkers, and imaging findings. In addition, it includes a set of clinical outcome assessments at 3, 6, and 12 months post-injury. The model is used to stratify patients into distinct groups in an unsupervised learning setting. We use the model to infer outcomes using input data, and show that the collection of input data reduces uncertainty of outcomes over a baseline approach. In addition, we quantify the performance of a likelihood scoring technique that can be used to self-evaluate the extrapolation risk of prognosis on unseen patients.

97 MATHEMATICS AND COMPUTING↗

Status Report on Regulatory Criteria Applicable to the Use of Artificial Intelligence (AI) and Machine Learning (ML)

Although the interest in the use of artificial intelligence (AI) and machine learning (ML) in nuclear energy is increasing rapidly, at present their implementation is limited. This rapid increase in interest is not surprising considering that implementing AI and ML technology would allow for continuous monitoring, facilitate the implementation of predictive maintenance with optimized staffing plans, enable automation and autonomy opportunities that could drastically reduce fixed operation and maintenance costs, and provide training for operations and maintenance. Other industries are using AI for construction, and in the nuclear arena AI could provide great benefit in decommissioning activities. The ability of AI and ML to operate in real time vastly increases their potential impact. Before AI can be used in design, operations, or as a regulatory tool, the specifics on the regulations applicable to the use of AI for nuclear power applications need to be established. The difficulty is that the specific use cases will dictate the applicability of regulations. For example, even within the application domain associated with operations, the regulations might vary if the AI is used to create a virtual reference for plant operations or is used for training, optimization of maintenance intervals, prioritization of maintenance activities, etc. Different still is if the AI is to be used for design or setting technical specifications, which will introduce additional requirements. US Nuclear Regulatory Commission (NRC) licensing reviews are based on an applicant’s design meeting its performance assessment based on (1) safety goals and objectives, (2) deterministic and/or probabilistic analysis of accident scenarios, and (3) quantitative assessment of design alternatives against the safety goals and objectives using accepted engineering tools, methodologies, and performance criteria. The current regulatory framework does not explicitly address AI or autonomous control. However, as implementing AI technology will require the use of a digital platform, it must meet the requirements of an instrumentation and control (I&C) system. The regulatory requirements for AI, which will be incorporated into the I&C system, will be very dependent on how it is used (i.e., its functionality, safety classification, etc.). The licensing process is primarily risk-based with the identification of components and systems as nonsafety, important to safety, or safety related. A risk-informed approach allows further gradation of components and systems based on risk metrics such as core damage frequency or large early release fractions. Thus, the use cases and the risk categorization of impacted systems and components will determine the regulatory requirements. Regardless of how AI is used it presents new opportunities for risk-informing operating, maintenance, and regulatory decisions. Trustworthiness, transparency, and the ability to validate and verify the results will be paramount in showing that the systems and plant still meet their performance requirements. This report describes the results of research to identify regulatory implications of AI technologies and their uses. Specifically, this report reviews current regulatory guidance relevant to the application of AI for design (including design changes or new designs including advanced reactors), construction, operations, training, maintenance, research, testing, and as a regulatory tool. AI can be automated at different levels from purely informative purposes to autonomous controls. The focus of this review included determination of constraints on the application of AI technology, identification of any regulatory gaps or uncertainties, and clarification of anticipated technical basis information likely to be important for regulatory acceptance of these technologies. Currently, any use of AI at nuclear power plants is focused on nonsafety-related applications. The NRC and other regulatory bodies are evaluating providing guidance to address gaps rather than create new regulations to address the use of AI and ML. This approach seems to be the best to encourage AI development without adding regulatory uncertainty.

97 MATHEMATICS AND COMPUTING↗

Fast Characterization of Inducible Regions of Atrial Fibrillation Models With Multi-Fidelity Gaussian Process Classification

Computational models of atrial fibrillation have successfully been used to predict optimal ablation sites. A critical step to assess the effect of an ablation pattern is to pace the model from different, potentially random, locations to determine whether arrhythmias can be induced in the atria. In this work, we propose to use multi-fidelity Gaussian process classification on Riemannian manifolds to efficiently determine the regions in the atria where arrhythmias are inducible. We build a probabilistic classifier that operates directly on the atrial surface. We take advantage of lower resolution models to explore the atrial surface and combine seamlessly with high-resolution models to identify regions of inducibility. We test our methodology in 9 different cases, with different levels of fibrosis and ablation treatments, totalling 1,800 high resolution and 900 low resolution simulations of atrial fibrillation. When trained with 40 samples, our multi-fidelity classifier that combines low and high resolution models, shows a balanced accuracy that is, on average, 5.7% higher than a nearest neighbor classifier. We hope that this new technique will allow faster and more precise clinical applications of computational models for atrial fibrillation. All data and code accompanying this manuscript will be made publicly available at: https://github.com/fsahli/AtrialMFclass.

59 BASIC BIOLOGICAL SCIENCES↗

A Causal Inference Model Based on Random Forests to Identify the Effect of Soil Moisture on Precipitation

Soil moisture influences precipitation mainly through its impact on land–atmosphere interactions. Understanding and correctly modeling soil moisture–precipitation (SM–P) coupling is crucial for improving weather forecasting and subseasonal to seasonal climate predictions, especially when predicting the persistence and magnitude of drought. However, the sign and spatial structure of SM–P feedback are still being debated in the climate research community, mainly due to the difficulty in establishing causal relationships and the high degree of nonlinearity in land–atmosphere processes. To this end, we developed a causal inference model based on the Granger causality analysis and a nonlinear machine learning model. This model includes three steps: nonlinear anomaly decomposition, nonlinear Granger causality analysis, and evaluation of the quality of SM–P feedback, which eliminates the nonlinear response of interannual and seasonal variability and the memory effects of climatic factors and isolates the causal relationship of local SM–P feedback. We applied this model by using National Climate Assessment–Land Data Assimilation System (NCA-LDAS) datasets over the United States. Here, the results highlight the importance of nonlinear atmosphere responses in land–atmosphere interactions. In addition, the strong feedback over the southwestern United States and the Great Plains both highlight the impacts of topographic factors rather than only the sensitivity of evapotranspiration to soil moisture. Furthermore, the SM–P index defined by our framework is used to benchmark Earth system models (ESMs), which provides a new metric for efficiently identifying potential model biases in modeling local land–atmosphere interactions and may help the development of ESMs in improving simulations of water cycle variability and extremes.

54 ENVIRONMENTAL SCIENCES↗

Leveraging Structured Biological Knowledge for Counterfactual Inference: A Case Study of Viral Pathogenesis

Counterfactual inference is a useful tool for comparing outcomes of interventions on complex systems. It requires us to represent the system in form of a structural causal model, complete with a causal diagram, probabilistic assumptions on exogenous variables, and functional assignments. Specifying such models can be extremely difficult in practice. The process requires substantial domain expertise, and does not scale easily to large systems, multiple systems, or novel system modifications. At the same time, many application domains, such as molecular biology, are rich in structured causal knowledge that is qualitative in nature. This manuscript proposes a general approach for querying a causal knowledge graph with a causal question and converting the qualitative result into a quantitative structural causal model that can learn from data to answer the question. Here, we demonstrate the feasibility, accuracy and versatility of this approach using two case studies in systems biology. The first demonstrates the appropriateness of the underlying assumptions and the accuracy of the results. The second demonstrates the versatility of the approach by querying a knowledge base for the molecular determinants of a severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2)-induced cytokine storm and performing counterfactual inference to predict the causal effect of medical countermeasures for severely ill COVID-19 patients.

60 APPLIED LIFE SCIENCES↗

Probabilistic Evaluation of Geoscientific Hypotheses with Geophysical Data: Application to Electrical Resistivity Imaging of a Fractured Bedrock Zone

As climate changes and populations grow, groundwater sustainability is becoming increasingly important. Groundwater models, based on a conceptual understanding of the subsurface structure, are crucial tools for making sustainable management decisions. Conceptual models of the subsurface are based on knowledge of geological processes, and, frequently, observations from geophysical data. A frequent problem in groundwater model development occurs when multiple geological phenomena could explain a single subsurface observation. Uncertainty in geophysical data makes it even more difficult to discern which explanations are consistent with the geophysics. Here, we present a framework for testing geological when a geological feature is observed in geophysical data, but its physical characteristics are uncertain. The framework builds on Popper-Bayes methods developed in previous work, and is applied to study a fractured bedrock zone in a mountainous watershed in southwest Colorado. First, we outline six hypotheses based on the geological history of the watershed. Then, using the proposed Popper-Bayes approach, we demonstrate that three of the six hypotheses are inconsistent with measured electrical resistivity data, even after accounting for uncertainty. Finally, we discuss the importance of the prior model, and how this framework for handling geophysical uncertainty can be applied in other settings.

54 ENVIRONMENTAL SCIENCES↗

Tropical Cyclone Super Resolution using conditional diffusion denoising probabilistic model from mesoscale simulation to LES

Accurate modeling of tropical cyclone wind fields is essential for the design, risk assessment, and operational planning of offshore energy infrastructure. While mesoscale simulations are widely used thanks to their computational efficiency, they lack the necessary resolution to capture key features such as wind shear and veer profiles as well as the distribution turbulent kinetic energy (TKE). High-fidelity large-eddy simulation (LES) models on the other hand, can resolve turbulent structures and provide a more accurate representation of the complex wind field, albeit at a higher computational cost. To address this modeling gap, we introduce a two-part generative framework to enhance the resolution and physics-capturing ability of mesoscale simulations. First, a reduced-order model based on Karhunen–Loève (KL) decomposition is used to extract dominant spatial modes from one-dimensional mean wind profiles. A multilayer perceptron (MLP) is trained to map mesoscale mode weights to their LES counterparts, enabling accurate reconstruction of vertical velocity profiles. Second, a conditional Diffusion Denoising Probabilistic Model (DDPM) is developed to super-resolve coarse and low-fidelity mesoscale velocity fields, recovering fine-scale turbulence structures and stress distributions. The framework is evaluated across different tropical cyclone intensity categories defined by the Saffir–Simpson scale and demonstrates strong performance in both interpolation and extrapolation tasks. The generated fields accurately reproduce spatial coherence, stress distributions, and spectral energy characteristics observed in LES data. By bridging the fidelity gap between mesoscale and LES outputs, this approach offers a scalable, data-driven solution for enhancing the representation of tropical cyclone wind fields, enabling more robust offshore energy infrastructure systems design in tropical-cyclone-prone areas.

17 WIND ENERGY↗

Emulator-Based Bayesian Calibration of the CISNET Colorectal Cancer Models

Purpose To calibrate Cancer Intervention and Surveillance Modeling Network (CISNET)'s SimCRC, MISCAN-Colon, and CRC-SPIN simulation models of the natural history colorectal cancer (CRC) with an emulator-based Bayesian algorithm and internally validate the model-predicted outcomes to calibration targets.Methods We used Latin hypercube sampling to sample up to 50,000 parameter sets for each CISNET-CRC model and generated the corresponding outputs. We trained multilayer perceptron artificial neural networks (ANNs) as emulators using the input and output samples for each CISNET-CRC model. We selected ANN structures with corresponding hyperparameters (i.e., number of hidden layers, nodes, activation functions, epochs, and optimizer) that minimize the predicted mean square error on the validation sample. We implemented the ANN emulators in a probabilistic programming language and calibrated the input parameters with Hamiltonian Monte Carlo-based algorithms to obtain the joint posterior distributions of the CISNET-CRC models' parameters. We internally validated each calibrated emulator by comparing the model-predicted posterior outputs against the calibration targets.Results The optimal ANN for SimCRC had 4 hidden layers and 360 hidden nodes, MISCAN-Colon had 4 hidden layers and 114 hidden nodes, and CRC-SPIN had 1 hidden layer and 140 hidden nodes. The total time for training and calibrating the emulators was 7.3, 4.0, and 0.66 h for SimCRC, MISCAN-Colon, and CRC-SPIN, respectively. The mean of the model-predicted outputs fell within the 95% confidence intervals of the calibration targets in 98 of 110 for SimCRC, 65 of 93 for MISCAN, and 31 of 41 targets for CRC-SPIN.Conclusions Using ANN emulators is a practical solution to reduce the computational burden and complexity for Bayesian calibration of individual-level simulation models used for policy analysis, such as the CISNET CRC models. In this work, we present a step-by-step guide to constructing emulators for calibrating 3 realistic CRC individual-level models using a Bayesian approach.

artificial neural networks↗

A data-driven method for modelling dissipation rates in stratified turbulence

We present a deep probabilistic convolutional neural network (PCNN) model for predicting local values of small-scale mixing properties in stratified turbulent flows, namely the dissipation rates of turbulent kinetic energy and density variance, $\varepsilon$ and $\chi$ . Inputs to the PCNN are vertical columns of velocity and density gradients, motivated by data typically available from microstructure profilers in the ocean. The architecture is designed to enable the model to capture several characteristic features of stratified turbulence, in particular the dependence of small-scale isotropy on the buoyancy Reynolds number $Re_b:=\varepsilon /(\nu N^2)$ , where $\nu$ is the kinematic viscosity and $N$ is the background buoyancy frequency, the correlation between suitably locally averaged density gradients and turbulence intensity and the importance of capturing the tails of the probability distribution functions of values of dissipation. Empirically modified versions of commonly used isotropic models for $\varepsilon$ and $\chi$ that depend only on vertical derivatives of density and velocity are proposed based on the asymptotic regimes $Re_b\ll 1$ and $Re_b\gg 1$ , and serve as an instructive benchmark for comparison with the data-driven approach. When trained and tested on a simulation of stratified decaying turbulence which accesses a range of turbulent regimes (associated with differing values of $Re_b$ ), the PCNN outperforms assumptions of isotropy significantly as $Re_b$ decreases, and additionally demonstrates improvements over the fitted empirical models. A differential sensitivity analysis of the PCNN facilitates a comparison with the theoretical models and provides a physical interpretation of the features enabling it to make improved predictions.

42 ENGINEERING↗

Harnessing Artificial Intelligence for Medical Diagnosis and Treatment During Space Exploration Missions

From May 8th to June 9th, 2023, I had the opportunity to participate in an experiential learning experience at Johnson Space Center in Houston, TX with Exploration Medical Capability (ExMC), an element of the NASA Human Research Program. During this research experience, I was not only able to work on the above titled research project, but also gain an immense exposure to the field of aerospace medicine, make numerous connections within the field, tour NASA facilities, as well as travel to the Aerospace Medical Association Annual Conference (AsMA) in New Orleans. To briefly introduce my project, it is well understood that the medical capabilities available to crew medical officers (CMOs) on the International Space Station will be different than the capabilities available and needed during deep space exploration missions to the Moon, Mars, and beyond. Ground support is particularly limited due to distance, communication delays (or lack of communication), and lack of resupply. Therefore, to support medical care by CMOs on these missions, robust clinical decision support systems (CDSSs) must be designed. The recent publication and public launch of generative artificial intelligence (AI) tools based upon large language models (LLM) such as ChatGPT provides the opportunity to create a smart assistant for onboard triage, diagnosis, and treatment of medical conditions. Ultimately, the overall purpose of the project was to research what AI tools currently exist or are in development, and to see how they might be implemented onboard during exploration class spaceflights of the future. The ExMC element is actively developing several tools to be used in preparation for and during deep space exploration missions. One of those tools, known as IMPACT, is a probabilistic risk assessment model which can be used to propose a desired medical system (based on mass and volume) and suggest the clinical outcomes likely to occur for a design reference mission (DRM). The group recently presented the IMPACT model and a DRM of interest titled “Modified Long Duration Lunar Orbital and Lunar Surface” (mLDLOLS) at the recent AsMA conference. The mLDLOLS mock mission is a 9 month and 6-day deep space exploration mission consisting of time in Moon’s orbit (3 months on the Gateway space station), on the lunar surface (3 months within habitat), and another 3 months on Gateway before return to Earth. For this DRM, IMPACT ultimately outlined a preferred medical system that was then associated with medical conditions considered to be most likely based on frequency, most likely to cause astronaut task time loss (TTL), most likely to cause return to definitive care (RTDC), and most likely cause loss of crew life (LOCL). IMPACT also highlighted the medical capabilities/skills that would be required to care for those medical conditions, such as performing a history of present illness or musculoskeletal exam with ultrasound. The primary objective of the project was to perform a survey of the AI tools and systems applicable to the conditions outlined for the proposed mLDLOLS mission. Using PubMed (including most relevant MeSH terms) and Google Scholar, we then created a robust annotated bibliography organized by condition. The 56-page and over 500 reference annotated bibliography was subsequently used to create a review outline that would become the basis for drafting of a future publication. For the review outline, we took those medical conditions researched within the annotated bibliography (condition-based approach) and deployed a systems-based approach, combining those medical conditions and related tools into ten categories. These categories included general/all-purpose CDSSs, tools to diagnose or manage respiratory, dermatologic, neurologic, auditory and vestibular, ophthalmic, musculoskeletal, infection-associated, and gynecologic conditions, as well as tools that could be deployed in the setting of trauma/emergency. With the completion of the 30-page outline, we then began drafting the review paper. To conclude the research experience, I presented the findings from our survey to the ExMC Clinical and Science team. With these objectives, I ultimately learned about the number of AI tools that exist today to assist medical professionals with the triage, diagnosis, and management of several medical conditions. These tools can span from chatbot assistants to help triage knee pain to vision transformer models that can identify ophthalmic conditions based on ocular surface images captured with a cell phone. We also highlighted the current gaps that exist in the literature alongside the advancements that are needed to make the desired CDSS for deep space exploration missions. With this experience, I certainly confirmed an existing career goal and identified several additional skills needed to become an aerospace medical doctor including knowledge of critical care in an extreme medicine setting, aerospace engineering and human integration systems, artificial intelligence, machine learning, and risk models. I also identified numerous transferable skills for this career goal including the basic knowledge of medicine (MD), deployment of the scientific method for critical thought about new scientific questions (PhD), review of published literature, including creating an annotated bibliography (PhD), as well as detailed scientific writing (PhD). The results of my research will likely guide the design of an all-encompassing onboard medical assistant for use during deep space exploration missions of the future. I plan on sharing the outcomes from this experience with my peers at a student seminar in the Fall semester on August 30th. During the seminar, I will detail the project, my experience at NASA and AsMA, as well as offer best practice guidelines for students entertaining similar experiences or careers. In conclusion, I would like to thank the WVU School of Medicine, Research and Graduate Education office, as well as NASA ExMC for the unwavering support of this life-changing experience.

Ryan A. Lacinski↗