Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Traditional Machine Learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

When Machine Learning Meets 2D Materials: A Review

The availability of an ever-expanding portfolio of 2D materials with rich internal degrees of freedom (spin, excitonic, valley, sublattice, and layer pseudospin) together with the unique ability to tailor heterostructures made layer by layer in a precisely chosen stacking sequence and relative crystallographic alignments, offers an unprecedented platform for realizing materials by design. However, the breadth of multi-dimensional parameter space and massive data sets involved is emblematic of complex, resource-intensive experimentation, which not only challenges the current state of the art but also renders exhaustive sampling untenable. To this end, machine learning, a very powerful data-driven approach and subset of artificial intelligence, is a potential game-changer, enabling a cheaper – yet more efficient – alternative to traditional computational strategies. It is also a new paradigm for autonomous experimentation for accelerated discovery and machine-assisted design of functional 2D materials and heterostructures. Here, the study reviews the recent progress and challenges of such endeavors, and highlight various emerging opportunities in this frontier research area.

2D materials↗

Assessment of Quantum ML Applicability for Climate Actions: Comparison of the Variational Quantum Classifier and the Quantum Support Vector Classifier with Classical ML Models

Climate change refers to significant and long-term alterations in the Earth’s climate patterns, typically resulting from human activities that increase greenhouse gas emissions. Addressing climate change is not merely an option but a necessity, demanding creative solutions and efforts from individuals, researchers, communities, and governments. Despite the capabilities of machine learning (ML) with data-driven solutions promising to combat climate change-related problems, they face challenges stemming from traditional computational methods and prolonged training times, impeding their practical utility. Recent strides in quantum computing have permeated diverse domains, spanning from manufacturing engineering and pharmaceutical discovery to the latest frontier of detecting climate anomalies. With the potential to substantially reduce time and computational complexity, quantum computing shows promise in addressing climate change impacts. Its distinctive features will enable the concurrent exploration of expansive solution spaces, making it well-suited for analyzing extensive climate datasets, simulating intricate climate models, optimizing resource allocation, and discerning patterns in climate data for mitigation and adaptation endeavors. This study explores the potential of using Quantum machine learning (QML) techniques on climate and weather data obtained from NASA Giovannis. We used two QML algorithms, the Quantum Support Vector Classifier (QSVC) and the Variational Quantum Classifier (VQC) models, using the IBM Qiskit ML 0.7.2 ecosystem. We used an actual 127-Qubit IBM Quantum Computer (IBM 127-qubit Eagle) in this study. The methodology and results sections describe the experiences gained from applying and evaluating quantum ML results on climate and weather data obtained from NASA satellites as a novel practical application of quantum computing.

Earth Observational Data↗

Detecting Masquerade Attacks in Controller Area Networks Using Graph Machine Learning

Modern vehicles rely on a myriad of electronic control units (ECUs) interconnected via controller area networks (CANs) for critical operations. Despite their ubiquitous use and reliability, CANs are susceptible to sophisticated cyberattacks, particularly masquerade attacks, which inject false data that mimic legitimate messages at the expected frequency. These attacks pose severe risks such as unintended acceleration, brake deactivation, and rogue steering. Traditional intrusion detection systems (IDS) often struggle to detect these subtle intrusions due to their seamless integration into normal traffic. This paper introduces a novel framework for detecting masquerade attacks in the CAN bus using graph machine learning (ML). We hypothesize that the integration of shallow graph embeddings with time series features derived from CAN frames enhances the detection of masquerade attacks. We show that by representing CAN bus frames as message sequence graphs (MSGs) and enriching each node with contextual statistical attributes from time series, we can enhance detection capabilities across various attack patterns compared to using graph-based features only. Our method ensures a comprehensive and dynamic analysis of CAN frame interactions, improving robustness and efficiency. Extensive experiments on the ROAD dataset validate the effectiveness of our approach, demonstrating statistically significant improvements in the detection rates of masquerade attacks compared to a baseline that uses graph-based features only as confirmed by Mann-Whitney U and Kolmogorov-Smirnov tests (p < 0.05) .

Marfo, William [Univ. of Texas, El Paso, TX (Unite↗

A Comprehensive Machine Learning Study to Classify Precipitation Type over Land from Global Precipitation Measurement Microwave Imager (GPM-GMI) Measurements

Precipitation type is a key parameter used for better retrieval of precipitation characteristics as well as to understand the cloud–convection–precipitation coupling processes. Ice crystals and water droplets inherently exhibit different characteristics in different precipitation regimes (e.g., convection, stratiform), which reflect on satellite remote sensing measurements that help us distinguish them. The Global Precipitation Measurement (GPM) Core Observatory’s microwave imager (GMI) and dual-frequency precipitation radar (DPR) together provide ample information on global precipitation characteristics. As an active sensor, the DPR provides an accurate precipitation type assignment, while passive sensors such as the GMI are traditionally only used for empirical understanding of precipitation regimes. Using collocated precipitation type flags from the DPR as the “truth”, this paper employs machine learning (ML) models to train and test the predictability and accuracy of using passive GMI-only observations together with ancillary information from a reanalysis and GMI surface emissivity retrieval products. Out of six ML models, four simple ones (support vector machine, neural network, random forest, and gradient boosting) and the 1-D convolutional neural network (CNN) model are identified to produce 90–94% prediction accuracy globally for five types of precipitation (convective, stratiform, mixture, no precipitation, and other precipitation), which is much more robust than previous similar effort. One novelty of this work is to introduce data augmentation (subsampling and bootstrapping) to handle extremely unbalanced samples in each category. A careful evaluation of the impact matrices demonstrates that the polarization difference (PD), brightness temperature (Tc) and surface emissivity at high-frequency channels dominate the decision process, which is consistent with the physical understanding of polarized microwave radiative transfer over different surface types, as well as in snow and liquid clouds with different microphysical properties. Furthermore, the view-angle dependency artifact that the DPR’s precipitation flag bears with does not propagate into the conical-viewing GMI retrievals. This work provides a new and promising way for future physics-based ML retrieval algorithm development.

machine learning/artificial intelligence↗

Applied Machine-Learning Models to Identify Spectral Sub-Types of M Dwarfs from Photometric Surveys

M dwarfs are the most abundant stars in the Solar Neighborhood and they are prime targets for searching for rocky planets in habitable zones. Consequently, a detailed characterization of these stars is in demand. The spectral sub-type is one of the parameters that is used for the characterization and it is traditionally derived from the observed spectra. However, obtaining the spectra of M dwarfs is expensive in terms of observation time and resources due to their intrinsic faintness. We study the performance of four machine-learning (ML) models—K-Nearest Neighbor (KNN), Random Forest (RF), Probabilistic Random Forest (PRF), and Multilayer Perceptron (MLP)—in identifying the spectral sub-types of M dwarfs at a grand scale by deploying broadband photometry in the optical and near-infrared. We trained the ML models by using the spectroscopically identified M dwarfs from the Sloan Digital Sky Survey (SDSS) Data Release (DR) 7, together with their photometric colors that were derived from the SDSS, Two-Micron All-Sky Survey, and Wide-field Infrared Survey Explorer. We found that the RF, PRF, and MLP give a comparable prediction accuracy, 74%, while the KNN provides slightly lower accuracy, 71%. We also found that these models can predict the spectral sub-type of M dwarfs with ~99% accuracy within ±1 sub-type. The five most useful features for the prediction are r - z, r - i, r - J, r - H , and g - z, and hence lacking data in all SDSS bands substantially reduces the prediction accuracy. However, we can achieve an accuracy of over 70% when the r and i magnitudes are available. Since the stars in this study are nearby (d ≲ 1300 pc for 95% of the stars), the dust extinction can reduce the prediction accuracy by only 3%. Finally, we used our optimized RF models to predict the spectral sub-types of M dwarfs from the Catalog of Cool Dwarf Targets for the Transiting Exoplanet Survey Satellite, and we provide the optimized RF models for public use.

79 ASTRONOMY AND ASTROPHYSICS↗

The confluence of machine learning and multiscale simulations

Multiscale modeling has a long history of use in structural biology, as computational biologists strive to overcome the time- and length-scale limits of atomistic molecular dynamics. Contemporary machine learning techniques, such as deep learning, have promoted advances in virtually every field of science and engineering and are revitalizing the traditional notions of multiscale modeling. Deep learning has found success in various approaches for distilling information from fine-scale models, such as building surrogate models and guiding the development of coarse-grained potentials. However, perhaps its most powerful use in multiscale modeling is in defining latent spaces that enable efficient exploration of conformational space. In conclusion, this confluence of machine learning and multiscale simulation with modern high-performance computing promises a new era of discovery and innovation in structural biology.

59 BASIC BIOLOGICAL SCIENCES↗

A Deep Potential model for liquid–vapor equilibrium and cavitation rates of water

Computational studies of liquid water and its phase transition into vapor have traditionally been performed using classical water models. Here, we utilize the Deep Potential methodology—a machine learning approach—to study this ubiquitous phase transition, starting from the phase diagram in the liquid–vapor coexistence regime. The machine learning model is trained on ab initio energies and forces based on the SCAN density functional, which has been previously shown to reproduce solid phases and other properties of water. Here, we compute the surface tension, saturation pressure, and enthalpy of vaporization for a range of temperatures spanning from 300 to 600 K and evaluate the Deep Potential model performance against experimental results and the semiempirical TIP4P/2005 classical model. Moreover, by employing the seeding technique, we evaluate the free energy barrier and nucleation rate at negative pressures for the isotherm of 296.4 K. Further, we find that the nucleation rates obtained from the Deep Potential model deviate from those computed for the TIP4P/2005 water model due to an underestimation in the surface tension from the Deep Potential model. From analysis of the seeding simulations, we also evaluate the Tolman length for the Deep Potential water model, which is (0.091 ± 0.008) nm at 296.4 K. Finally, we identify that water molecules display a preferential orientation in the liquid–vapor interface, in which H atoms tend to point toward the vapor phase to maximize the enthalpic gain of interfacial molecules. We find that this behavior is more pronounced for planar interfaces than for the curved interfaces in bubbles. This work represents the first application of Deep Potential models to the study of liquid–vapor coexistence and water cavitation.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Computational toolkit for predicting thickness of 2D materials using machine learning and autogenerated dataset by large language model

The thickness of 2D materials not only plays a crucial role in determining the performance of nanoelectronic and optoelectronic devices but also introduces complexities in predicting volume-dependent properties, such as energy storage capacity, due to the intrinsic vacuum within these materials. Although a plethora of experimental techniques, including but not limited to optical contrast, Raman spectroscopy, nonlinear optical spectroscopy, near-field optical imaging, and hyperspectral imaging, facilitate the measurement of 2D material thickness, comprehensive data for many materials remain elusive. Over the past decade, the exponential proliferation of 2D materials and their heterostructures has outstripped the capabilities of conventional experimental and computational approaches. In this evolving landscape, machine learning (ML) has emerged as an indispensable tool, offering a scalable approach to augment these traditional methodologies. Addressing the critical gap, we introduce THICK2D—Thickness Hierarchy Inference and Calculation Kit for 2D Materials. This Python-based computational framework harnesses an autogenerated thickness database, developed using large language models, and advanced ML algorithms to facilitate the rapid and scalable estimation of material thickness, relying solely on crystallographic data. To demonstrate the utility and robustness of THICK2D, we successfully used the toolkit to predict the thickness of more than 8000 2D-based materials, sourced from two extensive 2D materials databases. THICK2D is disseminated as an open-source utility, accessible on GitHub at https://github.com/gmp007/THICK2D, and archived on Zenodo at https://10.5281/zenodo.11216648.

Ekuma, Chinedu E. (ORCID:0000000258527556)↗

Deconvoluting experimental decay energy spectra: The O 26 case

In nuclear reaction experiments, the measured decay energy spectra can give insights into the shell structure of decaying systems. However, extracting the underlying physics from the measurements is challenging due to detector resolution and acceptance effects. The Richardson-Lucy (RL) algorithm, a deblurring method that is commonly used in optics and has proven to be a successful technique for restoring images, was applied to our experimental nuclear physics data. The only inputs to the method are the observed energy spectrum and the detector's response matrix also known as the transfer matrix. We demonstrate that the technique can help access information about the shell structure of particle-unbound systems from the measured decay energy spectrum that is not immediately accessible via traditional approaches such as χ-square fitting. For a similar purpose, we developed a machine learning model that uses a deep neural network (DNN) classifier to identify resonance states from the measured decay energy spectrum. We tested the performance of both methods on simulated data and experimental measurements. Then, we applied both algorithms to the decay energy spectrum of 26 O → 24 O + n + n measured via invariant mass spectroscopy. Here, the resonance states restored using the RL algorithm to deblur the measured decay energy spectrum agree with those found by the DNN classifier. Both deblurring and DNN approaches suggest that the raw decay energy spectrum of 26 O exhibits three peaks at approximately 0.15 MeV, 1.50 MeV, and 5.00 MeV, with half-widths of 0.29 MeV, 0.80 MeV, and 1.85 MeV, respectively.

Spectrometers & spectroscopic techniques↗

Supporting Risk-informed Decision-making During Reactor Accidents

Uncertainty in severe accident evolution and outcome is driven by event bifurcations that represent distinctive challenges to defensive layers and tend to promote the emergence of discrete classes of core damage and accident risk. This discrete set of "attractor" states arise from the complex networks of competing physical phenomena and conditional event cascades occurring as the overall system degrades – a process that yields increasing degrees of freedom and accident progression pathways. Characterization of these event spaces has proven elusive to more traditional data interrogation methods, but proves tractable by application of more advanced data collection and machine learning approaches. Through application of these approaches we demonstrate a conceptual framework that enables real-time/robust, risk-informed decision-making support to improve accident mitigation and encourage “graceful exits” during low probability, extreme events limiting accident consequences. In this analysis, we simulated over 8,000 short-term station blackout (STSBO) accidents with the state-of-the-art integral severe accident code, MELCOR, and demonstrate the potential for ML approaches to predict simulation outcomes. We chose to pair ML tools with interpretable and mechanistic event trees for the considered STSBO accident space to predict the likelihood of future event paths along the tree. In addition to the current state of the system, we use information from recent trajectories of temperature, pressure, and other physical features, combining both the current state and past trajectories to forecast future event paths. Finally, we simulate the random injection of variable amounts of water to quantify the efficacy of available actions at reducing risks along the many branches in the event tree. We identify scenarios and windows of opportunity to mitigate risk as well as scenarios in which such actions are unlikely to alter the accident end-state.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Rapid Bayesian High Entropy Alloy Designs Fabricated via Wire Arc Additive Manufacturing

Purpose: This project seeks to demonstrate a new high-throughput (rapid) alloy design technique applied to creating new high entropy alloys (HEAs) for extreme environments. High entropy alloys shift the design paradigm from being focused on a single principal element (e.g. nickel-based alloys) to target alloys that include high atomic fractions (X >10%) of multiple elements. These HEA materials can exhibit sluggish diffusion and enhanced corrosion resistance, ideal for potential applications in advanced ultra supercritical (A-USC) steam cycles for power generation. Scope: The addition of multiple elements in high atomic fractions creates an enormous design space that cannot easily be investigated by traditional material design strategies such as designed of experiments (DOE). This project utilizes a Bayesian machine learning algorithm that has been modified to work with calculation of phase diagrams (CALPHAD) software. This Bayesian algorithm reduces manual inputs and increase the likelihood of achieving an optimal solution. Compositional inputs to this algorithm will be assessed using existing material property models for high temperature strength and corrosion resistance. The target for alloy performance will be a 15% (~100 ⁰C) increase in allowable service temperature beyond heat-resistant stainless steels while maintaining or improving alloy cost and corrosion resistance. Haynes 230 was selected as a baseline, which is 57 wt% Ni with 22 wt% Cr 14 wt% W, and 2 wt% Mo as solid solution strengtheners. In addition to rapid design via Bayesian machine learning, the alloys were rapidly fabricated using a multi-wire arc additive manufacturing (mWAAM) technique which allows for precise control of alloy composition and assessing of alloy design “windows” to study composition effects. Build speeds for wire-arc additive processes are among the highest for additive technologies enabling rapid and reliable sample fabrication when compared to conventional methods such as arc button melting. The mWAAM samples will be rapidly characterized via instrumented indentation for room temperature modulus and strength and for elevated temperature strength via hot hardness tests. After being screened with hardness testing, potential alloys will be further evaluated with conventional microscopy techniques including scanning electron microscopy (SEM) and transmission electron microscopy (TEM) to assess agreement with modeling results. The most promising compositions will also be evaluated by printing full sized tensile specimens for mechanical behavior tests at elevated temperatures. Results: Bayesian machine learning of a single performance function was initially used to optimize five performance metrics: 1) single phase stability, 2) yield strength, 3) creep resistance (low diffusion coefficient), 4) freezing range (weldability), and 5) material cost. The single performance function was suboptimal as assumptions had to be made about the results while formulating the optimization. A goal-oriented Bayesian optimization strategy (Hanaoka, 2021) was implemented with CALPHAD for use with the five metrics above. This multi-objective Bayesian optimization (MOBO) enabled the design of NiCrCoFe alloys with V and W additions. A base composition of NiCoCr was selected as Ni provides a stable FCC matrix, Cr aids corrosion/oxidation resistance, and Co is a solid-solutions strengthener that also improves creep by increasing the activation energy. Fe helps reduce diffusion coefficients and cost. Finally, V and W were selected for their reasonable solubility and high atomic misfit to aid in solid solution strengthening. Cracking of the mWAAM specimens was an early issue, and the Easton solidification cracking model (Easton et al., 2014a) was selected for addition to the MOBO function. High performing alloys fabricated by mWAAM included Ni 28 Cr 25 Co 26 Fe 15 V 8 and Ni 62 Cr 18 Co 1 Fe 3 W 15 . It was observed that even after adapting the mWAAM process for W, the W did not fully dissolve. To fully evaluate the Ni 62 Cr 18 Co 1 Fe 3 W 15 composition, a cored wire (80-20 NiCr sheath/powder core) was manufactured and printed via WAAM, and HIP’ing was utilized to homogenize and densify the printed alloy. The V and W alloys produced met metrics 1 (solid solution), 4 (solidification cracking), and 5 (cost). However, an unmodeled mechanism of thermal stress cracking was identified in the WAAM produced materials, perhaps exacerbated by the lack of grain boundary strengthening elements (B, C). Conclusions & Recommendations: A high-throughput (rapid) alloy design technique was applied to designing and manufacturing new high entropy alloys (HEAs) for extreme environments utilizing MOBO and mWAAM. The developed process was rapid and effective in addressing the mechanisms included in the model. The lack of grain boundary strengthening element additions (e.g., B, C) was a simplification that likely produced thermal stress cracking that turned into a large part of the investigation. Additions on the order of 0.005 wt% B and 0.05 wt% C likely would have minimized thermal stress grain boundary cracking. Overall, the high throughput design strategy is promising for rapid design of metrics-driven alloys for advanced ultra supercritical (A-USC) steam cycles for power generation. The MOBO and mWAAM process could be commercialized to accelerate metrics-driven alloy design. In addition, the cored-wire process utilized for scale-up is a promising high-volume process for WAAM alloy development and scale-up.

36 MATERIALS SCIENCE↗

Machine learning overcomes human bias in the discovery of self-assembling peptides

Peptide materials have a wide array of functions, from tissue engineering and surface coatings to catalysis and sensing. Tuning the sequence of amino acids that comprise the peptide modulates peptide functionality, but a small increase in sequence length leads to a dramatic increase in the number of peptide candidates. Traditionally, peptide design is guided by human expertise and intuition and typically yields fewer than ten peptides per study, but these approaches are not easily scalable and are susceptible to human bias. Here, in this work, we introduce a machine learning workflow—AI-expert—that combines Monte Carlo tree search and random forest with molecular dynamics simulations to develop a fully autonomous computational search engine to discover peptide sequences with high potential for self-assembly. We demonstrate the efficacy of the AI-expert to efficiently search large spaces of tripeptides and pentapeptides. The predictability of AI-expert performs on par or better than our human experts and suggests several non-intuitive sequences with high self-assembly propensity, outlining its potential to overcome human bias and accelerate peptide discovery.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Secondary porosity prediction in complex carbonate reefs using 3D CT scan image analysis and machine learning

Development and preservation of carbonate porosity and permeability are critical to characterizing reservoirs. However, secondary porosity, such as vugs and fractures, are difficult to identify and require collection of expensive wireline tools or core sampling. Wireline logs and cores have traditionally been used to identify the presence of secondary porosity but fail to quantify the contribution to total reservoir porosity. Additionally, advanced wireline logs and core are not readily available for most wells. Dual energy CT scans were collected on whole core from the A-1 Carbonate and Brown Niagaran formations drawn from six wells in northern Michigan. 3D analysis techniques were applied to identify and isolate secondary porosity features. A series of machine learning and data analytics techniques were applied to the dataset to predict secondary porosity features on basic wireline logs. The predictive model successfully predicted secondary porosity with high confidence.

02 PETROLEUM↗

Characterizing vertical upper ocean temperature structures in the European Arctic through unsupervised machine learning

In-situ observations of subsurface ocean temperatures are, in many regions, inconsistently distributed in time and space. These spatio-temporal inconsistencies in the observational network lead to difficulties in utilizing those observations effectively for ocean model evaluation or understanding larger-scale ocean characteristics. Model accuracy of subsurface ocean characteristics is especially important within regions that contain complex ocean structures. One such region is the European Arctic which not only contains several types of water masses with unique characteristics, but also wintertime sea ice coverage and complex bathymetry. This study presents an unsupervised neural networking technique that can be used in combination with traditional ocean model evaluation techniques to provide additional information on the accuracy of modeled vertical ocean temperature profiles. Self-organizing maps is an unsupervised machine learning technique that we apply to approximately twenty thousand Argo and CTD temperature profiles from 2012 to 2020 in the European Arctic to categorize the observed vertical ocean temperature structures in the top 150 m. The observed ocean profile categories, or neurons, defined by the self-organizing map show strong spatial and temporal dependencies. We then use the neuron weights, or the learned temperature profile structure of each neuron, to validate the spatial and temporal variability of modeled vertical temperature structures. This analysis gives us new insights about the model’s capabilities to reproduce specific vertical structures of the top-most ocean layer within different regions and seasons. Mapping modeled ocean temperature profiles onto the neuron-space of the observationally-defined self organized map highlights the potential of this method to advance our understanding of model deficiencies in that region.

54 ENVIRONMENTAL SCIENCES↗

Benchmarking the performance of uncertainty quantification methods for neural network-based interatomic potentials

Machine-learned interatomic potentials (ML-IAPs) continue to gain popularity as accurate, computationally efficient replacements for traditional, physics-based interatomic potentials and expensive ab initio methods. Uncertainty quantification (UQ) of ML-IAPs is a growing area of research as UQ is critical in many applications of IAPs, such as developing curated datasets, active learning-based data augmentation, self-improving models, and estimating the uncertainty of molecular dynamics simulations. In this paper, we construct and benchmark a series of different neural network potentials (NNPs) with varying network architectures to determine the performance of these models with respect to both the mean and uncertainty calibration error. Each NNP method is specifically designed to predict either epistemic or aleatoric uncertainty with particular focus on the differences in behavior between the epistemic and aleatoric uncertainty estimates. We benchmark these methods using multiple datasets common in the ML-IAP literature. The results show that the aleatoric uncertainty from single-shot model architectures is a competitive alternative to ensemble-based epistemic uncertainty predictions in regions of sufficient data-density. However, in regions where the representative data is sparse, aleatoric uncertainty models tend to overpredict and epistemic methods tend to underpredict the actual model error. We conclude that the type of UQ is crucial when discussing performance of probabilistic model results as different methods have different performance characteristics depending on the regime in which they are evaluated. Therefore, the type of UQ method should be carefully evaluated against both the data characteristics and requirements for the intended application.

97 MATHEMATICS AND COMPUTING↗

Assessing Several Non-Traditional Data Sources for Value in Aviation Safety

The NASA System-Wide Safety (SWS) project and its predecessor projects have been developing Machine Learning (ML) algorithms for commercial aviation safety for many years. These algorithms have been applied to Flight Operations Quality Assurance (FOQA); radar track data (e.g., Threaded Track); and safety reports, including Aviation Safety Reporting System (ASRS) and Aviation Safety Action Plan (ASAP). SWS is working with partners to get access to other data that air carriers provide, such as maintenance data, and has been assisting carriers in working with other data, such as Line Operations Safety Audit (LOSA) data, using manual methods. However, the project has discussed whether there are other data that are not traditionally used in aviation safety analysis that may be useful. This paper discusses four sets of data and models that are not traditionally used in aviation safety but that have shown promise for such use. In the future, we plan to incorporate such data into ML algorithms to use with data that we have used before and determine the additional benefit that is actually achieved under different contexts from the inclusion of these non-traditional data sources.

Nikunj C. Oza↗

NASA Earth Systems Digital Twins (ESDT)

"Similarly to artificial intelligence, which is now revolutionizing many aspects of our daily lives, Earth system digital twin technologies have the potential to revolutionize the way Earth Science research will be conducted in the future, and how results and knowledge from this research will provide information to support decision making and yield impactful societal benefits. An Earth System Digital Twin or ESDT is a dynamic and interactive information system that first provides a digital replica of the past and current states of the Earth or Earth system as accurately and timely as possible; second, allows for computing forecasts of future states under nominal assumptions and based on the current replica; and third, offers the capability to investigate many hypothetical scenarios under varying impact assumptions. In other words, an ESDT provides the integrated What-Now, What-Next, and What-If pictures of the Earth or Earth system, by continuously ingesting newly observed data and by leveraging multiple interconnected models, machine learning as well advanced computing and visualization capabilities. Digital twins have been developed in engineering since 2002, but the interest in digital twins for the Earth domain is more recent and stems from the convergence of several developments: - The huge amount of diverse data that has now been collected continuously for more than 50 years, and that is becoming more and more difficult to access, understand, and utilize. - At the same time, because of climate change and its impacts the information produced by all of this data is becoming of interest to many new non-traditional users for analyzing and predicting various phenomena. - Because of advances in computational and visualization capabilities and the parallel unprecedented development of machine learning (ML), extracting relevant information from these large amounts of data and running complex models faster has become possible. As a result, it is becoming necessary and possible to build intuitive and interactive frameworks that will enable users with various skill levels and/or organizational hierarchy levels to easily access large amounts of targeted information along with the relevant tools and models (Earth system and human activity models), to support them in analyzing and visualizing this information, to help them understand interactions among models, to visualize the potential outcomes of various impacts, and to support decision or policy making. The full power of digital twins is that, through an integrated representation and standardized tools and software technologies, the same digital replica can address the needs of multiple users at various resolutions (spatial and temporal) and for various applications (science, economic, policy, etc.) – “from farmer to scientist”. With all these interests at stake, the challenges of building optimal digital twins are many and complex. The first challenge is to determine if a Digital Twin should be global or local, and multi-domain or thematic. For example, some domains such as Climate or Weather will require a global Digital Twin or Digital Twin capabilities while science areas such as Biodiversity might be more local. We can also envision that multiple thematic ESDTs, e.g., Air Quality, Wildfires, Hydrology could be federated or provide input to other ESDTs, either on a regional level or to a more global ESDT. Overall, we can imagine a future “web” of Digital Twins co-existing in a hierarchy or in a network, and capable of being connected or federated depending on the needs. This last point brings up the very important challenge of interoperability, including standards and protocols that will need to be built into these systems from the beginning. Each individual digital twin would have full flexibility in internal construction but would need standards-based interfaces (input and output) or hooks to make it compatible with others. Another challenge when building digital twins will be to decide how to organize each digital replica. Based on the applications targeted by the DT under implementation, various amounts and types of raw data, Analysis Ready Data (ARD) and information will need to be incorporated. Depending on the required latencies and needs of the users, various solutions can be considered, including Data Cubes, Data Lakes, pointers, or computing information on demand. We envision that each ESDT will choose a solution adapted to its specific objectives. Another important challenge is the type(s) of visualization that will be used, as well as the level of interactivity and refresh rate that will be required. Again, this will depend on the objectives of the ESDT, but also on the various users’ needs. In most cases, several types of visualizations and human interfaces will need to be offered depending on the projected users of that system. In parallel to the challenges highlighted above, there are also many tools and technologies that will need to be developed or improved for all types of digital twins. Among those are improved machine learning technologies, for example providing explainability, but also ML techniques for causality and providing a better integration of physics models. Additionally, reliable uncertainty quantification methods will be needed for all ESDT components, from validating data fusion and assimilation to assessing the accuracy of ML models and weighing the values of decisions supported by those systems. This presentation introduces the ESDT concept, presents several ESDT use cases, and a proposed ESDT architecture framework, as well as various technologies being developed by the Advanced Information Systems Technology (AIST) Program."

Earth Science Remote Sensing; Information Systems↗