Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Predictive Data analytics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

An AI-Enabled MODEX Framework for Improving Predictability of Subsurface Water Storage across Local and Continental Scales

Focal Area: (2) Predictive modeling through the use of AI techniques and AI-derived model components. (3) Insight gleaned from complex data using AI, big data analytics, and other advanced methods. We propose an AI-enabled model-experiment (MODEX) framework to improve the predictability of subsurface water storage (SWS) from local to conus scales in a changing environment by taking advantage of DOE’s observation and simulation capabilities, as well as to inform the model and the observation development.

54 ENVIRONMENTAL SCIENCES↗

Method for simultaneous characterization and expansion of reference libraries for small molecule identification

A variational autoencoder (VAE) has been developed to learn a continuous numerical, or latent, representation of molecular structure to expand reference libraries for small molecule identification. The VAE has been extended to include a chemical property decoder, trained as a multitask network, to shape the latent representation such that it assembles according to desired chemical properties. The approach is unique in its application to metabolomics and small molecule identification, focused on properties that are obtained from experimental measurements (m/z, CCS) paired with its training paradigm, which involves a cascade of transfer learning iterations. First, molecular representation is learned from a large dataset of structures with m/z labels. Next, in silico property values are used to continue training. Finally, the network is further refined by being trained with the experimental data. The trained network is used to predict chemical properties directly from structure and generate candidate structures with desired chemical properties. The network is extensible to other training data and molecular representations, and for use with other analytical platforms, for both chemical property and feature prediction as well as molecular structure generation.

Colby, Sean M.↗

Data for A Hybrid Biophysical-Machine Learning Framework for Diurnal Surface Energy Flux Estimation Using Proximal Sensing

Thermal infrared-based remote sensing of surface energy fluxes has traditionally relied on high spatial resolution satellite data with revisit frequencies on the order of weeks. In this study, we evaluate a biophysics-based analytical surface energy balance model for predicting latent energy (LE) and sensible heat (H) fluxes using proximal sensing observations. The Surface Temperature Initiated Closure (STIC1.2) model has been extensively validated across a wide range of spatial and temporal scales using various satellite-derived thermal infrared data sets. Here we extend this validation by applying STIC at sub-hourly temporal resolution over multiple growing seasons for four distinct agricultural systems. We further develop and evaluate novel STIC variants that incorporate machine learning (ML) techniques to eliminate the need for surface energy balance observations, specifically net radiation and soil heat flux, thereby enhancing model applicability in data-sparse settings. The integration of a ML component to estimate surface available energy is shown to have strong predictive performance for both LE (R2 = 0.81–0.94) and H (R2 = 0.46–0.72) across all agricultural systems examined here, demonstrating the potential of hybrid biophysical-machine learning approaches for surface energy balance modeling with minimal data requirements. This study concludes with a novel application of explainable machine learning (exML) to diagnose sources of model error. This exML framework attributes residual prediction errors to both model input variables and environmental drivers not explicitly included in the simulation experiments. This approach provides a new pathway for improving model design and integrating previously overlooked yet influential variables into future model iterations.

AI/ML↗

Learning to simulate high energy particle collisions from unlabeled data

In many scientific fields which rely on statistical inference, simulations are often used to map from theoretical models to experimental data, allowing scientists to test model predictions against experimental results. Experimental data is often reconstructed from indirect measurements causing the aggregate transformation from theoretical models to experimental data to be poorly-described analytically. Instead, numerical simulations are used at great computational cost. We introduce Optimal-Transport-based Unfolding and Simulation (OTUS), a fast simulator based on unsupervised machine-learning that is capable of predicting experimental data from theoretical models. Without the aid of current simulation information, OTUS trains a probabilistic autoencoder to transform directly between theoretical models and experimental data. Identifying the probabilistic autoencoder’s latent space with the space of theoretical models causes the decoder network to become a fast, predictive simulator with the potential to replace current, computationally-costly simulators. Here, we provide proof-of-principle results on two particle physics examples, Z-boson and top-quark decays, but stress that OTUS can be widely applied to other fields.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Temperature Dependence of Band Gap Renormalization in High-T Sensor Materials via First-Principles and Experimental Corroboration

Understanding the temperature dependence of functional properties of high-T gas sensing materials is vital for their applications in combustion environments. The electron-phonon coupling that derives the electronic structure change with temperatures is a key property of interest as it affects other sensing responses. Herein, we assess the temperature dependence of band gap renormalization in metal oxides and perovskites by employing Allen-Heine-Cardona theory with first-principles simulations and corroborate with experimental observation. The calculated temperature-dependent band gap changes of these materials studied are in good agreement with in-house experimental data, proving that the theory can adequately predict renormalization on the band gap in the system of interest. The predicted and measured band gap variations are characterized using an analytical model, which can provide useful insights on the simulated zero-temperature band gaps. Based on the available data, a set of 53 metal oxides and perovskites were identified as potential high-T gas sensors. A machine learning model has been developed to predict the band-gap change by capturing the overall trend of the empirical parameters with respect to a reduced feature obtained by transforming the set of available physical features.

Park, Jongwoo↗

Facilitating better and faster simulations of aerosol-cloud interactions in Earth system models

Focal Area(s): 1. Predictive modeling through the use of AI techniques and AI-derived model components; the use of AI and other tools to design a prediction system comprising a hierarchy of models. 2. Insight gleaned from complex data (both observed and simulated) using AI, big data analytics, and other advanced methods, including explainable AI and physics- or knowledge-guided AI. Science Challenge: One major challenge that Earth system models (ESMs) face in providing credible prediction of the Earth system and its water cycle characteristics (e.g., mean state, variability, and extreme events) is to accurately simulate aerosol-cloud interactions (ACI). The physical, chemical, and dynamical processes affecting ACI are extremely complex and they range from nanoscale to planetary scale. In each model development cycle, scientists spend significant efforts investigating model deficiencies and uncertainties associated with aerosols (e.g., emissions, chemical processes, aerosol microphysics, and transport) and clouds (e.g., macrophysics, microphysics, turbulence, and large-scale circulation) in order to develop improved treatments. However, despite decades of active research, ACI is still a major source of uncertainty in climate projections, even though great progress has been made. Specific scientific challenges include: (i) Parameterizations are developed based on limited data; (ii) The complexity of a parameterization required for accurate predictions is not understood; (iii) Incomplete and unknown physics leads to errors in the fully coupled Earth system; and (iv) Complex physics is computationally too expensive to employ in ESMs.

54 ENVIRONMENTAL SCIENCES↗

Data-driven studies of magnetic two-dimensional materials

We use a data-driven approach to study the magnetic and thermodynamic properties of van der Waals (vdW) layered materials. We investigate monolayers of the form A 2 B 2 X 6 , based on the known material Cr 2 Ge 2 Te 6 , using density functional theory (DFT) calculations and machine learning methods to determine their magnetic properties, such as magnetic order and magnetic moment. We also examine formation energies and use them as a proxy for chemical stability. We show that machine learning tools, combined with DFT calculations, can provide a computationally efficient means to predict properties of such two-dimensional (2D) magnetic materials. Our data analytics approach provides insights into the microscopic origins of magnetic ordering in these systems. For instance, we find that the X site strongly affects the magnetic coupling between neighboring A sites, which drives the magnetic ordering. Our approach opens new ways for rapid discovery of chemically stable vdW materials that exhibit magnetic behavior.

42 ENGINEERING↗

Can machine learning accelerate process understanding and decision‐relevant predictions of river water quality?

Abstract The global decline of water quality in rivers and streams has resulted in a pressing need to design new watershed management strategies. Water quality can be affected by multiple stressors including population growth, land use change, global warming, and extreme events, with repercussions on human and ecosystem health. A scientific understanding of factors affecting riverine water quality and predictions at local to regional scales, and at sub‐daily to decadal timescales are needed for optimal management of watersheds and river basins. Here, we discuss how machine learning (ML) can enable development of more accurate, computationally tractable, and scalable models for analysis and predictions of river water quality. We review relevant state‐of‐the art applications of ML for water quality models and discuss opportunities to improve the use of ML with emerging computational and mathematical methods for model selection, hyperparameter optimization, incorporating process knowledge into ML models, improving explainablity, uncertainty quantification, and model‐data integration. We then present considerations for using ML to address water quality problems given their scale and complexity, available data and computational resources, and stakeholder needs. When combined with decades of process understanding, interdisciplinary advances in knowledge‐guided ML, information theory, data integration, and analytics can help address fundamental science questions and enable decision‐relevant predictions of riverine water quality.

54 ENVIRONMENTAL SCIENCES↗

Leveraging Hydropower Multi-Sensor Data for Inference and Age-Informed Modeling

Increased demand of operational flexibility such as faster ramp up/down in generation, and more frequent start/stops are putting hydropower plants and their associated components in unprecedented stress. Consequently, these plants are at the high risk of extended and more frequent outage to accommodate unscheduled, and unexpected maintenance. Therefore, hydropower plants are in critical need of data driven and age-informed analysis for their regular and unscheduled operation. Yet not all hydropower plants are exhaustively equipped with sensors and/or measurement streams for their respective components – demanding solutions on how to detect, identify, and locate the cause of any event from the unobservable. Idaho National Laboratory (INL) analyzed the anonymized measurements and event records from the Hydropower Research Institute (HRI) to address this issue, as part of the Water Power Technologies Office (WPTO) funded one year multi-lab project. First, we investigated how time series of multiple sensor measurements can be leveraged to identify an event “root cause” as well as to develop an inference (i.e., estimate the unobservable) problem. INL also investigated how individual hydropower components’ reaction or response times vary across the pre-event, during event, and post-event conditions – enabling the hydropower dynamic models to be age-informed. Finally, the impact of clustering multi-sensor time series on short-term vibration prediction is analyzed. INL will present key findings from these analyses and recommend next steps for stakeholder adoption.

13 HYDRO ENERGY↗

Perspectives on Machine Learning-assisted Plasma Medicine: Towards Automated Plasma Treatment

Cold atmospheric plasmas (CAPs) have shown great promise for medical applications through their synergistic chemical, electrical, and thermal effects, which can induce therapeutic outcomes. However, safe and reproducible plasma treatment of complex biological surfaces poses a major hurdle to the widespread adoption of CAPs for medical applications. Predictive modeling of the mutual interactions between the plasma and biological surfaces and, thus, systematic approaches to quantify and predict plasma treatment outcomes remain largely elusive due to the lack of mechanistic understanding of plasma-surface interactions that can span across vastly different length-and time-scales. In addition, real-time sensing capabilities in biomedical CAP devices are often limited, which can be detrimental to plasma treatment due to the intrinsic plasma and surface variability during the treatment, as well as sensitivity to external perturbations. All of these challenges can make reproducible and effective plasma treatment of biological surfaces difficult to realize, which is further compounded by errors due to human operation of hand-held CAP devices. Machine learning and data-driven approaches can be particularly useful in addressing these challenges in three major ways: (i) data-driven modeling of hard-to-model plasma-surface interactions and plasma treatment outcomes; (ii) learning data analytics for plasma and surface diagnostics in real-time; and (iii) developing predictive controllers that enable reliable and effective CAP treatments. Furthermore, this paper discusses the promise of machine learning to accelerate plasma medicine research in these areas, toward machine learning-assisted and automated CAP treatment of complex biological surfaces.

60 APPLIED LIFE SCIENCES↗

In Situ Inference for Earth System Predictability

Focal Area: Focal Area 3: Insight gleaned from complex simulated data using AI, big data analytics, and other advanced methods, including explainable AI and physics- or knowledge-guided AI. Science Challenge: An understanding of future evolution in precipitation extremes is critical to numerous DOE mission questions. Extreme events are by nature short time-scale events that are difficult to diagnose in available model data. Accurate modeling of extreme events necessarily requires high spatial resolution at the storm scale locally. However, the environment in which storms grow is dependent on global, remote, processes. These complex spatiotemporal relationships are impossible to diagnose at resolutions required to accurately model storms responsible for extreme precipitation. At exascale, climate simulations will produce results at fine enough resolution to investigate these relationships. However, the resulting data from these simulations will be far too large to save for post-simulation analysis. We advocate for fitting statistical models inside the simulations as they run, a context known as in situ, which will facilitate scientific investigations using the full fine-scale data stream. Figure 1 shows an example of the type of model we could consider, a Bayesian hierarchical spatial regression model. Precipitation extremes at each grid cell are modeled using extreme value distributions. Since extremes are rare, fitting models to individual grid cells can result in high variance and poor estimates. Instead, the model can be made more robust by smoothing the parameters of the extreme value model across space. Additionally, the parameters themselves can be functionally linked to other variables elsewhere in the simulation. Thus, we can use the fine-scale data to build more robust models for extremes that link extreme behavior to other climate patterns.

54 ENVIRONMENTAL SCIENCES↗

Structural stability of thin overhanging walls during material extrusion additive manufacturing of thermoset-based ink

Recent developments have enabled material extrusion additive manufacturing of thermoset-based composite inks on the large scale. In addition, printing out-of-plane components is of broad interest to the polymer material extrusion community. Here, we address some of the challenges associated with both large-scale and out-of-plane thermoset material extrusion additive manufacturing by studying the height at which thin overhanging walls collapse. Walls at a range of overhang angles were printed until they collapsed. An optical camera captured the profile of each wall throughout the print, allowing the collapse height to be identified and the geometric fidelity to the programmed angle to be evaluated. Using previously measured rheological properties, predictive models were generated to approximate the collapse height and profile of the deflected walls. First, an analytical model was created to predict the height at which the walls would yield. The analytical model assumes the walls exhibit a perfectly linear profile; however, experiments proved this assumption to be false. Therefore, a finite element simulation was developed to account for the elastic deflection that occurs during printing. The finite element simulation predicts both the yield height and the deflected profile after the deposition of each layer. For the properties of the thermoset ink used here, the yield height predicted by the analytical model and finite element simulation are virtually identical. These predictions match experimental data reasonably well, but minor errors are observed. Accounting for the fully plastic moment appears to explain the small mismatch between experimental data and predictions. Additionally, the finite element simulation provides an excellent prediction of the deflected profile before the wall begins to collapse. Finally, by demonstrating that the collapse height and deflected profile of thin overhanging walls can be predicted, this work illustrates how the soft viscoelastic properties of thermoset-based composite inks limit the scale of a key feature required to print some nonplanar components. It also provides a basis to tailor in-process curing systems to suppress deflection and collapse of thin overhanging walls.

36 MATERIALS SCIENCE↗

Envisioning U.S. Climate Predictions and Projections to Meet New Challenges

In the face of a changing climate, the understanding, predictions, and projections of natural and human systems are increasingly crucial to prepare and cope with extremes and cascading hazards, determine unexpected feedbacks and potential tipping points, inform long-term adaptation strategies, and guide mitigation approaches. Increasingly complex socio-economic systems require enhanced predictive information to support advanced practices. Such new predictive challenges drive the need to fully capitalize on ambitious scientific and technological opportunities. These include the unrealized potential for very high-resolution modeling of global-to-local Earth system processes across timescales, reduction of model biases, enhanced integration of human systems and the Earth Systems, better quantification of predictability and uncertainties; expedited science-to-service pathways, and co-production of actionable information with stakeholders. Enabling technological opportunities include exascale computing, advanced data storage, novel observations and powerful data analytics, including artificial intelligence and machine learning. Looking to generate community discussions on how to accelerate progress on U.S. climate predictions and projections, representatives of Federally-funded U.S. modeling groups outline here perspectives on a six-pillar national approach grounded in climate science that builds on the strengths of the U.S. modeling community and agency goals. This calls for an unprecedented level of coordination to capitalize on transformative opportunities, augmenting and complementing current modeling center capabilities and plans to support agency missions. Tangible outcomes include projections with horizontal spatial resolutions finer than 10 km, representing extremes and associated risks in greater detail, reduced model errors, better predictability estimates, and more customized projections to support next generation climate services.

54 ENVIRONMENTAL SCIENCES↗

The L-CAPE Project at FNAL

The controls system at FNAL records data asynchronously from several thousand Linac devices at their respective cadences, ranging from 15Hz down to once per minute. In case of downtimes, current operations are mostly reactive, investigating the cause of an outage and labeling it after the fact. However, as one of the most upstream systems at the FNAL accelerator complex, the Linac’s foreknowledge of an impending downtime as well as its duration could prompt downstream systems to go into standby, potentially leading to energy savings. The goals of the Linac Condition Anomaly Prediction of Emergence (L-CAPE) project that started in late 2020 are (1) to apply data-analytic methods to improve the information that is available to operators in the control room, and (2) to use machine learning to automate the labeling of outage types as they occur and discover patterns in the data that could lead to the prediction of outages. We present an overview of the challenges in dealing with time-series data from 2000+ devices, our approach to developing an ML-based automated outage labeling system, and the status of augmenting operations by identifying the most likely devices predicting an outage.

43 PARTICLE ACCELERATORS↗

Prediction of hemiwicking dynamics in micropillar arrays

Dynamic hemiwicking behavior is observable in both nature and a wide range of industrial applications ranging from biomedical devices to thermal management. We present a semi-analytical modeling framework (without empirical fitting coefficients) to predict transient capillary-driven hemiwicking behavior of a liquid through a nano/microstructured surface, specifically a micropillar array. In our model framework, the liquid domain is discretized into micropillar unit cells to enable the time marching of the hemiwicking front. A simplified linear pressure drop is assumed along the hemiwicking length such that the local meniscus curvature, contact angle, and effective liquid height are determined at each time step in our transient model. This semi-analytical model is validated with experimental data from our own experiments and from published literature for different fluids. Our model predicts hemiwicking dynamics with <20% error over a broad range of micropillar geometries with height-to-pitch ratio ranging between ≈0.34 and 6.7 and diameter-to-pitch ratio in the range of ≈0.25–0.7 and without any fitting parameters. For lower diameter-to-pitch ratio data points related to sparse micropillar array arrangements, we suggest modifications to the semi-analytical model. This work sheds light on complex and dynamic solid–liquid–vapor interfacial interactions which could serve as a guide for the design of textured surfaces for wicking enhancement in multi-phase thermal and mass transport technologies and applications.

Mechanics↗

Machine Learning in the Context of Laser-Induced Breakdown Spectroscopy

The integration of machine learning (ML) with Laser-Induced Breakdown Spectroscopy (LIBS) has revolutionized the analytical capabilities of LIBS. The combi-nation of both methods enables more accurate and efficient data analysis. While LIBS itself is a powerful technique for elemental analysis, the vast amount of spectral data it generates can be hard to interpret. Machine learning addresses these challenges by leveraging algorithms that can learn from data, identify patterns, and make predictions without explicit programming for the interpretation of each specific task. In LIBS application, ML techniques are used to enhance various analytical processes. For example, ML algorithms can classify materials based on their spectral fingerprints, predict the concentration of elements in a sample, and identify underlying patterns within complex datasets. Here, this application improves the precision of LIBS analyses while significantly reducing the time required for data processing and interpretation. In this chapter, the fundamental concepts of ML will be discussed first. Following this, the process of data splitting and the importance of feature selection will be examined. Several machine learning methods will then be closely examined, exploring how each can benefit LIBS analysis and highlighting their respective advantages and shortcomings. This structured approach will provide a comprehensive understanding of the integration of ML in the context of LIBS analysis.

47 OTHER INSTRUMENTATION↗

A machine learning approach for determining temperature-dependent bandgap of metal oxides utilizing Allen–Heine–Cardona theory and O’Donnell model parameterization

To evaluate the high temperature sensing properties of metal oxide and perovskite materials suitable for use in combustion environments, it is necessary to understand the temperature dependence of their bandgaps. Although such temperature-driven changes can be calculated via the Allen–Heine–Cardona (AHC) theory, which assesses electron–phonon coupling for the bandgap correction at given temperatures, this approach is computationally demanding. Another approach to predict bandgap temperature-dependence is the O’Donnell model, which uses analytical expressions with multiple fitting parameters that require bandgap information at 0 K. This work employs data-driven Gaussian process regression (GPR) to predict the parameters employed in the O’Donnell model from a set of physical features. We use a sample of 54 metal oxides for which density functional theory has been performed to calculate the bandgap at 0 K, and the AHC calculations have been carried out to determine the shift in the bandgap at non-zero temperatures. As the AHC calculations are impractical for high-throughput screening of materials, the developed GPR model attempts to alleviate this issue by predicting the O'Donnell parameters purely from physical features. To mitigate the reliability issues arising from the very small size of the dataset, we apply a Bayesian technique to improve the generalizability of the data-driven models as well as quantify the uncertainty associated with the predictions. The method captures well the overall trend of the O’Donnell parameters with respect to a reduced feature set obtained by transforming the available physical features. Quantifying the associated uncertainty helps us understand the reliability of the predictions of the O’Donnell parameters and, therefore, the bandgap as a function of temperature for any novel material.

36 MATERIALS SCIENCE↗

The Lorentzian inversion formula and the spectrum of the 3d O(2) CFT

We study the spectrum and OPE coefficients of the three-dimensional critical O(2) model, using four-point functions of the leading scalars with charges 0, 1, and 2 ( s , $\phi$, and t ). We obtain numerical predictions for low-twist OPE data in several charge sectors using the extremal functional method. We compare the results to analytical estimates using the Lorentzian inversion formula and a small amount of numerical input. We find agreement between the analytic and numerical predictions. We also give evidence that certain scalar operators lie on double-twist Regge trajectories and obtain estimates for the leading Regge intercepts of the O(2) model.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗