Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Predictive Data analytics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Ensemble methods for quantification of potassium oxide in ChemCam Mars and laboratory spectra

In this paper we test new approaches for predicting the amount of element oxides in rock samples from the ChemCam instrument suite onboard the NASA Curiosity rover by focusing on K 2 O. Using the expanded dataset compiled by Gasda et al. (2021) with and without the Earth to Mars (E2M and NoE2M) transformation discussed in Clegg et al. (2017) we trained blended submodels using the “double blending” technique and compared these to ensemble methods (Random Forest, ExtraTrees, and Gradient Boosting Regression). We found that ensemble methods performed similar to blended submodels when looking at RMSE-P on the laboratory spectra and provided significant advantages when looking at spectra coming from Mars. For the full model, blended submodels achieved an RMSE-P of 0.62 and 0.60 (E2M and NoE2M respectively) while Gradient Boosting Regression resulted in a slightly improved RMSE-P of 0.59 and 0.60. More importantly, by employing a local RMSE-P estimation technique where model performance is evaluated based on nearby test samples we found that using ensemble methods can lower the quantification limit for K 2 O from the current value of ≈0.6 wt% to ≈0.08 wt% using Extra Trees and Random Forest. This would allow for a much larger range of K 2 O values to be quantified on Mars with greater certainty given that most targets seen on Mars tend to have <1 wt% K2O. Finally, we used both Mean Decrease in Impurity (MDI) and permutation importance techniques to investigate the wavelengths used by the ensemble methods and found that they correspond to known potassium emission lines. This suggests that ensemble methods can provide an easier to train and improved alternative to blended submodels for predicting potassium compositions from Laser Induced Breakdown Spectroscopy (LIBS) data.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Exploring Structure-Sensitive Relations for Small Species Adsorption Using Machine Learning

Accurate prediction of adsorption energies on heterogeneous catalyst surfaces is crucial to predicting reactivity and screening materials. Adsorption linear scaling relations have been developed extensively but often lack accuracy and apply to one adsorbate and a single binding site type at a time. These facts undermine their ability to predict structure sensitivity and optimal catalyst structure. Using machine learning on nearly 300 density functional theory calculations, we demonstrate that generalized coordination number scaling relations hold well for oxygen- and high-valency carbon-binding species but fail for others. Here we reveal that the valency and the electronic coupling of a species with the surface, along with the site type and its coordination environment, are critical for small species adsorption. The model simultaneously predicts the adsorption energy and preferred site and significantly outperforms linear scalings in accuracy. It can expose the structure sensitivity of chemical reactions and enable enhanced catalyst activity via engineering particle shape and facet defects. The generality of our methodology is validated by training the model with transition metal data and transferring it to predict adsorption energies on single-atom alloys.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A thermoelectric materials database auto-generated from the scientific literature using ChemDataExtractor

An auto-generated thermoelectric-materials database is presented, containing 22,805 data records, automatically generated from the scientific literature, spanning 10,641 unique extracted chemical names. Each record contains a chemical entity and one of the seminal thermoelectric properties: thermoelectric figure of merit, ZT; thermal conductivity, κ; Seebeck coefficient, S; electrical conductivity, σ; power factor, PF; each linked to their corresponding recorded temperature, T. The database was auto-generated using the automatic sentence-parsing capabilities of the chemistry-aware, natural language processing toolkit, ChemDataExtractor 2.0, adapted for application in the thermoelectric-materials domain, following a rule-based sentence-simplification step. Data were mined from the text of 60,843 scientific papers that were sourced from three scientific publishers: Elsevier, the Royal Society of Chemistry, and Springer. To the best of our knowledge, this is the first automatically-generated database of thermoelectric materials and their properties from existing literature. The database was evaluated to have a precision of 82.25% and has been made publicly available to facilitate the application of data science in the thermoelectric-materials domain, for analysis, design, and prediction.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

The importance of cycle-by-cycle data in performing rapid battery technology development and validation

Lithium-ion battery (LiB) technology is playing a crucial role in transforming the predominantly fossil fuel-based transportation and stationary storage sectors to achieve a low-carbon economy. Rapid innovation in the LiB materials to electrode to cell design is happening to satisfy the performance, life, and safety metrics required by those myriads of applications. Lately, advanced analytics, such as machine-learning or artificial intelligence (ML/AI) techniques, are being used more frequently to aid in expedited LiB technology development, performance validation, and life prediction. The success of these techniques often relies on a large volume of well-defined and high-quality battery test data. On the other hand, most battery developers and research and development (R&D) communities are still following a classical approach to develop batteries, which is running calendar- and/or cycle-aging tests, performing reference performance tests (RPTs), and conducting post-mortem analyses periodically without paying attention to the wealth of data often not collected during the calendar or cycle life aging tests. This sparse data collection approach is time- and resource-intensive, requiring data capture and evaluation of months to years of RPT data to diagnose accurate battery state of performance, health, and safety. Even so, the underlying aging modes and mechanisms can be missed. If collected properly, battery test data during cycling or calendaring can be efficiently combined with ML/AI techniques to create powerful tools in the rapid diagnosis of battery state of performance, health, and safety along with insights into underlying aging modes and mechanisms. In this report, we discuss the importance of effective cycle-by-cycle (CBC) data collection with example case studies. Within a reasonable timeframe, RPT data are often inadequate in capturing many of the crucial battery aging dynamics, which often predominantly show up in CBC test data. Finally, we also show examples of ML/AI techniques that use CBC data in rapid diagnosis and projection of LiB state of health (SOH) to motivate the scientific community in collecting and using CBC data to facilitate expeditious technology development and validation.

25 ENERGY STORAGE↗

Integration of Condition-Based, Diagnostic, Prognostic, And Anomaly Detection Data into Reliability Models to Support a Predictive Maintenance Context

Reliability data employed in plant reliability models are an approximated integral representation of the past industrywide operational experience, and they neglect the present asset health status (available, for example, from online monitoring data and diagnostic assessments) and forecasted health projection (when available from prognostic models). Ideally, in a predictive maintenance context, system reliability models should support decision making by propagating actual health information from the asset to the system level in order to provide a quantitative snapshot of system health and identify the most critical assets. Asset health should be informed solely by that specific asset’s current and historical performance data and should not be an approximated integral representation of the past industrywide operational experience (as currently performed by system reliability models through Bayesian updating processes). This paper proposes a reliability modeling approach that relies on asset diagnostic and prognostic assessments, along with monitoring data to measure asset health. We show how state-of-the art condition-based, diagnostic, prognostic, and anomaly detection models can be linked to system reliability models not in probability terms, but in terms of margin where margin is defined as the “distance” between the present status and an undesired event (e.g., failure or unacceptable performance). Then, we show how the propagation of margin data from the asset to the system level is performed through classical reliability models such as fault trees or reliability block diagrams. The described method is in fact able to propagate heterogenous health data from the asset to the system level in order to analytically assess system health.

97 MATHEMATICS AND COMPUTING↗

Evaluating the Impact of Proprietary Oil & Gas Data on Machine Learning Model Performance Using a Quasi-Experimental Analytical Approach

This study implements a data-intensive supervised ML approach through a quasi-experimental framework with the objective of quantifying the impact of oil and gas operator-specific proprietary data on ML-based predictive model performance relative to using oil and gas datasets that may be more commonly publicly available. The models are designed to jointly predict daily oil, gas, and water production for horizontal wells as a function of bottom-hole pressure drawdown, spatial placement across the study domain, and well completion attributes. Model performance is quantified on holdout test data to evaluate how each dataset affects resulting model variant performance.

02 PETROLEUM↗

Review—Effects of Solution and Alloy Composition on Critical Crevice Temperature

Critical temperature for localized corrosion can be a good design parameter because localized corrosion is not likely to occur below that temperature. The critical temperature depends on alloy composition, microstructure, and environment chemistry (including its redox potential). This paper reviews the literature on critical temperature for localized corrosion, expressed either as Critical Pitting Temperature (CPT) or Critical Crevice Temperature (CCT). A history of various testing methods is presented. Different approaches for modeling the temperature of transition to active pit growth are reviewed, including probabilistic aspects of critical temperature. A semi-empirical, electrolyte-based, model is described that can be useful in predicting CCT in service environments that differ from standard laboratory test environments. The model predictions are compared to experimental data for various alloys. The effect of solvent on CCT/CPT is described briefly and future avenues of research are recommended.

02 PETROLEUM↗

A machine learning approach for clinker quality prediction and nonlinear model predictive control design for a rotary cement kiln

Abstract Cement manufacturing is energy‐intensive (5Gj/t) and comprises a significant portion of the energy footprint of concrete systems. Incorporating modern monitoring, simulation and control systems will allow lower energy use, lower environmental impact, and lower costs of this widely used construction material. One of the goals of the CESMII roadmap project on the Smart Manufacturing of Cement included developing an analytical process model for clinker quality that includes the chemistry of the kiln feed and accounts for critical process variables. This predictive model will be used in nonlinear model predictive control system designed to significantly reduce process energy use while maintaining or improving product quality. In the cement manufacturing plant used in this study, the kiln feed (meal) is tested every 12 h and used to estimate the mineral composition of the cement kiln output (clinker) using the stoichiometry‐based Bogue's model and the expertise of the plant operators. During kiln operation, kiln output (clinker) is sampled and tested every 2 h to measure its chemical and mineral composition. The predicted and measured values of the clinker composition are used by the plant operators to adjust the kiln input stream and the production process characteristics to maintain stable operation and uniform product quality. However, the time delay between prediction and testing, along with inaccuracies inherent in the Bogue's model have made any process changes designed to minimize energy use problematic, especially in‐light of potential clinker quality issues that process changes often pose. A new analytical model that integrates quality information and process operation information has been developed from data collected from 2 years of production from an operating cement facility. To make the model fuel‐type‐independent, consumed heat energy was computed in the model instead of fuel type and amount. A Feedforward Network was trained and tailored from collected data. Many data‐based simulations were conducted to quantitatively evaluate the proposed model and the 5‐fold cross‐validation procedure was used to test the models. The resulting predictive model was shown to have a low root mean square error (MSE) with respect to the estimated clinker mineral composition compared to that using the industry standard “Bogue’ model”. The end goal of this work was to develop a single machine learning tool that allows the use of quality control data and process control variables to improve energy efficiency of the process in a continuous fashion. The proposed nonlinear model predictive control system (NMPC) can generate predicted kiln production characteristics based on manipulated variables in manner that accurately follows the target product quality values. Simulation results also show that the proposed model produced accurate predictions of kiln outputs that fell within the required constraints, while manipulating control variables within typical operational ranges.

Ali, Asem M.↗

Correcting for filter-based aerosol light absorption biases at the Atmospheric Radiation Measurement program's Southern Great Plains site using photoacoustic measurements and machine learning

Abstract. Measurement of light absorption of solar radiation by aerosols is vital for assessing direct aerosol radiative forcing, which affects local and global climate. Low-cost and easy-to-operate filter-based instruments, such as the Particle Soot Absorption Photometer (PSAP), that collect aerosols on a filter and measure light attenuation through the filter are widely used to infer aerosol light absorption. However, filter-based absorption measurements are subject to artifacts that are difficult to quantify. These artifacts are associated with the presence of the filter medium and the complex interactions between the filter fibers and accumulated aerosols. Various correction algorithms have been introduced to correct for the filter-based absorption coefficient measurements toward predicting the particle-phase absorption coefficient (Babs). However, the inability of these algorithms to incorporate into their formulations the complex matrix of influencing parameters such as particle asymmetry parameter, particle size, and particle penetration depth results in prediction of particle-phase absorption coefficients with relatively low accuracy. The analytical forms of corrections also suffer from a lack of universal applicability: different corrections are required for rural and urban sites across the world. In this study, we analyzed and compared 3 months of high-time-resolution ambient aerosol absorption data collected synchronously using a three-wavelength photoacoustic absorption spectrometer (PASS) and PSAP. Both instruments were operated on the same sampling inlet at the Department of Energy's Atmospheric Radiation Measurement program's Southern Great Plains (SGP) user facility in Oklahoma. We implemented the two most commonly used analytical correction algorithms, namely, Virkkula (2010) and the average of Virkkula (2010) and Ogren (2010)–Bond et al. (1999) as well as a random forest regression (RFR) machine learning algorithm to predict Babs values from the PSAP's filter-based measurements. The predicted Babs was compared against the reference Babs measured by the PASS. The RFR algorithm performed the best by yielding the lowest root mean square error of prediction. The algorithm was trained using input datasets from the PSAP (transmission and uncorrected absorption coefficient), a co-located nephelometer (scattering coefficients), and the Aerosol Chemical Speciation Monitor (mass concentration of non-refractory aerosol particles). A revised form of the Virkkula (2010) algorithm suitable for the SGP site has been proposed; however, its performance yields approximately 2-fold errors when compared to the RFR algorithm. To generalize the accuracy and applicability of our proposed RFR algorithm, we trained and tested it on a dataset of laboratory measurements of combustion aerosols. Input variables to the algorithm included the aerosol number size distribution from the Scanning Mobility Particle Sizer, absorption coefficients from the filter-based Tricolor Absorption Photometer, and scattering coefficients from a multiwavelength nephelometer. The RFR algorithm predicted Babs values within 5 % of the reference Babs measured by the multiwavelength PASS during the laboratory experiments. Thus, we show that machine learning approaches offer a promising path to correct for biases in long-term filter-based absorption datasets and accurately quantify their variability and trends needed for robust radiative forcing determination.

54 ENVIRONMENTAL SCIENCES↗

Robust Data-Driven Predictive Run-to-Run Control for Automated Serial Sectioning

This letter presents a one-step predictive run-to-run controller (R2R-MPC) for the automation of mechanical serial sectioning (MSS), a destructive material analysis process. To address the inherent uncertainty and disturbances in the MSS process, a robust closed-loop approach is presented. Here, the robust R2R-MPC models the uncertainty of the MSS process using a linear differential inclusion. As an analytical model of the MSS process is unavailable, the differential inclusion is identified from historical data. The R2R-MPC is posed as an optimization problem that computes incremental changes to the control input which minimize the worst-case material removal errors. This optimization-based controller is combined with a run-to-run controller to provide integral action that rejects constant disturbances and tracks constant reference removal rates. To demonstrate the efficacy of our robust R2R-MPC, we present simulation results which compare the presented controller with a conventional non-robust R2R.

42 ENGINEERING↗

Precise measurements of W - and Z -boson transverse momentum spectra with the ATLAS detector using pp collisions at $\sqrt{s} = 5.02$ TeV and 13 TeV

This paper describes measurements of the transverse momentum spectra of W and Z bosons produced in proton–proton collisions at centre-of-mass energies of $\sqrt{s}$ = 5.02 TeV and $\sqrt{s}$ = 13 TeV with the ATLAS experiment at the Large Hadron Collider. Measurements are performed in the electron and muon channels, W → $\ell$$v$ and Z → $\ell$$\ell$ ($\ell$ = e or μ), and for W events further separated by charge. The data were collected in 2017 and 2018, in dedicated runs with reduced instantaneous luminosity, and correspond to 255 and 338 pb -1 at $\sqrt{s}$ = 5.02 TeV and 13 TeV, respectively. These conditions optimise the reconstruction of the W-boson transverse momentum. The distributions observed in the electron and muon channels are unfolded, combined, and compared to QCD calculations based on parton shower Monte Carlo event generators and analytical resummation. The description of the transverse momentum distributions by Monte Carlo event generators is imperfect and shows significant differences largely common to W - , W + and Z production. The agreement is better at $\sqrt{s}$ = 5.02 TeV, especially for predictions that were tuned to Z production data at $\sqrt{s}$ = 7 TeV. Higher-order, resummed predictions based on DYTurbo generally match the data best across the spectra. Distribution ratios are also presented and test the understanding of differences between the production processes.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

QES-Plume v1.0: a Lagrangian dispersion model

Low-cost simulations providing accurate predictions of transport of airborne material in urban areas, vegetative canopies, and complex terrain are demanding because of the small-scale heterogeneity of the features influencing the mean flow and turbulence fields. Common models used to predict turbulent transport of passive scalars are based on the Lagrangian stochastic dispersion model. The Quick Environmental Simulation (QES) tool is a low-computational-cost framework developed to provide high-resolution wind and concentration fields in a variety of complex atmospheric-boundary-layer environments. Part of the framework, QES-Plume, is a Lagrangian dispersion code that uses a time-implicit integration scheme to solve the generalized Langevin equations which require mean flow and turbulence fields. Here, QES-Plume is driven by QES-Winds, a 3D fast-response model that computes mass-consistent wind fields around buildings, vegetation, and hills using empirical parameterizations, and QES-Turb, a local-mixing-length turbulence model. In this paper, the particle dispersion model is presented and validated against analytical solutions to examine QES-Plume’s performance under idealized conditions. In particular, QES-Plume is evaluated against a classical Gaussian plume model for an elevated continuous point-source release in uniform flow, the Lagrangian scaling of dispersion in isotropic turbulence, and a non-Gaussian plume model for an elevated continuous point-source release in a power-law boundary-layer flow. In these cases, QES-Plume yields a maximum relative error below 6 % when compared with analytical solutions. In addition, the model is tested against wind-tunnel data for a uniform array of cubical buildings. QES-Plume exhibits good agreement with the experiment with 99 % of matched zeros and 59 % of the predicted concentrations falling within a factor of 2 of the experimental concentrations. Furthermore, results also emphasize the importance of using high-quality turbulence models for particle dispersion in complex environments. Finally, QES-Plume demonstrates excellent computational performance.

58 GEOSCIENCES↗

Preventing Reverse Engineering of Critical Industrial Data with DIOD

Business analytics augmented by artificial intelligence and machine learning (AI/ML) have revolutionized the role of data in the modern world. In recent years, businesses have incorporated data into their decision-making process for better prediction, risk-assessment, content creation, etc. While such businesses often seek to leverage the full use of their data through third-party AI/ML services, they are often hampered by the risks of data leaks, reverse-engineering, stolen technology, etc. that often have disastrous consequences for businesses and their stakeholders alike. Thus, there arises a need for data masking prior to its transmission that obfuscates proprietary information while preserving the information relevant for AI/ML applications. In order to meet the needs of industrial data which are significantly different from those of data warehouses, previous work proposed an efficient time and space-scalable data masking paradigm known as the deceptive infusion of data (DIOD) methodology. The present work expands upon this work by leveraging existing reverse-engineering capabilities to facilitate the decomposition of industrial data into its proprietary and AI/ML-relevant parts, referred to as fundamental and inference metadata respectively. Both sets of metadata are further obfuscated in accordance with the DIOD methodology to create the DIOD rendition of the industrial data, which is rendered immune to reverse-engineering by discarding proprietary information and only preserving AI/ML-relevant information. Additionally, constraints of the original DIOD manuscript are relaxed using mutual information by configuring the methodology to the target AI/ML application to unlock the full potential of the DIOD methodology. As an example, data from a nuclear reactor is transformed into that from a nonlinear spring-mass system with different levels of data masking as required by the generic system and the target application.

97 MATHEMATICS AND COMPUTING↗

Heart Shape to Fracture Distance: Characterizing Hydraulic Fracture Propagation before Hits

Estimating the distance from the hydraulic fracture tip to the monitor well can be useful for fracture characterization, well spacing optimization, and preventing parent-child well interference. A heart-shaped signal is referred to as the extensional precursor of a fracture hit recorded by crosswell strain measurements and can serve as a vital tool for such estimation. This study incorporates the 3D displacement discontinuity method (DDM) to understand the impact of fracture geometry and monitor well offset on the heart-shaped signal’s characteristics. Results from numerical simulation and analytical solutions reveal a strong linear correlation between the spatial extent of the heart-shaped signal and the fracture tip distance. This relationship was further developed to predict tip distance using field data from the Hydraulic Fracture Test Site 2 (HFTS2). A reasonable approximation result from field data further validates the methodology. In addition, it is worth noting that the estimation accuracy depends on the ratio between fracture dimension and tip distance. The findings of this study offer a novel approach for real-time monitoring and characterizing hydraulic fracture propagation, which can be further used for well spacing optimization in unconventional and enhanced geothermal system reservoir development, as well as caprock integrity monitoring for carbon sequestration projects.

58 GEOSCIENCES↗

Integrating Applied Energy and BER Smart Data Capabilities to Develop a DOE Data Fabric for Energy-Water R&D

Focal Area(s): 1) Data acquisition and assimilation enabled by machine learning, AI, and advanced methods including experimental/network design/optimization, unsupervised learning (including deep learning), and hardware-related efforts involving AI (e.g., edge computing). Science Challenge: DOE R&D, including DOE’s Basic Energy Research (BER)’s Environmental Systems Science Division (EESSD) program and DOE’s applied energy research (AER) programs (EERE, FE, and NE) are producers and consumers of Earth systems datasets. This white paper focuses on the first topic area from the call in relation to how crosscutting resources and innovations from DOE’s EESSD and AER can be brought to bear to mutual benefit and more efficient energy-water, Earth system data resources through improved. The overarching challenge posed by this call focuses on how DOE can directly leverage artificial intelligence (AI) to engineer a substantial (paradigm-changing) improvement in Earth System Predictability? While stemming from DOE BER’s EESSD program, this is a challenge that is faced and also being addressed by DOE’s AER programs. Over the past decade plus, FE, EERE, and NE programs have made important strides towards addressing this need. These strides are in many ways highly complementary to EESSD’s MODEX efforts. Energy water systems spanning metocean to groundwater to surface water systems all are data driven whether for basic energy or applied energy. These are remote, multi-variate, complex natural, and in many cases engineered, systems. Key needs and challenges of both EESSD and AER include developing data-focused tools to enhance data search and discovery to fill in knowledge gaps (address sparse data challenge), and rapidly transform datasets, including disparate and multi-source data. Leveraging DOE on-premise computing (HPC, exascale) infrastructure supports the computing-intensive algorithms required to execute these data acquisition and transformation processes to derive enriched knowledge and data, driving AI/ML and big data analytics for these systems. The opportunity lies in combining BER and AER efforts to provide a more robust, advanced, efficient and complete computing data fabric to address energy-water data acquisition and assimilation needs which currently pose significant impediments to AI/ML predictions and research.

54 ENVIRONMENTAL SCIENCES↗

Faint and Fading Tails: The Fate of Stripped H i Gas in Virgo Cluster Galaxies

Although many galaxies in the Virgo cluster are known to have lost significant amounts of H i gas, only about a dozen features are known where the H i extends significantly outside its parent galaxy. Previous numerical simulations have predicted that H i removed by ram pressure stripping should have column densities far in excess of the sensitivity limits of observational surveys. We construct a simple model to try and quantify how many streams we might expect to detect. This accounts for the expected random orientation of the streams in position and velocity space as well as the expected stream length and mass of stripped H i. Using archival data from the Arecibo Galaxy Environment Survey, we search for any streams that might previously have been missed in earlier analyses. We report the confident detection of 10 streams as well as 16 other less-certain detections. We show that these well match our analytic predictions for which galaxies should be actively losing gas; however, the mass of the streams is typically far below the amount of missing H i in their parent galaxies, implying that a phase change and/or dispersal renders the gas undetectable. By estimating the orbital timescales, we estimate that dissolution rates of 1–10 M {sub ⊙} yr{sup −1} are able to explain both the presence of a few long, massive streams and the greater number of shorter, less-massive features.

79 ASTRONOMY AND ASTROPHYSICS↗