Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “machine learning tools”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

MLtool Python Code

Machine Learning (ML) is a subfield of Artificial Intelligence that gives computers the ability to learn from past data without being explicitly programmed. The predictive capabilities of ML models have already been used to facilitate several scientific breakthroughs. However, the practical application of ML is often limited due to the gaps in technical knowledge of its users. The common issue faced by many scientific researchers is the inability to choose the appropriate ML pipelines that are needed to treat real-world data, which is often sparse and noisy. To solve this problem, we have developed an automated Machine Learning tool (MLtool) that includes a set of ML algorithms and approaches to aid scientific researchers. The current version of MLtool is implemented as an object-oriented Python code that is easily extensible. It includes 44 different regression algorithms used to model data. MLtool helps users select the best model for their data, based on the scoring metrics used. Besides regression algorithms, MLtool also includes a suite of pre- and post-processing techniques such as missing value imputation, categorical variable encoding, input feature normalization, uncertainty quantification, exploratory data analysis (EDA), etc. MLtool was tested on several publicly available multi-dimensional data sets and was found capable of making accurate predictions.

Machine Learning↗

A traffic accident dataset for Chattanooga, Tennessee

This publication presents an annotated accident dataset which fuses traffic data from radar detection sensors, weather condition data, and light condition data with traffic accident data (as illustrated in Fig. 1) in a format that is easy to process using machine learning tools, databases, or data workflows. The purpose of this data is to analyze, predict, and detect traffic patterns when accidents occur. Each file contains a timeseries of traffic speeds, flows, and occupancies at the sensor nearest to the accident, as well as 5 neighboring sensors upstream and downstream. It also contains information about the accident type, date, and time. In addition to the accident data, we provide baseline data for typical traffic patterns during a given time of day. Overall, the dataset contains 6 months of annotated traffic data from November 2020 to April 2021. During this timeframe, and 361 accidents occurred in the monitored area around Chattanooga, Tennessee. This dataset served as the basis for a study on topology-aware automated accident detection for a companion publication [1].

97 MATHEMATICS AND COMPUTING↗

Graph theory inspired anomaly detection at the LHC

Designing model-independent anomaly detection algorithms for analyzing LHC data remains a central challenge in the search for new physics, due to the high dimensionality of collider events. In this work, we develop a graph autoencoder as an unsupervised, model-agnostic tool for anomaly detection, using the LHC Olympics dataset as a benchmark. By representing jet constituents as a graph, we introduce a method to systematically control the information available to the model through sparse graph constructions that serve as physically motivated inductive biases. Specifically, (1) we construct graph autoencoders based on locally rigid Laman graphs and globally rigid unique graphs, and (2) we explore the clustering of jet constituents into subjets to interpolate between high- and low-level input representations. We obtain the best performance, measured in terms of the Significance Improvement Characteristic curve for an intermediate level of subjet clustering and certain sparse unique graph constructions. We further investigate the role of graph connectivity in jet classification tasks. Our results demonstrate the potential of leveraging graph-theoretic insights to refine and increase the interpretability of machine learning tools for collider experiments.

Automation↗

Data-driven recovery of hidden physics in reduced order modeling of fluid flows

In this article, we introduce a modular hybrid analysis and modeling (HAM) approach to account for hidden physics in reduced order modeling (ROM) of parameterized systems relevant to fluid dynamics. The hybrid ROM framework is based on using first principles to model the known physics in conjunction with utilizing the data-driven machine learning tools to model the remaining residual that is hidden in data. This framework employs proper orthogonal decomposition as a compression tool to construct orthonormal bases and Galerkin projection (GP) as a model to build the dynamical core of the system. Our proposed methodology hence compensates structural or epistemic uncertainties in models and utilizes the observed data snapshots to compute true modal coefficients spanned by these bases. The GP model is then corrected at every time step with a data-driven rectification using a long short-term memory (LSTM) neural network architecture to incorporate hidden physics. A Grassmannian manifold approach is also adopted for interpolating basis functions to unseen parametric conditions. The control parameter governing the system's behavior is thus implicitly considered through true modal coefficients as input features to the LSTM network. The effectiveness of the HAM approach is then discussed through illustrative examples that are generated synthetically to take hidden physics into account. Furthermore, our approach thus provides insights addressing a fundamental limitation of the physics-based models when the governing equations are incomplete to represent underlying physical processes.

42 ENGINEERING↗

Analysis of Random Forest Modeling Strategies for Multi-Step Wind Speed Forecasting

Although the random forest (RF) model is a powerful machine learning tool that has been utilized in many wind speed/power forecasting studies, there has been no consensus on optimal RF modeling strategies. This study investigates three basic questions which aim to assist in the discernment and quantification of the effects of individual model properties, namely: (1) using a standalone RF model versus using RF as a correction mechanism for the persistence approach, (2) utilizing a recursive versus direct multi-step forecasting strategy, and (3) training data availability on model forecasting accuracy from one to six hours ahead. These questions are investigated utilizing data from the FINO1 offshore platform and Atmospheric Radiation Measurement (ARM) Southern Great Plains (SGP) C1 site, and testing results are compared to the persistence method. At FINO1, due to the presence of multiple wind farms and high inter-annual variability, RF is more effective as an error-correction mechanism for the persistence approach. The direct forecasting strategy is seen to slightly outperform the recursive strategy, specifically for forecasts three or more steps ahead. Finally, increased data availability (up to ~8 equivalent years of hourly training data) appears to continually improve forecasting accuracy, although changing environmental flow patterns have the potential to negate such improvement. We hope that the findings of this study will assist future researchers and industry professionals to construct accurate, reliable RF models for wind speed forecasting.

54 ENVIRONMENTAL SCIENCES↗

Gene network centrality analysis identifies key regulators coordinating day-night metabolic transitions in Synechococcus elongatus PCC 7942 despite limited accuracy in predicting direct regulator-gene interactions

Synechococcus elongatus PCC 7942 is a model organism for studying circadian regulation and bioproduction, where precise temporal control of metabolism significantly impacts photosynthetic efficiency and CO 2 -to-bioproduct conversion. Despite extensive research on core clock components, our understanding of the broader regulatory network orchestrating genome-wide metabolic transitions remains incomplete. We address this gap by applying machine learning tools and network analysis to investigate the transcriptional architecture governing circadian-controlled gene expression. While our approach showed moderate accuracy in predicting individual transcription factor-gene interactions - a common challenge with real expression data - network-level topological analysis successfully revealed the organizational principles of circadian regulation. Our analysis identified distinct regulatory modules coordinating day-night metabolic transitions, with photosynthesis and carbon/nitrogen metabolism controlled by day-phase regulators, while nighttime modules orchestrate glycogen mobilization and redox metabolism. Through network centrality analysis, we identified potentially significant but previously understudied transcriptional regulators: HimA as a putative DNA architecture regulator, and TetR and SrrB as potential coordinators of nighttime metabolism, working alongside established global regulators RpaA and RpaB. This work demonstrates how network-level analysis can extract biologically meaningful insights despite limitations in predicting direct regulatory interactions. The regulatory principles uncovered here advance our understanding of how cyanobacteria coordinate complex metabolic transitions and may inform metabolic engineering strategies for enhanced photosynthetic bioproduction from CO 2 .

59 BASIC BIOLOGICAL SCIENCES↗

PythonFOAM: In-situ data analyses with OpenFOAM and Python

Here, we outline the development of a general-purpose Python-based data analysis tool for OpenFOAM. Our implementation relies on the construction of OpenFOAM applications that have bindings to data analysis libraries in Python. Double precision data in OpenFOAM is cast to a NumPy array using the NumPy C-API and Python modules may then be used for arbitrary data analysis and manipulation on flow-field information. We highlight how the proposed wrapper may be used for an in-situ online singular value decomposition (SVD) implemented in Python and accessed from the OpenFOAM solver PimpleFOAM. Here, 'in-situ' refers to a programming paradigm that allows for a concurrent computation of the data analysis on the same computational resources utilized for the partial differential equation solver. In addition, to demonstrate parallel deployments, we deploy a distributed SVD, which collects snapshot data across the ranks of a distributed simulation to compute the global left singular vectors. Crucially, both OpenFOAM and Python share the same message passing interface (MPI) communicator for this deployment which allows Python objects and functions to exchange NumPy arrays across ranks. Subsequently, we provide scaling assessments of this distributed SVD on multiple nodes of Intel Broadwell and KNL architectures for canonical test cases such as the large eddy simulations of a backward facing step and a channel flow at friction Reynolds number of 395. Finally, we demonstrate the deployment of a deep neural network for compressing the flow-field information using an autoencoder to demonstrate an ability to use state-of-the-art machine learning tools in the Python ecosystem.

97 MATHEMATICS AND COMPUTING↗

Data‐Driven Safety Risk Prediction of Lithium‐Ion Battery

Abstract Inevitable safety issues have pushed battery engineers to become more conservative in battery system design; however, battery‐involved accidents still frequently are reported in headlines. Identifying, understanding, and predicting safety risks have become priorities to further accelerate technology and industry development. However, diverse loading scenarios, significantly varied stress‐induced short circuit mechanisms, and highly coupled mechanical–electrochemical safety behaviors have remained grand challenges. Herein, the safety risk is termed as the probability of the mechanical triggering of an internal short circuit, to reflect the safety related behaviors of lithium‐ion batteries. Based on a mechanical model and experimental results, a sufficient dataset is generated consisting of strain states and their corresponding safety risks, covering both cylindrical and pouch cells, various states of charges, and loading conditions. Machine‐learning tools combined with the established finite element mechanical model are applied to predict the safety risks of the cells. The results achieve a high level of accuracy on the test data (the relative error of the average short circuit prediction deviation is less than 6.2%.). This work underpins the safety risk concept and highlights the promise of physics combined with data‐driven modeling methodology to predict the safety behaviors of energy storage systems.

Jia, Yikai↗

A machine learning approach for clinker quality prediction and nonlinear model predictive control design for a rotary cement kiln

Abstract Cement manufacturing is energy‐intensive (5Gj/t) and comprises a significant portion of the energy footprint of concrete systems. Incorporating modern monitoring, simulation and control systems will allow lower energy use, lower environmental impact, and lower costs of this widely used construction material. One of the goals of the CESMII roadmap project on the Smart Manufacturing of Cement included developing an analytical process model for clinker quality that includes the chemistry of the kiln feed and accounts for critical process variables. This predictive model will be used in nonlinear model predictive control system designed to significantly reduce process energy use while maintaining or improving product quality. In the cement manufacturing plant used in this study, the kiln feed (meal) is tested every 12 h and used to estimate the mineral composition of the cement kiln output (clinker) using the stoichiometry‐based Bogue's model and the expertise of the plant operators. During kiln operation, kiln output (clinker) is sampled and tested every 2 h to measure its chemical and mineral composition. The predicted and measured values of the clinker composition are used by the plant operators to adjust the kiln input stream and the production process characteristics to maintain stable operation and uniform product quality. However, the time delay between prediction and testing, along with inaccuracies inherent in the Bogue's model have made any process changes designed to minimize energy use problematic, especially in‐light of potential clinker quality issues that process changes often pose. A new analytical model that integrates quality information and process operation information has been developed from data collected from 2 years of production from an operating cement facility. To make the model fuel‐type‐independent, consumed heat energy was computed in the model instead of fuel type and amount. A Feedforward Network was trained and tailored from collected data. Many data‐based simulations were conducted to quantitatively evaluate the proposed model and the 5‐fold cross‐validation procedure was used to test the models. The resulting predictive model was shown to have a low root mean square error (MSE) with respect to the estimated clinker mineral composition compared to that using the industry standard “Bogue’ model”. The end goal of this work was to develop a single machine learning tool that allows the use of quality control data and process control variables to improve energy efficiency of the process in a continuous fashion. The proposed nonlinear model predictive control system (NMPC) can generate predicted kiln production characteristics based on manipulated variables in manner that accurately follows the target product quality values. Simulation results also show that the proposed model produced accurate predictions of kiln outputs that fell within the required constraints, while manipulating control variables within typical operational ranges.

Ali, Asem M.↗

Search for heavy scalar resonances decaying to Lorentz-boosted Higgs and Higgs-like bosons in the $\mathrm{b}\overline{\mathrm{b}}4\mathrm{q}$ final state at $\sqrt{s}=13$ TeV

A search is performed for a heavy scalar resonance X decaying to a Higgs boson (H) and a Higgs-like scalar boson (Y) in the two bottom quark (H → $b\bar{b}$) and four quark (Y → VV → 4q) final state, where V denotes a W or Z boson. Masses of the X between 900 and 4000 GeV and the Y between 60 and 2800 GeV are considered. The search is performed in data collected by the CMS experiment at the CERN LHC from proton-proton collisions at 13 TeV center-of-mass energy, with a data set corresponding to a total integrated luminosity of 138 fb −1 . It targets the Lorentz-boosted regime, in which the products of the H → $b\bar{b}$ decay can be reconstructed as a single large-area jet, and those from the Y → VV → 4q decay as either one Y → 4q or two jets V → H → $q\bar{q}$. Jet identification and mass reconstruction exploit machine-learning tools, including a novel attention-based “particle transformer” for Y → 4q identification. No significant excess is observed in the data above the standard model background expectation. Upper limits on the product of production cross section and branching fraction as low as 0.2 fb are derived at 95% confidence level for various mass hypotheses. This is the first search at the LHC for scalar resonances in the all-hadronic H → $b\bar{b}$VV decay channel.

Beyond Standard Model↗

Anticipating gelation and vitrification with medium amplitude parallel superposition (MAPS) rheology and artificial neural networks

Abstract Anticipating qualitative changes in the rheological response of complex fluids (e.g., a gelation or vitrification transition) is an important capability for processing operations that utilize such materials in real-world environments. One class of complex fluids that exhibits distinct rheological states are soft glassy materials such as colloidal gels and clay dispersions, which can be well characterized by the soft glassy rheology (SGR) model. We first solve the model equations for the time-dependent, weakly nonlinear response of the SGR model. With this analytical solution, we show that the weak nonlinearities measured via medium amplitude parallel superposition (MAPS) rheology can be used to anticipate the rheological aging transitions in the linear response of soft glassy materials. This is a rheological version of a technique called structural health monitoring used widely in civil and aerospace engineering. We design and train artificial neural networks (ANNs) that are capable of quickly inferring the parameters of the SGR model from the results of sequential MAPS experiments. The combination of these data-rich experiments and machine learning tools to provide a surrogate for computationally expensive viscoelastic constitutive equations allows for rapid experimental characterization of the rheological state of soft glassy materials. We apply this technique to an aging dispersion of Laponite ® clay particles approaching the gel point and demonstrate that a trained ANN can provide real-time detection of transitions in the nonlinear response well in advance of incipient changes in the linear viscoelastic response of the system.

Lennon, Kyle R. (ORCID:0000000212515461)↗

Long-time integration of parametric evolution equations with physics-informed DeepONets

Ordinary and partial differential equations (ODEs/PDEs) play a paramount role in analyzing and simulating complex dynamic processes across all corners of science and engineering. In recent years machine learning tools are aspiring to introduce new effective ways of simulating such equations, however existing approaches are not able to reliably return stable and accurate predictions across long temporal horizons. We aim to address this challenge by introducing an effective framework for learning evolution operators that map random initial conditions to associated ODE/PDE solutions within a short time interval. Such operators can be parametrized by deep neural networks that are trained in an entirely self-supervised manner without requiring one to generate any paired input-output observations. Global long-time predictions across a range of initial conditions can be then obtained by iteratively evaluating the trained model using each prediction as the initial condition for the next evaluation step. Here, this introduces a new approach to temporal domain decomposition that is shown to be effective in performing accurate long-time simulations for a wide range of parametric ODE and PDE systems, from wave propagation, to reaction-diffusion dynamics and stiff chemical kinetics, introducing a new way of rapidly emulating non-equilibrium processes in science and engineering.

97 MATHEMATICS AND COMPUTING↗

Model calibration of the liquid mercury spallation target using evolutionary neural networks and sparse polynomial expansions

The mercury constitutive model predicting the strain and stress in the target vessel plays a central role in improving the lifetime prediction and future target designs of the mercury targets at the Spallation Neutron Source. We leverage the experiment strain data collected over multiple years to improve the mercury constitutive model through a combination of large scale simulations of the target behavior and the use of machine learning tools for parameter estimation. We present two interdisciplinary approaches for surrogate-based model calibration of expensive simulations using evolutionary neural networks and sparse polynomial expansions. The newly calibrated simulations achieve 7% average improvement on the prediction accuracy and 8% reduction in mean absolute error compared to previously reported reference parameters, with some individual sensors experiencing up to 30% improvement. The calibrated simulations can aid in fatigue analysis to estimate the mercury target lifetime, which reduces abrupt failure and saves tremendous amount of costs.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Dynamical coupled-channel models for hadron dynamics

Dynamical coupled-channel (DCC) approaches parametrize the interactions and dynamics of two and more hadrons and their response to different electroweak probes. The inclusion of unitarity, three-body channels, and other properties from scattering theory allows for a reliable extraction of resonance spectra and their properties from data. Here, we review the formalism and application of the ANL-Osaka, the Juelich-Bonn-Washington, and other DCC approaches in the context of light baryon resonances from meson, (virtual) photon, and neutrino-induced reactions, as well as production reactions, strange baryons, light mesons, heavy meson systems, exotics, and baryon-baryon interactions. Finally, we also provide a connection of the formalism to study finite-volume spectra obtained in Lattice QCD, and review applications involving modern statistical and machine learning tools.

Amplitude analysis↗

Machine-learning-aided density functional theory calculations of stacking fault energies in steel

A combined large-scale first principles approach with machine learning and materials informatics is proposed to quickly sweep the chemistry-composition space of advanced high strength steels (AHSS). AHSS are composed of iron and key alloying elements such as aluminum and manganese. A systematic exploration of the distribution of aluminum and manganese atoms in iron is used to investigate low stacking fault energies configurations using first principles calculations. To overcome the computational cost of exploring the composition space, this process is sped up using an automated machine learning tool: DeepHyper. Here our results predict that it is energetically favorable for Al to stay away from a stacking fault, but Mn atoms do not affect the stacking fault energy and can stay in the vicinity of the fault. The distribution of Al and Mn atoms in systems containing stacking faults and the effects of their interactions on the equilibrium distribution are systematically analyzed.

36 MATERIALS SCIENCE↗

Polymers for Extreme Conditions Designed Using Syntax-Directed Variational Autoencoders

We report the design/discovery of new materials is highly nontrivial owing to the near-infinite possibilities of material candidates and multiple required property/performance objectives. Thus, machine learning tools are now commonly employed to virtually screen material candidates with desired properties by learning a theoretical mapping from material-to-property space, referred to as the forward problem. However, this approach is inefficient and severely constrained by the candidates that the human imagination can conceive. Thus, in this work on polymers, we tackle the materials discovery challenge by solving the inverse problem: directly generating candidates that satisfy desired property/performance objectives. We utilize syntax-directed variational autoencoders (VAE) in tandem with Gaussian process regression (GPR) models to discover polymers expected to be robust under three extreme conditions: (1) high temperatures, (2) high electric field, and (3) high temperature and high electric field, useful for critical structural, electrical, and energy storage applications. This approach to learn from and augment) human ingenuity is general and can be extended to discover polymers with other targeted properties and performance measures.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

DancePartner: Python Package to Mine Multiomics Relationship Networks from Literature and Databases

A goal of multi-omics experiments is to understand how mechanistic molecular biology is altered between conditions, typically a control group and experimental groups. Oftentimes this involves studying changes in biomolecule relationships (e.g. interactions, metabolic relationships) of several types of biomolecules (e.g. proteins, lipids, metabolites). Though several databases contain relationships between biomolecules, understudied species may have little to no relationship information in databases and thus must be mined from literature. There are several challenges to literature mining, including automated full-text extraction, duplicate biomolecule term collapsing, and implementing complex machine learning tools. To make relationship extraction more accessible to the community, a python package called DancePartner was developed to allow for the extraction of relationships from literature and databases, with functions to map biomolecule synonyms to standardized identifiers and visualize and characterize the resulting multi-omics network. Here, in this study, an example dataset involving Caenorhabditis elegans is presented, where relationships are mined from 1443 publications using DancePartner. These relationships are combined with relationships from KEGG, WikiPathways, UniProt, and LipidMaps, and visualized.

BERT↗

Dynamic Multiplexed Control and Modeling of Optogenetic Systems Using the High-Throughput Optogenetic Platform, Lustro

The ability to control cellular processes using optogenetics is inducer-limited, with most optogenetic systems responding to blue light. To address this limitation, we leverage an integrated framework combining Lustro, a powerful high-throughput optogenetics platform, and machine learning tools to enable multiplexed control over blue light-sensitive optogenetic systems. Specifically, we identify light induction conditions for sequential activation as well as preferential activation and switching between pairs of light-sensitive split transcription factors in the budding yeast, Saccharomyces cerevisiae. We use the high-throughput data generated from Lustro to build a Bayesian optimization framework that incorporates data-driven learning, uncertainty quantification, and experimental design to enable the prediction of system behavior and the identification of optimal conditions for multiplexed control. This work lays the foundation for designing more advanced synthetic biological circuits incorporating optogenetics, where multiple circuit components can be controlled using designer light induction programs, with broad implications for biotechnology and bioengineering.

59 BASIC BIOLOGICAL SCIENCES↗