Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “machine learning tools”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Polymers for Extreme Conditions Designed Using Syntax-Directed Variational Autoencoders

We report the design/discovery of new materials is highly nontrivial owing to the near-infinite possibilities of material candidates and multiple required property/performance objectives. Thus, machine learning tools are now commonly employed to virtually screen material candidates with desired properties by learning a theoretical mapping from material-to-property space, referred to as the forward problem. However, this approach is inefficient and severely constrained by the candidates that the human imagination can conceive. Thus, in this work on polymers, we tackle the materials discovery challenge by solving the inverse problem: directly generating candidates that satisfy desired property/performance objectives. We utilize syntax-directed variational autoencoders (VAE) in tandem with Gaussian process regression (GPR) models to discover polymers expected to be robust under three extreme conditions: (1) high temperatures, (2) high electric field, and (3) high temperature and high electric field, useful for critical structural, electrical, and energy storage applications. This approach to learn from and augment) human ingenuity is general and can be extended to discover polymers with other targeted properties and performance measures.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

DancePartner: Python Package to Mine Multiomics Relationship Networks from Literature and Databases

A goal of multi-omics experiments is to understand how mechanistic molecular biology is altered between conditions, typically a control group and experimental groups. Oftentimes this involves studying changes in biomolecule relationships (e.g. interactions, metabolic relationships) of several types of biomolecules (e.g. proteins, lipids, metabolites). Though several databases contain relationships between biomolecules, understudied species may have little to no relationship information in databases and thus must be mined from literature. There are several challenges to literature mining, including automated full-text extraction, duplicate biomolecule term collapsing, and implementing complex machine learning tools. To make relationship extraction more accessible to the community, a python package called DancePartner was developed to allow for the extraction of relationships from literature and databases, with functions to map biomolecule synonyms to standardized identifiers and visualize and characterize the resulting multi-omics network. Here, in this study, an example dataset involving Caenorhabditis elegans is presented, where relationships are mined from 1443 publications using DancePartner. These relationships are combined with relationships from KEGG, WikiPathways, UniProt, and LipidMaps, and visualized.

BERT↗

Dynamic Multiplexed Control and Modeling of Optogenetic Systems Using the High-Throughput Optogenetic Platform, Lustro

The ability to control cellular processes using optogenetics is inducer-limited, with most optogenetic systems responding to blue light. To address this limitation, we leverage an integrated framework combining Lustro, a powerful high-throughput optogenetics platform, and machine learning tools to enable multiplexed control over blue light-sensitive optogenetic systems. Specifically, we identify light induction conditions for sequential activation as well as preferential activation and switching between pairs of light-sensitive split transcription factors in the budding yeast, Saccharomyces cerevisiae. We use the high-throughput data generated from Lustro to build a Bayesian optimization framework that incorporates data-driven learning, uncertainty quantification, and experimental design to enable the prediction of system behavior and the identification of optimal conditions for multiplexed control. This work lays the foundation for designing more advanced synthetic biological circuits incorporating optogenetics, where multiple circuit components can be controlled using designer light induction programs, with broad implications for biotechnology and bioengineering.

59 BASIC BIOLOGICAL SCIENCES↗

An Efficient Bayesian Approach to Learning Droplet Collision Kernels: Proof of Concept Using “Cloudy,” a New n -Moment Bulk Microphysics Scheme

The small-scale microphysical processes governing the formation of precipitation particles cannot be resolved explicitly by cloud resolving and climate models. Instead, they are represented by microphysics schemes that are based on a combination of theoretical knowledge, statistical assumptions, and fitting to data (“tuning”). Historically, tuning was done in an ad hoc fashion, leading to parameter choices that are not explainable or repeatable. Recent work has treated it as an inverse problem that can be solved by Bayesian inference. The posterior distribution of the parameters given the data—the solution of Bayesian inference—is found through computationally expensive sampling methods, which require over $\mathcal{O}$(10 5 ) evaluations of the forward model; this is prohibitive for many models. We present a proof of concept of Bayesian learning applied to a new bulk microphysics scheme named “Cloudy,” using the recently developed Calibrate-Emulate-Sample (CES) algorithm. Cloudy models collision-coalescence and collisional breakup of cloud droplets with an adjustable number of prognostic moments and with easily modifiable assumptions for the cloud droplet mass distribution and the collision kernel. The CES algorithm uses machine learning tools to accelerate Bayesian inference by reducing the number of forward evaluations needed to $\mathcal{O}$(10 2 ). It also exhibits a smoothing effect when forward evaluations are polluted by noise. In a suite of perfect-model experiments, we show that CES enables computationally efficient Bayesian inference of parameters in Cloudy from noisy observations of moments of the droplet mass distribution. In an additional imperfect-model experiment, a collision kernel parameter is successfully learned from output generated by a Lagrangian particle-based microphysics model.

54 ENVIRONMENTAL SCIENCES↗

Learning perturbation-inducible cell states from observability analysis of transcriptome dynamics

Abstract A major challenge in biotechnology and biomanufacturing is the identification of a set of biomarkers for perturbations and metabolites of interest. Here, we develop a data-driven, transcriptome-wide approach to rank perturbation-inducible genes from time-series RNA sequencing data for the discovery of analyte-responsive promoters. This provides a set of biomarkers that act as a proxy for the transcriptional state referred to as cell state. We construct low-dimensional models of gene expression dynamics and rank genes by their ability to capture the perturbation-specific cell state using a novel observability analysis. Using this ranking, we extract 15 analyte-responsive promoters for the organophosphate malathion in the underutilized host organism Pseudomonas fluorescens SBW25. We develop synthetic genetic reporters from each analyte-responsive promoter and characterize their response to malathion. Furthermore, we enhance malathion reporting through the aggregation of the response of individual reporters with a synthetic consortium approach, and we exemplify the library’s ability to be useful outside the lab by detecting malathion in the environment. The engineered host cell, a living malathion sensor, can be optimized for use in environmental diagnostics while the developed machine learning tool can be applied to discover perturbation-inducible gene expression systems in the compendium of host organisms.

59 BASIC BIOLOGICAL SCIENCES↗

Neural-network decoders for measurement induced phase transitions

Open quantum systems have been shown to host a plethora of exotic dynamical phases. Measurement-induced entanglement phase transitions in monitored quantum systems are a striking example of this phenomena. However, naive realizations of such phase transitions requires an exponential number of repetitions of the experiment which is practically unfeasible on large systems. Recently, it has been proposed that these phase transitions can be probed locally via entangling reference qubits and studying their purification dynamics. In this work, we leverage modern machine learning tools to devise a neural network decoder to determine the state of the reference qubits conditioned on the measurement outcomes. We show that the entanglement phase transition manifests itself as a stark change in the learnability of the decoder function. We study the complexity and scalability of this approach in both Clifford and Haar random circuits and discuss how it can be utilized to detect entanglement phase transitions in generic experiments.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

HypoRiPPAtlas as an Atlas of hypothetical natural products for mass spectrometry database search

Abstract Recent analyses of public microbial genomes have found over a million biosynthetic gene clusters, the natural products of the majority of which remain unknown. Additionally, GNPS harbors billions of mass spectra of natural products without known structures and biosynthetic genes. We bridge the gap between large-scale genome mining and mass spectral datasets for natural product discovery by developing HypoRiPPAtlas, an Atlas of hypothetical natural product structures, which is ready-to-use for in silico database search of tandem mass spectra. HypoRiPPAtlas is constructed by mining genomes using seq2ripp, a machine-learning tool for the prediction of ribosomally synthesized and post-translationally modified peptides (RiPPs). In HypoRiPPAtlas, we identify RiPPs in microbes and plants. HypoRiPPAtlas could be extended to other natural product classes in the future by implementing corresponding biosynthetic logic. This study paves the way for large-scale explorations of biosynthetic pathways and chemical structures of microbial and plant RiPP classes.

60 APPLIED LIFE SCIENCES↗

Data-driven studies of magnetic two-dimensional materials

We use a data-driven approach to study the magnetic and thermodynamic properties of van der Waals (vdW) layered materials. We investigate monolayers of the form A 2 B 2 X 6 , based on the known material Cr 2 Ge 2 Te 6 , using density functional theory (DFT) calculations and machine learning methods to determine their magnetic properties, such as magnetic order and magnetic moment. We also examine formation energies and use them as a proxy for chemical stability. We show that machine learning tools, combined with DFT calculations, can provide a computationally efficient means to predict properties of such two-dimensional (2D) magnetic materials. Our data analytics approach provides insights into the microscopic origins of magnetic ordering in these systems. For instance, we find that the X site strongly affects the magnetic coupling between neighboring A sites, which drives the magnetic ordering. Our approach opens new ways for rapid discovery of chemically stable vdW materials that exhibit magnetic behavior.

42 ENGINEERING↗

Active Learning-driven Quantitative Synthesis-Structure-Property Relations for Improving Performance and Revealing Active Sites of Nitrogen-Doped Carbon for the Hydrogen Evolution Reaction

While quantitative structure-properties relations (QSPRs) have been developed successfully in multiple fields, catalyst synthesis affects structure and in turn performance, making simple QSPRs inadequate. Furthermore, catalysts often have multiple active sites preventing one from obtaining insights into structure-property relations. Here, we develop a data-driven quantitative synthesis-structure-property relation (QS2PRs) methodology to elucidate correlations between catalyst synthesis conditions, structural properties as well as observed performance and to provide fundamental insights into active sites and a systematic way to optimize practical catalysts. Here, we demonstrate the approach to the synthesis of nitrogen-doped catalysts (NDC) made via pyrolysis for the performance of the electrochemical hydrogen evolution reaction (HER), quantified by the onset potential and the current density. We determine crystallinity, nitrogen species type and fraction, surface area, and pore structure of the NDC’s using XRD, XPS, and BET characterization. We demonstrated that an active learning-based optimization combined with various elementary machine learning tools (regression, principal component analysis, partial least squares) can efficiently identify optimum pyrolysis conditions to tune structural characteristics and performance with concomitant savings in materials and experimental time. Unlike previous reports on the importance of pyridinic or graphitic nitrogen, we discover that the electrochemical performance is not driven by a single catalyst property; rather, it arises from a multivariate influence of nitrogen dopants, pore structure and disorder in the NDC materials. Identification of active sites can help mechanistic understanding and further catalyst improvement.

42 ENGINEERING↗

Investigating magnetic van der Waals materials using data-driven approaches

In this work, we investigate magnetic monolayers of the form A i A ii B 4 X 8 based on the well-known intrinsic topological magnetic van der Waals (vdW) material MnBi 2 Te 4 (MBT) using first-principles calculations and machine learning techniques. We select an initial subset of structures to calculate the thermodynamic properties, electronic properties, such as the band gap, and magnetic properties, such as the magnetic moment and magnetic order using density functional theory (DFT). Data analytics approaches are used to gain insight into the microscopic origin of materials’ properties. The dependence of materials’ properties on chemical composition is also explored. For example, we find that the formation energy and magnetic moment depend largely on A and B sites whereas the band gap depends on all three sites. Finally, we employ machine learning tools to accelerate the search for novel vdW magnets in the MBT family with optimized properties. Finally, this study creates avenues for rapidly predicting novel materials with desirable properties that could enable applications in spintronics, optoelectronics, and quantum computing.

36 MATERIALS SCIENCE↗

Polyhydroxyalkanoates in emerging recycling technologies for a circular materials economy

Circular polymer systems, specifically polyesters operating through chemical and biological technologies, are approaching a critical moment of industrial adoption and scale-up feasibility. At the same time, polyhydroxyalkanoate (PHA) production, scale-up, and resulting material development is converging toward commodity applications. The current PHA end-of-life philosophy, however, focalizes leveraging inherent biodegradability to circumvent plastic waste accumulation. If indeed a substantial replacement of incumbent single-use plastics with PHA alternatives is to be met in commercial manufacture, we emphasize the importance of linking PHA development with feasible polymer recycling technologies. In other words, a PHA materials economy is significantly more carbon- and cost-favorable when efficient mechanical (reprocessing), chemical (deconstruction, depolymerization), or biological (enzymatic) recycling is prioritized over biodegradation or composting. In this perspective, we discuss strategies for PHA recyclable-by-design principles, guidable by developing machine learning tools, as well as material compatibility with closed-loop recycling technologies. Additionally, we posit compelling life-cycle assessment incentives for adopting polymer reclamation over competing pathways. Ultimately, we hope this narrative further inspires the alignment between PHA design with growing calls for a circular material economy.

36 MATERIALS SCIENCE↗

How can a diverse set of integral and semi-integral measurements inform identification of discrepant nuclear data?

Nuclear data are used for a variety of applications, including criticality safety, reactor performance, and material safeguards. Despite the breadth of use-cases, the effective neutron multiplication factor, keff, of ICSBEP critical assemblies are primarily used for nuclear data validation; these are sensitive to specific energy regions and nuclides and are unable to uniquely constrain nuclear data. As a consequence, general-purpose nuclear data libraries, such as ENDF/B-VIII.0, may have deficiencies that, while not apparent in criticality applications, negatively impact other applications, such as non-destructive analysis of special nuclear material and neutron diagnosed subcritical experiments. Recent work by the Experiments Underpinned by Computational Learning for Improvements in Nuclear Data (EUCLID) project developed a machine learning tool, RAFIEKI, which uses random forests and the SHAP metric to determine which nuclear data contribute most to predicted bias between measured and simulated responses (e.g. keff). This paper contrasts RAFIEKI analysis applied to keff only against RAFIEKI analysis with keff paired with either LLNL pulsed sphere measurements or subcritical benchmarks. Two examples show that a) including pulsed sphere measurements substantially increases 9Be nuclear data importance to bias between 2 and 15 MeV, and b) including subcritical benchmarks has the potential for disentangling compensating errors between 240Pu (n,el) and (n,il) cross-sections between 0.1 and 10 MeV. These results show that RAFIEKI analysis applied to response sets that include, but go beyond, keff can aid nuclear data evaluators in identifying issues in nuclear data.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Uncovering Where Compensating Errors Could Hide in ENDF/B-VIII.0

Unconstrained physics spaces between two or more nuclear data observables in a library occur when their values can be simultaneously adjusted without violating the uncertainties in either differential information or simulations of relevant integral experiments. Differential data are often too imprecise to fully bound all nuclear data observables of interest for application simulations. Integral data are simulated with combinations of nuclear data so that an error in one observable may be hidden by a counterbalancing error in another. In this manner compensating errors may lurk within nuclear data libraries and these errors have the potential to undermine the predictive power of neutron transport simulations, particularly in situations where there is no conclusive validation experiment that resembles the application of interest. The EUCLID project (Experiments Underpinned by Computational Learning for Improvements in Nuclear Data) developed a preliminary workflow to identify these unconstrained physics spaces by bringing together results from a large collection of integral experiments with their simulated counter-parts as well as differential information that have a one-to-one correspondence to nuclear data. This wealth of information is processed by machine learning tools for subsequent refinement by human experts. Here, we show how the EUCLID work-flow is executed by applying it first to 239 Pu and then to 9 Be nuclear data in ENDF/B-VIII.0.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Online Electron Reconstruction at CLAS12

Online reconstruction plays a crucial role in monitoring and in real-time analysis of high energy and nuclear physics experiments. A vital aspect of reconstruction algorithms is particle identification, which combines information from various detector components to determine the type of particle. Electron identification is particularly significant in electro-production nuclear physics experiments like the CLAS12 spectrometer at Jefferson Laboratory as it is essential in data recording. A machine learning approach has been developed for CLAS12 experiments to reconstruct and identify electrons by combining raw signals from multiple detector components at the data acquisition level. This method achieves high electron identification purity while maintaining nearly 100% efficiency. Furthermore, the machine learning tools operate at rates exceeding data acquisition speed, enabling the real-time electron reconstruction. This advancement significantly improves online analyses and monitoring capabilities for CLAS12 experiments.

Tyson,, Richard [Thomas Jefferson National Acceler↗

Regression Analysis with the Directed Infusion of Data

Integrating artificial intelligence and machine learning tools into industry necessitates large-scale collaborative efforts that ensure the robust and accurate execution of downstream analytics such as time series prediction, uncertainty quantification, grid optimization, and condition monitoring. However, concerns related to data privacy pervade the nuclear industry due to the proprietary nature of its data and the possibility of data leakage. Legacy techniques such as encryption often require the explicit transmission of data to trustworthy parties, thereby inviting data leakage concerns. The ideal collaboration scenario avoids the explicit dissemination of data/code while maintaining experimental fidelity, which is currently accomplished using various techniques such as trusted execution environments, homomorphic encryption, differential privacy, and multimatrix masking. These techniques, however, often necessitate a trade-off between trust, efficiency, and utility. This article extends a previously proposed technique called the directed infusion of data (DIOD) that ensures data privacy, allows for scalable obfuscation, and combats the risk of data leakage without compromising utility. The experiments discussed in this article examine a regression-type scenario using DIOD with the goal of preserving the inferential link between two variables. Using the point-kinetics equations, regression experiments compare the performance of a model trained using the original data to that of a model trained using the obfuscated data, which produced identical results. Our claim is further strengthened by an information theoretic proof and experiment, which showed that the inferential content between variables remains the same after obfuscation, thereby avoiding the required communication of the proprietary data.

47 - OTHER INSTRUMENTATION↗

The LHC Olympics 2020 a community challenge for anomaly detection in high energy physics

A new paradigm for data-driven, model-agnostic new physics searches at colliders is emerging, and aims to leverage recent breakthroughs in anomaly detection and machine learning. In order to develop and benchmark new anomaly detection methods within this framework, it is essential to have standard datasets. To this end, we have created the LHC Olympics 2020, a community challenge accompanied by a set of simulated collider events. Participants in these Olympics have developed their methods using an R&D dataset and then tested them on black boxes: datasets with an unknown anomaly (or not). Furthermore, methods made use of modern machine learning tools and were based on unsupervised learning (autoencoders, generative adversarial networks, normalizing flows), weakly supervised learning, and semi-supervised learning. This paper will review the LHC Olympics 2020 challenge, including an overview of the competition, a description of methods deployed in the competition, lessons learned from the experience, and implications for data analyses with future datasets as well as future colliders.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Pre-training Vision Models for the Classification of Alerts from Wide-field Time-domain Surveys

Modern wide-field time-domain surveys facilitate the study of transient, variable and moving phenomena by conducting image differencing and relaying alerts to their communities. Machine learning tools have been used on data from these surveys and their precursors for more than a decade, and convolutional neural networks (CNNs), which make predictions directly from input images, saw particularly broad adoption through the 2010s. Since then, continually rapid advances in computer vision have transformed the standard practices around using such models. It is now commonplace to use standardized architectures pre-trained on large corpora of everyday images (e.g., ImageNet). In contrast, time-domain astronomy studies still typically design custom CNN architectures and train them from scratch. Here, we explore the effects of adopting various pre-training regimens and standardized model architectures on the performance of alert classification. We find that the resulting models match or outperform a custom, specialized CNN like what is typically used for filtering alerts. Moreover, our results show that pre-training on galaxy images from Galaxy Zoo tends to yield better performance than pre-training on ImageNet or training from scratch. We observe that the design of standardized architectures are much better optimized than the custom CNN baseline, requiring significantly less time and memory for inference despite having more trainable parameters. On the eve of the Legacy Survey of Space and Time and other image-differencing surveys, these findings advocate for a paradigm shift in the creation of vision models for alerts, demonstrating that greater performance and efficiency, in time and in data, can be achieved by adopting the latest practices from the computer vision field.

79 ASTRONOMY AND ASTROPHYSICS↗

Selecting durable building envelope systems with machine learning assisted hygrothermal simulations database

Hygrothermal simulations provide insight into the energy performance and moisture durability of building envelope components under dynamic conditions. The inputs required for hygrothermal simulations are extensive, and carrying out simulations and analyses requires expert knowledge. An expert system, the Building Science Advisor (BSA), has been developed to predict the performance and select the energy-efficient and durable building envelope systems for different climates. The BSA consists of decision rules based on expert opinions and thousands of parametric simulation results for selected wall systems. The number of potential wall systems results in millions, too many to simulate all of them. We present how machine learning can help predict durability data, such as mold growth, while minimizing the number of simulations needed to run. The simulation results are used for training and validation of machine learning tools for predicting wall durability. We tested Artificial Neural Network (ANN) and Gradient Boosted Decision Trees (GBDT) for their applicability and model accuracy. Models developed with both methods showed adequate prediction performance (root mean square error of 0.195 and 0.209, respectively). Finally, we introduce how the information supports guidance for envelope design via an easy-to-use web-based tool that does not require the end-user to run hygrothermal simulations.

Salonvaara, Mikael↗