Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “ML T&E”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

FluxRETAP: a REaction TArget Prioritization genome-scale modeling technique for selecting genetic targets

MOTIVATION: Metabolic engineering is rapidly evolving as a result of new advances in synthetic biology tools and automation platforms that enable high throughput strain construction, as well as the development of machine learning tools (ML) for biology. However, selecting genetic engineering targets that effectively guide the metabolic engineering process is still challenging. ML can provide predictive power for synthetic biology, but current technical limitations prevent the independent use of ML approaches without previous biological knowledge. RESULTS: Here, we present FluxRETAP, a simple and computationally inexpensive method that leverages the prior mechanistic knowledge embedded in genome-scale models for suggesting targets for genetic overexpression, downregulation or deletion, with the final goal of increasing the production of a desired metabolite. This method can provide a list of desirable engineering targets that can be combined with current ML pipelines. FluxRETAP captured 100% of reaction targets experimentally verified to improve Escherichia coli isoprenol production, 50% of targets that experimentally improved taxadiene production in E. coli and ∼60% of genetic targets from a verified minimal constrained cut-set in Pseudomonas putida, while providing additional high priority targets that could be tested. Overall, FluxRETAP is an efficient algorithm for identifying a prioritized list of testable genetic and reaction targets. AVAILABILITY AND IMPLEMENTATION: FluxRETAP is implemented in python and released under the creative commons license. The implementation and code are freely available at: https://github.com/JBEI/FluxRETAP.

Czajka, Jeffrey J↗

Increasing the Reproducibility and Replicability of Supervised AI/ML in the Earth Systems Science by Leveraging Social Science Methods

Artificial intelligence (AI) and machine learning (ML) pose a challenge for achieving science that is both reproducible and replicable. The challenge is compounded in supervised models that depend on manually labeled training data, as they introduce additional decision-making and processes that require thorough documentation and reporting. We address these limitations by providing an approach to hand labeling training data for supervised ML that integrates quantitative content analysis (QCA)—a method from social science research. The QCA approach provides a rigorous and well-documented hand labeling procedure to improve the replicability and reproducibility of supervised ML applications in Earth systems science (ESS), as well as the ability to evaluate them. Specifically, the approach requires (a) the articulation and documentation of the exact decision-making process used for assigning hand labels in a “codebook” and (b) an empirical evaluation of the reliability” of the hand labelers. In this paper, we outline the contributions of QCA to the field, along with an overview of the general approach. We then provide a case study to further demonstrate how this framework has and can be applied when developing supervised ML models for applications in ESS. With this approach, we provide an actionable path forward for addressing ethical considerations and goals outlined by recent AGU work on ML ethics in ESS.

58 GEOSCIENCES↗

Highly Reversible Sodium Ion Batteries Enabled by Stable Electrolyte-Electrode Interphases

Sodium (Na) ion battery is a very promising technology for the alternative energy storage systems because of the abundance and low cost of Na element in the Earth’s crust. However, the limited cycle life and safety concerns still hinder its large-scale applications. Here, we report a nonflammable localized high concentration electrolyte (sodium bis(fluorosulfonyl)imide - triethyl phosphate/1,1,2,2-tetrafluoroethyl-2,2,3,3-tetrafluoropropyl ether (1:1.5:2 in molar ratio)), which enables a very high initial Coulombic efficiency (CE) of 97.8% for Na||Na-CNFM (O3-NaCu1/9Ni2/9Fe1/3Mn1/3O2) cells and stable cycling of Na||hard carbon (HC) cells with a capacity retention of 95.4% after 500 cycles. The HC||Na-CNFM full cells using this electrolyte retain 82.5% capacity after 200 cycles with a CE of ~99.9% compared to 48.4% capacity retention in the carbonate electrolyte (1 M NaPF6/EC+DMC (1:1 in weight)). The extremely high CE and stability of HC||Na-CNFM cells in this electrolyte can be attributed to the stable interphase layers formed on both HC anode and Na-CNFM cathode. These layers minimize undesirable reaction between HC and electrolyte, and block the dissolution of transition metal from cathode. The insight obtained in this work can be used to further improve cycling stability and safety of rechargeable batteries.

Jin, Yan↗

Machine Learning Prediction of Tritium‐Helium Groundwater Ages in the Central Valley, California, USA

Abstract Groundwater ages provides insight into recharge rates, flow velocities, and vulnerability to contaminants. The ability to predict groundwater ages based on more accessible parameters via Machine Learning (ML) would advance our ability to guide sustainable management of groundwater resources. In this study, ML models were trained and tested on a large data set of tritium concentrations and tritium‐helium groundwater ages from the California Central Valley, a large groundwater basin with complex land use, irrigation, and water management practices. The ML models were trained on 63 features, including location, well construction information, landscape characteristics, and climate variables, water chemistry, and stable isotopes. The Bagging regressor method can accurately classify (F1‐score = 0.91) groundwater samples as either modern or pre‐modern whereas the accuracy of the ML prediction of continuous tritium‐helium groundwater ages is limited and explains only of the variability in this data set. In general, ML groundwater age prediction relies mostly on features related to (a) the source of groundwater recharge, (b) contaminant history, (c) aquifer materials, (d) well construction, and (e) geochemical reactions along flow paths.

54 ENVIRONMENTAL SCIENCES↗

Machine learning for ultrasonic nondestructive examination of welding defects: A systematic review

Recent years have seen a substantial increase in the application of machine learning (ML) for automated analysis of nondestructive examination (NDE) data. One of the applications of interest is the use of ML for the analysis of data from in-service inspection of welds in nuclear power and other industries. These types of inspections are performed in accordance with criteria described in the ASME Boiler and Pressure Vessel Code and require the use of reliable NDE techniques. The rapid growth in ML methods and the diversity of possible approaches indicate a need to assess the current capabilities of ML and automated data analysis for NDE and identify any gaps or shortcomings in current ML technologies as applied to the automated analysis of NDE data. In particular, there is a need to determine the impact of ML on the NDE reliability. This paper discusses the findings from a literature survey on the current state of ML for the automated analysis of data from ultrasonic NDE of weld flaws. It discusses an overview of ultrasonic NDE as used for weld inspections in nuclear power and other industries. Herein, data sets and ML models used in the literature are summarized, along with a generally applicable workflow for ML. Findings on the capabilities, limitations and potential gaps in feature selection, data selection, and ML model optimization are discussed. The paper identified several needs for quantifying and validating the performance of ML methods for ultrasonic NDE, including the need for common data sets.

36 MATERIALS SCIENCE↗

Imaging biomarkers in neurodegeneration: current and future practices

There is an increasing role for biological markers (biomarkers) in the understanding and diagnosis of neurodegenerative disorders. The application of imaging biomarkers specifically for the in vivo investigation of neurodegenerative disorders has increased substantially over the past decades and continues to provide further benefits both to the diagnosis and understanding of these diseases. This review forms part of a series of articles which stem from the University College London/University of Gothenburg course “Biomarkers in neurodegenerative diseases”. In this review, we focus on neuroimaging, specifically positron emission tomography (PET) and magnetic resonance imaging (MRI), giving an overview of the current established practices clinically and in research as well as new techniques being developed. We will also discuss the use of machine learning (ML) techniques within these fields to provide additional insights to early diagnosis and multimodal analysis.

60 APPLIED LIFE SCIENCES↗

A dilatometer for volume-temperature determinations of liquids.

An instrument is described for the continuous volume measurement of small (3.7 ml) samples of liquids as a function of temperature. The sample is sealed in a stainless steel bellows chamber. Volume changes are measured with a linear variable differential transformer while temperature is measured by means of a thermistor and associated circuitry. Volume changes can be determined to better than .1 microliter ml over a temperature range of 10-60 C. The maximum sample volume change is 7%.

Rothman, J. E.↗

Maximum likelihood identification of aircraft stability and control derivatives

Application of a generalized identification method to flight test data analysis. The method is based on the maximum likelihood (ML) criterion and includes output error and equation error methods as special cases. Both the linear and nonlinear models with and without process noise are considered. The flight test data from lateral maneuvers of HL-10 and M2/F3 lifting bodies are processed to determine the lateral stability and control derivatives, instrumentation accuracies, and biases. A comparison is made between the results of the output error method and the ML method for M2/F3 data containing gusts. It is shown that better fits to time histories are obtained by using the ML method. The nonlinear model considered corresponds to the longitudinal equations of the X-22 VTOL aircraft. The data are obtained from a computer simulation and contain both process and measurement noise. The applicability of the ML method to nonlinear models with both process and measurement noise is demonstrated.

Mehra, R. K.↗

Challenges in the Verification of Reinforcement Learning Algorithms

Machine learning (ML) is increasingly being applied to a wide array of domains from search engines to autonomous vehicles. These algorithms, however, are notoriously complex and hard to verify. This work looks at the assumptions underlying machine learning algorithms as well as some of the challenges in trying to verify ML algorithms. Furthermore, we focus on the specific challenges of verifying reinforcement learning algorithms. These are highlighted using a specific example. Ultimately, we do not offer a solution to the complex problem of ML verification, but point out possible approaches for verification and interesting research opportunities.

Van Wesel, Perry↗

The ANI-1ccx and ANI-1x data sets, coupled-cluster and density functional theory properties for molecules

Abstract Maximum diversification of data is a central theme in building generalized and accurate machine learning (ML) models. In chemistry, ML has been used to develop models for predicting molecular properties, for example quantum mechanics (QM) calculated potential energy surfaces and atomic charge models. The ANI-1x and ANI-1ccx ML-based general-purpose potentials for organic molecules were developed through active learning; an automated data diversification process. Here, we describe the ANI-1x and ANI-1ccx data sets. To demonstrate data diversity, we visualize it with a dimensionality reduction scheme, and contrast against existing data sets. The ANI-1x data set contains multiple QM properties from 5 M density functional theory calculations, while the ANI-1ccx data set contains 500 k data points obtained with an accurate CCSD(T)/CBS extrapolation. Approximately 14 million CPU core-hours were expended to generate this data. Multiple QM calculated properties for the chemical elements C, H, N, and O are provided: energies, atomic forces, multipole moments, atomic charges, etc. We provide this data to the community to aid research and development of ML models for chemistry.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Predicting Small Molecule Transfer Free Energies by Combining Molecular Dynamics Simulations and Deep Learning

Accurately predicting small molecule partitioning and hydrophobicity is critical in the drug discovery process. There are many heterogeneous chemical environments within a cell and entire human body. For example, drugs must be able to cross the hydrophobic cellular membrane to reach their intracellular targets, and hydrophobicity is an important driving force for drug–protein binding. Atomistic molecular dynamics (MD) simulations are routinely used to calculate free energies of small molecules binding to proteins, crossing lipid membranes, and solvation but are computationally expensive. Machine learning (ML) and empirical methods are also used throughout drug discovery but rely on experimental data, limiting the domain of applicability. We present atomistic MD simulations calculating 15,000 small molecule free energies of transfer from water to cyclohexane. This large data set is used to train ML models that predict the free energies of transfer. We show that a spatial graph neural network model achieves the highest accuracy, followed closely by a 3D-convolutional neural network, and shallow learning based on the chemical fingerprint is significantly less accurate. A mean absolute error of ~4 kJ/mol compared to the MD calculations was achieved for our best ML model. We also show that including data from the MD simulation improves the predictions, tests the transferability of each model to a diverse set of molecules, and show multitask learning improves the predictions. This work provides insight into the hydrophobicity of small molecules and ML cheminformatics modeling, and our data set will be useful for designing and testing future ML cheminformatics methods.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Learning a general model of single phase flow in complex 3D porous media

Modeling effective transport properties of 3D porous media, such as permeability, at multiple scales is challenging as a result of the combined complexity of the pore structures and fluid physics—in particular, confinement effects which vary across the nanoscale to the microscale. While numerical simulation is possible, the computational cost is prohibitive for realistic domains, which are large and complex. Although machine learning (ML) models have been proposed to circumvent simulation, none so far has simultaneously accounted for heterogeneous 3D structures, fluid confinement effects, and multiple simulation resolutions. By utilizing numerous computer science techniques to improve the scalability of training, we have for the first time developed a general flow model that accounts for the pore-structure and corresponding physical phenomena at scales from Angstrom to the micrometer. Using synthetic computational domains for training, our ML model exhibits strong performance (R 2 = 0.9) when tested on extremely diverse real domains at multiple scales.

36 MATERIALS SCIENCE↗

A Review of Recent and Emerging Machine Learning Applications for Climate Variability and Weather Phenomena

Abstract Climate variability and weather phenomena can cause extremes and pose significant risk to society and ecosystems, making continued advances in our physical understanding of such events of utmost importance for regional and global security. Advances in machine learning (ML) have been leveraged for applications in climate variability and weather, empowering scientists to approach questions using big data in new ways. Growing interest across the scientific community in these areas has motivated coordination between the physical and computer science disciplines to further advance the state of the science and tackle pressing challenges. During a recently held workshop that had participants across academia, private industry, and research laboratories, it became clear that a comprehensive review of recent and emerging ML applications for climate variability and weather phenomena that can cause extremes was needed. This article aims to fulfill this need by discussing recent advances, challenges, and research priorities in the following topics: sources of predictability for modes of climate variability, feature detection, extreme weather and climate prediction and precursors, observation–model integration, downscaling, and bias correction. This article provides a review for domain scientists seeking to incorporate ML into their research. It also provides a review for those with some ML experience seeking to broaden their knowledge of ML applications for climate variability and weather.

54 ENVIRONMENTAL SCIENCES↗

Scientific machine learning for closure models in multiscale problems: A review

Here, closure problems are omnipresent when simulating multiscale systems, where some quantities and processes cannot be fully prescribed despite their effects on the simulation's accuracy. Recently, scientific machine learning approaches have been proposed as a way to tackle the closure problem, combining traditional (physics-based) modeling with data-driven (machine-learned) techniques, typically through enriching differential equations with neural networks. This paper reviews the different reduced model forms, distinguished by the degree to which they include known physics, and the different objectives of a priori and a posteriori learning. The importance of adhering to physical laws (such as symmetries and conservation laws) in choosing the reduced model form and choosing the learning method is discussed. The effect of spatial and temporal discretization and recent trends toward discretization-invariant models are reviewed. In addition, we make the connections between closure problems and several other research disciplines: inverse problems, Mori-Zwanzig theory, and multi-fidelity methods. In conclusion, much progress has been made with scientific machine learning approaches for solving closure problems, but many challenges remain. In particular, the generalizability and interpretability of learned models is a major issue that needs to be addressed further.

97 MATHEMATICS AND COMPUTING↗

Aqueous nitrite ion determination by selective reduction and gas phase nitric oxide chemiluminescence

An improved method of flow injection analysis for aqueous nitrite ion exploits the sensitivity and selectivity of the nitric oxide (NO) chemilluminescence detector. Trace analysis of nitrite ion in a small sample (5-160 microL) is accomplished by conversion of nitrite ion to NO by aqueous iodide in acid. The resulting NO is transported to the gas phase through a semipermeable membrane and subsequently detected by monitoring the photoemission of the reaction between NO and ozone (O3). Chemiluminescence detection is selective for measurement of NO, and, since the detection occurs in the gas-phase, neither sample coloration nor turbidity interfere. The detection limit for a 100-microL sample is 0.04 ppb of nitrite ion. The precision at the 10 ppb level is 2% relative standard deviation, and 60-180 samples can be analyzed per hour. Samples of human saliva and food extracts were analyzed; the results from a standard colorimetric measurement are compared with those from the new chemiluminescence method in order to further validate the latter method. A high degree of selectivity is obtained due to the three discriminating steps in the process: (1) the nitrite ion to NO conversion conditions are virtually specific for nitrite ion, (2) only volatile products of the conversion will be swept to the gas phase (avoiding turbidity or color in spectrophotometric methods), and (3) the NO chemiluminescence detector selectively detects the emission from the NO + O3 reaction. The method is free of interferences, offers detection limits of low parts per billion of nitrite ion, and allows the analysis of up to 180 microL-sized samples per hour, with little sample preparation and no chromatographic separation. Much smaller samples can be analyzed by this method than in previously reported batch analysis methods, which typically require 5 mL or more of sample and often need chromatographic separations as well.

NASA Discipline Environmental Health↗

Flux REaction TArget Prioritization (Flux RETAP) v1

Metabolic engineering is evolving rapidly as a result of new advances in synthetic biology and automation, as well as the irruption of machine learning (ML). ML has been shown to provide the predictive power synthetic biology lacked and needed, and to be able to effectively guide the metabolic engineering process. However, current technical limitations prevent the independent application of ML approaches to metabolic engineering without the use of previous biological knowledge in the form of a prioritized list of desirable engineering targets. Here, we present FluxRETAP, a simple and computationally inexpensive method that leverages the prior mechanistic knowledge embedded in genome-scale metabolic models (GSMs) for suggesting targets for genetic overexpression, downregulation or deletion, with the final goal of increasing metabolite production. FluxRETAP captured 100% of reaction targets experimentally verified to improve Escherichia coli isoprenol production in the literature accessible to us, 50% of targets that experimentally improved taxadiene production in E. coli and ~60% of genetic targets from a verified minimal constrained cut-set in Pseudomonas putida while providing additional high priority targets that could be tested. Overall, FluxRETAP is an efficient algorithm for identifying a prioritized list of testable genetic and reaction targets which can also be utilized in ML pipelines.

Czajka, Jeffrey [Battelle Memorial Institute, Paci↗

Earth System Reanalysis in Support of Climate Model Improvements

Recent climate model developments, established through increased model resolution, have led to substantial improvements in model simulations of the time-evolving, coupled Earth system and its subcomponents. However, regardless of resolution, climate models will always produce climate features and variability that differ from the real world and will be prone to biases. This is due to many remaining uncertainties, such as in parametric and structural model uncertainty, in the initial conditions prescribed, and in the prescribed (scenario) forcing which varies on decadal to centennial timescales. Further model improvements are expected to arise specifically from improved representation of physical processes realized through model-data fusion. This will create an unprecedented opportunity to better exploit a large array of Earth observations, from in situ measurements to weather radars and satellite observations, as the resolved scales of the models approach those of the observations. For this, climate DA will be the central tool to bring models and observations into consistency, by improving initial conditions, inferring uncertain model parameters and structure, and quantifying uncertainty. Generally, there will be advantages and complementarities of adjoint-based smoother approaches, ensemble-based filter approaches, or new ML-inspired approaches. Yet, the ever-increasing model resolution will present growing challenges arising from computational cost, calling for new ways of performing data assimilation and model optimization. Using the complementarity in a hybrid approach, blending tools and concepts from variational, ensemble and ML methods might be what is required in the future. In this context ML could be important to handle non-linear responses, and to better approximate non-Gaussian distributions.

54 ENVIRONMENTAL SCIENCES↗

Concentration transient analysis of antimony surface segregation during Si(100) molecular beam epitaxy

Antimony surface segregation during Si(100) molecular beam epitaxy (MBE) was investigated at temperatures T(sub s) = 515 - 800 C using concentration transient analysis (CTA). The dopant surface coverage Theta, bulk fraction gamma, and incorporation probability sigma during MBE were determined from secondary-ion mass spectrometry depth profiles of modulation-doped films. Programmed T(sub s) changes during growth were used to trap the surface-segregated dopant overlayer, producing concentration spikes whose integrated area corresponds to Theta. Thermal antimony doping by coevaporation was found to result in segregation strongly dependent on T(sub s) with Theta(sub Sb) values up to 0.9 monolayers (ML): in films doped with Sb(+) ions accelerated by 100 V, Theta(sub Sb) was less than or equal to 4 x 10(exp -3) ML. Surface segregation of coevaporated antimony was kinematically limited for the film growth conditions in these experiments.

Markert, L. C.↗