Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “machine learning algorithms”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Review of machine learning and deep learning models for toxicity prediction

The ever-increasing number of chemicals has raised public concerns due to their adverse effects on human health and the environment. To protect public health and the environment, it is critical to assess the toxicity of these chemicals. Traditional in vitro and in vivo toxicity assays are complicated, costly, and time-consuming and may face ethical issues. These constraints raise the need for alternative methods for assessing the toxicity of chemicals. Recently, due to the advancement of machine learning algorithms and the increase in computational power, many toxicity prediction models have been developed using various machine learning and deep learning algorithms such as support vector machine, random forest, k-nearest neighbors, ensemble learning, and deep neural network. This review summarizes the machine learning- and deep learning-based toxicity prediction models developed in recent years. Support vector machine and random forest are the most popular machine learning algorithms, and hepatotoxicity, cardiotoxicity, and carcinogenicity are the frequently modeled toxicity endpoints in predictive toxicology. It is known that datasets impact model performance. The quality of datasets used in the development of toxicity prediction models using machine learning and deep learning is vital to the performance of the developed models. The different toxicity assignments for the same chemicals among different datasets of the same type of toxicity have been observed, indicating benchmarking datasets is needed for developing reliable toxicity prediction models using machine learning and deep learning algorithms. This review provides insights into current machine learning models in predictive toxicology, which are expected to promote the development and application of toxicity prediction models in the future.

Research & Experimental Medicine↗

Materials Learning Algorithms (MALA): Scalable machine learning for electronic structure calculations in large-scale atomistic simulations

We present the Materials Learning Algorithms (MALA) package, a scalable machine learning framework designed to accelerate density functional theory (DFT) calculations suitable for large-scale atomistic simulations. Using local descriptors of the atomic environment, MALA models efficiently predict key electronic observables, including local density of states, electronic density, density of states, and total energy. The package integrates data sampling, model training and scalable inference into a unified library, while ensuring compatibility with standard DFT and molecular dynamics codes. We demonstrate MALA's capabilities with examples including boron clusters, aluminum across its solid-liquid phase boundary, and predicting the electronic structure of a stacking fault in a large beryllium slab. Scaling analyses reveal MALA's computational efficiency and identify bottlenecks for future optimization. With its ability to model electronic structures at scales far beyond standard DFT, MALA is well suited for modeling complex material systems, making it a versatile tool for advanced materials research.

Density functional theory↗

Research Data Alliance: Understanding Big Data Analytics Applications in Earth Science

The Research Data Alliance (RDA) enables data to be shared across barriers through focused working groups and interest groups, formed of experts from around the world - from academia, industry and government. Its Big Data Analytics (BDA) interest groups seeks to develop community based recommendations on feasible data analytics approaches to address scientific community needs of utilizing large quantities of data. BDA seeks to analyze different scientific domain applications (e.g. earth science use cases) and their potential use of various big data analytics techniques. These techniques reach from hardware deployment models up to various different algorithms (e.g. machine learning algorithms such as support vector machines for classification). A systematic classification of feasible combinations of analysis algorithms, analytical tools, data and resource characteristics and scientific queries will be covered in these recommendations. This contribution will outline initial parts of such a classification and recommendations in the specific context of the field of Earth Sciences. Given lessons learned and experiences are based on a survey of use cases and also providing insights in a few use cases in detail.

Riedel, Morris↗

GOOML (Geothermal Operational Optimization with Machine Learning) [SWR-23-01]

The Geothermal Operational Optimization with Machine Learning (GOOML) is a partnership between NREL and Upflow, NZ, awarded in response to the U.S. Department of Energy's Geothermal Technologies Office's Funding Opportunity Announcement (FOA) to expand the role of advanced analytics and automation in geothermal operations through machine learning. Partnering with industry (Contact Energy Limited ("Contact"), Ngati Tuwharetoa Geothermal Assets Limited ("NTGA"), Ormat Technologies Inc. ("Ormat") and Flow State Solutions Limited ("FSS"), GOOML was created to improve the operational efficiency of geothermal power plant steam fields through the analysis of historical operational data and the application of custom machine learning algorithms. NREL's contributions include machine learning, coding, and data management expertise as well as access to high-performance compute solutions. GOOML can increase geothermal operational efficiency through development of a digital system twin that can be utilized to provide optimal geothermal operating conditions for real-world geothermal fields. GOOML allows users to analyze field production histories in detail, develop models, and train machine learning algorithms to identify opportunities for increased geothermal efficiency, detect potential trouble, and allow predictive scenario modeling. Preliminary experiments have demonstrated a potential to increase total generation by as much as 12% through ML optimization of the utilization of existing steam field resources.

Buster, Grant↗

Improving the Freight Productivity of a Heavy-Duty, Battery Electric Truck by Intelligent Energy Management

This project aimed to enhance the range and reduce the operating costs of battery electric Class 8 trucks traveling over 250 miles daily. This was achieved through the development and implementation of an intelligent-Energy Management System (i-EMS) that leverages vehicle and operations data, physics-aware machine learning algorithms, and vehicle-to-cloud (V2C) connectivity. The project hypothesized that advanced machine learning algorithms and real-time data analytics could significantly improve the energy efficiency and range of these trucks. Key objectives included developing a physics-aware machine learning algorithm, implementing an i-EMS with V2C connectivity and physics-aware spatial data analytics (PSDA), and validating the system’s effectiveness with fleet partners HEB Companies and Murphy Logistics. Extensive data collection from vehicle operations, including vehicle characteristics, road conditions, and payload, was conducted. A machine learning algorithm was developed to predict energy consumption and enable proactive decision-making. The i-EMS was implemented on two Volvo VNR BEVs, with operators receiving charging and routing recommendations. Charging stations were installed at depot locations in Texas and Minnesota, with an additional on-route charger in Minnesota. Significant findings included a 14% range improvement for Murphy Logistics on a highway-driving eco-route and a 22% range improvement for HEB Companies on a city-driving eco-route. The i-EMS utilized rule-based methods and physics-based algorithms to predict and reduce energy consumption, with real-time monitoring and analysis through V2C connectivity enabling proactive decision-making. The project demonstrated the feasibility and economic viability of battery electric Class 8 trucks for long-haul operations, showcasing the potential of physics-aware machine learning in optimizing energy management. The successful implementation of the i-EMS in real-world scenarios validates its practical application and effectiveness, paving the way for the widespread adoption of battery electric vehicles in the freight transportation industry.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Modeling the Swift Bat Trigger Algorithm with Machine Learning

To draw inferences about gamma-ray burst (GRB) source populations based on Swift observations, it is essential to understand the detection efficiency of the Swift burst alert telescope (BAT). This study considers the problem of modeling the Swift / BAT triggering algorithm for long GRBs, a computationally expensive procedure, and models it using machine learning algorithms. A large sample of simulated GRBs from Lien et al. is used to train various models: random forests, boosted decision trees (with AdaBoost), support vector machines, and artificial neural networks. The best models have accuracies of greater than or equal to 97 percent (less than or equal to 3 percent error), which is a significant improvement on a cut in GRB flux, which has an accuracy of 89.6 percent (10.4 percent error). These models are then used to measure the detection efficiency of Swift as a function of redshift z, which is used to perform Bayesian parameter estimation on the GRB rate distribution. We find a local GRB rate density of n (sub 0) approaching 0.48 (sup plus 0.41) (sub minus 0.23) per cubic gigaparsecs per year with power-law indices of n (sub 1) approaching 1.7 (sup plus 0.6) (sub minus 0.5) and n (sub 2) approaching minus 5.9 (sup plus 5.7) (sub minus 0.1) for GRBs above and below a break point of z (redshift) (sub 1) approaching 6.8 (sup plus 2.8) (sub minus 3.2). This methodology is able to improve upon earlier studies by more accurately modeling Swift detection and using this for fully Bayesian model fitting.

gamma-ray burst: general – gamma-rays: general â↗

Modeling the Swift BAT Trigger Algorithm with Machine Learning

To draw inferences about gamma-ray burst (GRB) source populations based on Swift observations, it is essential to understand the detection efficiency of the Swift burst alert telescope (BAT). This study considers the problem of modeling the Swift BAT triggering algorithm for long GRBs, a computationally expensive procedure, and models it using machine learning algorithms. A large sample of simulated GRBs from Lien et al. (2014) is used to train various models: random forests, boosted decision trees (with AdaBoost), support vector machines, and artificial neural networks. The best models have accuracies of approximately greater than 97% (approximately less than 3% error), which is a significant improvement on a cut in GRB flux which has an accuracy of 89:6% (10:4% error). These models are then used to measure the detection efficiency of Swift as a function of redshift z, which is used to perform Bayesian parameter estimation on the GRB rate distribution. We find a local GRB rate density of eta(sub 0) approximately 0.48(+0.41/-0.23) Gpc(exp -3) yr(exp -1) with power-law indices of eta(sub 1) approximately 1.7(+0.6/-0.5) and eta(sub 2) approximately -5.9(+5.7/-0.1) for GRBs above and below a break point of z(sub 1) approximately 6.8(+2.8/-3.2). This methodology is able to improve upon earlier studies by more accurately modeling Swift detection and using this for fully Bayesian model fitting. The code used in this is analysis is publicly available online.

gamma rays: general↗

FY21 Progress Report: SRNL Analysis of ICCWR LCM and WAMS data for Corrosion and Cracking

The development of algorithms for machine learning and data analysis for the 3013 Surveillance Program is a collaborative effort by the Savannah River National Laboratory (SRNL) and the University of South Carolina (USC). For corrosion detection, Laser Confocal Microscope (LCM) or Wide Area 3D Measurement System (WAMS) data is extracted from large binary files, with software written to convert the data to physical attributes (e.g., height, color and grayscale values; all as functions of a location in a plane projection). A user-friendly Matlab Graphical User Interface (GUI) that reads data from either LCM or WAMS files was developed to integrate input data with software developed for processing and evaluation. The GUI can selectively download binary data, interrogate data attributes, label data, flag significant features, execute Machine Learning (ML) algorithms, output parameters for trained ML algorithms, report ML model accuracy with respect to labeled data, and generate graphical representations for various analyses. Features can be called out by user-specified thresholds, manual labeling or machine learning algorithms when they have been completed. The ability to rapidly label data is important because of the volume of data required for training machine learning algorithms. The GUI has the flexibility to allow addition of improved ML algorithms, methods for data visualization, and statistical computations. Statistical analyses via the GUI include areas of pits within a defined range of pit depths, correlations between Red-Green-Blue (RGB) or grayscale intensity and relative surface height, covariances between values associated with features, and feature histograms. The development of supervised machine learning algorithms, however, has been hindered by a lack of training data. The machine learning algorithms for crack identification are being refined but require improvements to the true positive rate for crack detection. This shortcoming is an artifact of the limited training data currently available, perhaps more so than the structure of the neural networks. At present, the best results are had from a consensus over an ensemble of randomly generated Deep Neural Network (DNN) or Convolutional Neural Network (CNN) algorithms. Although the consensus accuracy method has yielded optimum true positive and true negative rates in excess of 80%, additional validation testing is necessary. In addition to the suite of LCM data that was initially used, and which represents the majority of the work presented in this report, WAMS image data was also reviewed at a preliminary level. The review included a comparison between image resolution and dynamic range for each method. WAMS (ZON file) image data was found to have a pixel pitch of 3.69μm compared to 1 μm for the LCM (vk4 file) data, which implies a lower resolution for the WAMS images. Conversely, the ratio of dynamic range of the WAMS data to the LCM data was approximately 41:20 for height data, suggesting that information from WAMS should more accurately determine the depth of pits. At present, the significance of the greater dynamic range of the WAMS data relative to the LCM data has not yet been evaluated.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Findings on Subtask 3.1 - Bakken Rich Gas Enhanced Oil Recovery Project

Total in-place oil for the Bakken petroleum system (BPS) (which includes the Bakken and Three Forks Formations) has been estimated to be 600 billion barrels (bbl). However, BPS wells have decline rates as high as 85% over the first 3 years of their lives, and primary recovery factors typically range from 3% to 10% of original oil in place. Given the low initial recovery rates, even small incremental productivity improvements could dramatically increase technically recoverable oil in the BPS. One potential solution is enhanced oil recovery (EOR) using gas injection, such as carbon dioxide (CO2) or hydrocarbon (HC) gases. While commonly used in conventional reservoirs, CO2 EOR in unconventional tight oil reservoirs has been limited to pilot tests. EOR using rich gas (mixture of methane, ethane, and propane) has also been employed in numerous pilots in several unconventional plays and has recently been successfully applied in the Eagle Ford play. If successful, large-scale gas-based EOR in the BPS could dramatically increase oil productivity and recovery factors and extend the life of the play for decades. While CO2 may be a technically suitable working fluid for EOR in the BPS, supplies are limited and costs for using CO2 in EOR pilots are prohibitively high. Meanwhile, produced gas flaring has presented challenges for BPS operators in North Dakota. Analysis conducted by the North Dakota Pipeline Authority indicates that the current gas-gathering infrastructure in North Dakota is insufficient to accommodate all of the associated gas that is produced from the BPS. The geographically isolated location of North Dakota relative to large natural gas markets, combined with suppressed natural gas prices, has made it economically challenging for industry to invest capital in expanding gas-gathering infrastructure in the state. These circumstances led to a research program conducted by the Energy & Environmental Research Center (EERC) in partnership with Liberty Resources Management Company LLC (LR) to examine the potential to use rich gas injection for EOR and mitigate flaring. A rich gas EOR pilot test was designed and executed by LR at its Stomping Horse development area in Williams County, North Dakota. From July 2018 through May 2019, a total of 160 million standard cubic feet (MMscf) of rich produced gas was injected into the BPS using five different wells in a sequential injection strategy. LR’s Leon–Gohrick drill spacing unit (DSU) was used as the test site. Regulatory oversight was provided by the North Dakota Industrial Commission (NDIC). Technical support was provided by the EERC through a series of laboratory, modeling, and field-based activities, and additional post-pilot research activities incorporated learnings from the test, developed new laboratory data, improved fracture modeling methods, and developed machine learning and big data analytics. The results from the Stomping Horse rich gas EOR pilot activities indicate that developing an effective, economical EOR approach for the BPS will require more field tests. Another key lesson learned from the Stomping Horse tests is that detailed pre- and posttest data on reservoir conditions and fluids production are essential. Robust reservoir characterization provides information that is crucial to creating realistic geomodels and conducting valid dynamic simulations of potential EOR scenarios. A detailed understanding of the completions and production history of offset wells is also necessary for valid test result interpretations. This knowledge is essential to designing the operational parameters of injectivity tests and interpreting the results. A conformance control strategy is also essential to success. Laboratory-based examinations of rich gas interactions with reservoir fluids and rocks were conducted, with an emphasis on determining the ability to mobilize oil in the tight reservoir rocks and shales of the BPS. Injection fluid composition was shown to have a positive impact on reducing reservoir oil minimum miscibility pressure (MMP), reducing interfacial tension (IFT), and altering wettability. IFT and contact angle measurements demonstrated that wettability can be altered in the presence of rich gas, suggesting the potential to improve oil recovery. Iterative modeling of surface infrastructure and reservoir performance using data generated by the various project activities was conducted. A geologic model of the Stomping Horse area was built; history-matched oil, gas, and water production was used in simulations of various EOR scenarios. Early programmatic modeling results were used to support LR’s design and operation of the EOR pilot and to provide insight regarding optimization of future commercial-scale BPS EOR design and operations. Post-pilot modeling focused on alternative methods of understanding complex fracture networks and accelerating simulation time. These led to improved simulation run times and provide excellent history-matching results. Several of these iterative models were used as the bases for developing algorithms into machine learning and big data analytics. History matching in reservoir simulation is time-consuming and computer processing-intensive. Machine learning algorithms were created, and an automated history-matching tool was developed. A large set of synthetic reservoir simulations were created to generate well responses (oil, gas, and water production, well bottomhole pressure [BHP], and tracer or propane breakthrough) for a set of EOR operating parameters that included offset well status (open or closed), injectate (rich gas or propane), injection rate, and injection well BHP. A user interface was developed to provide real-time visualization. Machine learning-based models were developed to provide rapid forecasting of well performance given a set of user-defined EOR operating parameters. These predictive models allow the user to modify the offset well status, injection rate, and injection well BHP and rapidly forecast future production performance. The combination of real-time visualization tools with real-time forecasting tools provides a framework for real-time control—operational changes that the EOR site operator can enact (e.g., changing gas injection rates) to affect the observed performance and potentially improve the EOR outcome. There is great reason to be optimistic about the future of EOR in the Bakken. The results of the laboratory studies suggest significant potential for high rates of oil mobilization using produced field gas injection under the right conditions. The results of the lab studies, combined with rigorous statistical analysis of well production data and associated modeling efforts, confirm the notion that fluid mobility within the reservoir is controlled by fractures. As more knowledge is gained about the nature and distribution of fracture networks in the Bakken, the industry will be in a better position to predict and, ultimately, influence fluid mobility. New field tests are necessary to develop a more complete understanding of those conditions. Thoughtful and creatively engineered field tests within a well-characterized geologic setting will yield the fundamental knowledge needed to take Bakken oil production to the next level. This subtask was cofunded through the EERC–U.S. Department of Energy Joint Program on Research and Development for Fossil Energy-Related Resources Cooperative Agreement No. DE-FE0024233. Nonfederal funding was provided by the North Dakota Industrial Commission’s Oil and Gas Research Program and Computer Modelling Group.

04 OIL SHALES AND TAR SANDS↗

Empirical Analysis and Automated Classification of Security Bug Reports

With the ever expanding amount of sensitive data being placed into computer systems, the need for effective cybersecurity is of utmost importance. However, there is a shortage of detailed empirical studies of security vulnerabilities from which cybersecurity metrics and best practices could be determined. This thesis has two main research goals: (1) to explore the distribution and characteristics of security vulnerabilities based on the information provided in bug tracking systems and (2) to develop data analytics approaches for automatic classification of bug reports as security or non-security related. This work is based on using three NASA datasets as case studies. The empirical analysis showed that the majority of software vulnerabilities belong only to a small number of types. Addressing these types of vulnerabilities will consequently lead to cost efficient improvement of software security. Since this analysis requires labeling of each bug report in the bug tracking system, we explored using machine learning to automate the classification of each bug report as a security or non-security related (two-class classification), as well as each security related bug report as specific security type (multiclass classification). In addition to using supervised machine learning algorithms, a novel unsupervised machine learning approach is proposed. An ac- curacy of 92%, recall of 96%, precision of 92%, probability of false alarm of 4%, F-Score of 81% and G-Score of 90% were the best results achieved during two-class classification. Furthermore, an accuracy of 80%, recall of 80%, precision of 94%, and F-score of 85% were the best results achieved during multiclass classification.

Cybersecurity↗

Evaluation of global terrestrial evapotranspiration using state-of-the-art approaches in remote sensing, machine learning and land surface modeling

Evapotranspiration (ET) is critical in linking global water, carbon and energy cycles. However, direct measurement of global terrestrial ET is not feasible. Here, we first reviewed the basic theory and state-of-the-art approaches for estimating global terrestrial ET, including remote-sensing-based physical models, machine-learning algorithms and land surface models (LSMs). We then utilized 4 remote-sensing-based physical models, 2 machine-learning algorithms and 14 LSMs to analyze the spatial and temporal variations in global terrestrial ET. The results showed that the ensemble means of annual global terrestrial ET estimated by these three categories of approaches agreed well, with values ranging from 589.6 mm/yr (6.56×10^4 cu.km/yr) to 617.1 mm/yr (6.87×10^4 cu.km/yr). For the period from 1982 to 2011, both the ensembles of remote-sensing-based physical models and machine-learning algorithms suggested increasing trends in global terrestrial ET (0.62 mm/sq.yr with a significance level of p<0.05 and 0.38 mm yr−2 with a significance level of p<0.05, respectively). In contrast, the ensemble mean of the LSMs showed no statistically significant change (0.23 mm/sq.yr, p>0.05), although many of the individual LSMs reproduced an increasing trend. Nevertheless, all 20 models used in this study showed that anthropogenic Earth greening had a positive role in increasing terrestrial ET. The concurrent small interannual variability, i.e., relative stability, found in all estimates of global terrestrial ET, suggests that a potential planetary boundary exists in regulating global terrestrial ET, with the value of this boundary being around 600 mm/yr. Uncertainties among approaches were identified in specific regions, particularly in the Amazon Basin and arid/semiarid regions. Improvements in parameterizing water stress and canopy dynamics, the utilization of new available satellite retrievals and deep-learning methods, and model–data fusion will advance our predictive understanding of global terrestrial ET.

surface modeling↗