Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “machine learning algorithms”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Using Machine Learning Algorithm to Detect Blowing Snow and Fog in Antarctica Based on Ceilometer and Surface Meteorology Systems

Blowing snow is a common weather phenomenon in Antarctica and plays an important role in the water vapor cycle and ice sheet mass balance. Although it has a significant impact on the climate of Antarctica, people do not know much about this process. Fog events are difficult to distinguish from blowing snow events using existing detection algorithms by a ceilometer. In this study, based on ceilometer, the meteorological parameters observed by surface meteorology systems are further combined to detect blowing snow and fog using the AdaBoost algorithm. The weather phenomena recorded by human observers are ‘true’. The dataset is collected from 1 January 2016 to 31 December 2016 at the AWARE site. Among them, three-quarters of the data are used as the training set and the rest of the data as the testing set. The classification accuracy of the proposed algorithm for the testing set is about 94%. Compared with the Loeb method, the proposed algorithm can detect 89.12% of blowing snow events and 76.10% of fog events, while the Loeb method can only identify 64.29% of blowing snow events and 31.87% of fog events.

54 ENVIRONMENTAL SCIENCES↗

Review of machine learning and deep learning models for toxicity prediction

The ever-increasing number of chemicals has raised public concerns due to their adverse effects on human health and the environment. To protect public health and the environment, it is critical to assess the toxicity of these chemicals. Traditional in vitro and in vivo toxicity assays are complicated, costly, and time-consuming and may face ethical issues. These constraints raise the need for alternative methods for assessing the toxicity of chemicals. Recently, due to the advancement of machine learning algorithms and the increase in computational power, many toxicity prediction models have been developed using various machine learning and deep learning algorithms such as support vector machine, random forest, k-nearest neighbors, ensemble learning, and deep neural network. This review summarizes the machine learning- and deep learning-based toxicity prediction models developed in recent years. Support vector machine and random forest are the most popular machine learning algorithms, and hepatotoxicity, cardiotoxicity, and carcinogenicity are the frequently modeled toxicity endpoints in predictive toxicology. It is known that datasets impact model performance. The quality of datasets used in the development of toxicity prediction models using machine learning and deep learning is vital to the performance of the developed models. The different toxicity assignments for the same chemicals among different datasets of the same type of toxicity have been observed, indicating benchmarking datasets is needed for developing reliable toxicity prediction models using machine learning and deep learning algorithms. This review provides insights into current machine learning models in predictive toxicology, which are expected to promote the development and application of toxicity prediction models in the future.

Research & Experimental Medicine↗

Materials Learning Algorithms (MALA): Scalable machine learning for electronic structure calculations in large-scale atomistic simulations

We present the Materials Learning Algorithms (MALA) package, a scalable machine learning framework designed to accelerate density functional theory (DFT) calculations suitable for large-scale atomistic simulations. Using local descriptors of the atomic environment, MALA models efficiently predict key electronic observables, including local density of states, electronic density, density of states, and total energy. The package integrates data sampling, model training and scalable inference into a unified library, while ensuring compatibility with standard DFT and molecular dynamics codes. We demonstrate MALA's capabilities with examples including boron clusters, aluminum across its solid-liquid phase boundary, and predicting the electronic structure of a stacking fault in a large beryllium slab. Scaling analyses reveal MALA's computational efficiency and identify bottlenecks for future optimization. With its ability to model electronic structures at scales far beyond standard DFT, MALA is well suited for modeling complex material systems, making it a versatile tool for advanced materials research.

Density functional theory↗

GOOML (Geothermal Operational Optimization with Machine Learning) [SWR-23-01]

The Geothermal Operational Optimization with Machine Learning (GOOML) is a partnership between NREL and Upflow, NZ, awarded in response to the U.S. Department of Energy's Geothermal Technologies Office's Funding Opportunity Announcement (FOA) to expand the role of advanced analytics and automation in geothermal operations through machine learning. Partnering with industry (Contact Energy Limited ("Contact"), Ngati Tuwharetoa Geothermal Assets Limited ("NTGA"), Ormat Technologies Inc. ("Ormat") and Flow State Solutions Limited ("FSS"), GOOML was created to improve the operational efficiency of geothermal power plant steam fields through the analysis of historical operational data and the application of custom machine learning algorithms. NREL's contributions include machine learning, coding, and data management expertise as well as access to high-performance compute solutions. GOOML can increase geothermal operational efficiency through development of a digital system twin that can be utilized to provide optimal geothermal operating conditions for real-world geothermal fields. GOOML allows users to analyze field production histories in detail, develop models, and train machine learning algorithms to identify opportunities for increased geothermal efficiency, detect potential trouble, and allow predictive scenario modeling. Preliminary experiments have demonstrated a potential to increase total generation by as much as 12% through ML optimization of the utilization of existing steam field resources.

Buster, Grant↗

Improving the Freight Productivity of a Heavy-Duty, Battery Electric Truck by Intelligent Energy Management

This project aimed to enhance the range and reduce the operating costs of battery electric Class 8 trucks traveling over 250 miles daily. This was achieved through the development and implementation of an intelligent-Energy Management System (i-EMS) that leverages vehicle and operations data, physics-aware machine learning algorithms, and vehicle-to-cloud (V2C) connectivity. The project hypothesized that advanced machine learning algorithms and real-time data analytics could significantly improve the energy efficiency and range of these trucks. Key objectives included developing a physics-aware machine learning algorithm, implementing an i-EMS with V2C connectivity and physics-aware spatial data analytics (PSDA), and validating the system’s effectiveness with fleet partners HEB Companies and Murphy Logistics. Extensive data collection from vehicle operations, including vehicle characteristics, road conditions, and payload, was conducted. A machine learning algorithm was developed to predict energy consumption and enable proactive decision-making. The i-EMS was implemented on two Volvo VNR BEVs, with operators receiving charging and routing recommendations. Charging stations were installed at depot locations in Texas and Minnesota, with an additional on-route charger in Minnesota. Significant findings included a 14% range improvement for Murphy Logistics on a highway-driving eco-route and a 22% range improvement for HEB Companies on a city-driving eco-route. The i-EMS utilized rule-based methods and physics-based algorithms to predict and reduce energy consumption, with real-time monitoring and analysis through V2C connectivity enabling proactive decision-making. The project demonstrated the feasibility and economic viability of battery electric Class 8 trucks for long-haul operations, showcasing the potential of physics-aware machine learning in optimizing energy management. The successful implementation of the i-EMS in real-world scenarios validates its practical application and effectiveness, paving the way for the widespread adoption of battery electric vehicles in the freight transportation industry.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Artificial Diversity and Defense Security (ADDSec)

Artificial Diversity and Defense Security (ADDSec) machine learning algorithms are used to classify and cluster threats so that an appropriate response can be initiated as a mitigation strategy. The package includes an ensemble of machine learning algorithms such as Support Vector Machines, naïve bayes, logistic regression, and random forest that evolve with the data to recognize anomalous behavior at the host and network levels. Inputs into the machine learning algorithms include end host system calls, system utilization, packet captures, and syslog messages. The machine learning algorithms can be retrained based on user defined intervals or on the number of packets received. ADDSEC's threat responses include Internet Protocol (IP) Address randomization, application port number randomization, and application library randomization. The IP randomization implementation is built on top of a Software Defined Networking (SDN) framework. The SDN controller installs flows on each of the SDN switches with randomized source and destination IP addresses. The application port numbers are randomized using iptables. The application library randomization is created with a LLVM compiler. All randomization schemes are transparent to the endpoints on the network. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525. SAND2021-3379 O

Cox, RebeccaE.↗

FY21 Progress Report: SRNL Analysis of ICCWR LCM and WAMS data for Corrosion and Cracking

The development of algorithms for machine learning and data analysis for the 3013 Surveillance Program is a collaborative effort by the Savannah River National Laboratory (SRNL) and the University of South Carolina (USC). For corrosion detection, Laser Confocal Microscope (LCM) or Wide Area 3D Measurement System (WAMS) data is extracted from large binary files, with software written to convert the data to physical attributes (e.g., height, color and grayscale values; all as functions of a location in a plane projection). A user-friendly Matlab Graphical User Interface (GUI) that reads data from either LCM or WAMS files was developed to integrate input data with software developed for processing and evaluation. The GUI can selectively download binary data, interrogate data attributes, label data, flag significant features, execute Machine Learning (ML) algorithms, output parameters for trained ML algorithms, report ML model accuracy with respect to labeled data, and generate graphical representations for various analyses. Features can be called out by user-specified thresholds, manual labeling or machine learning algorithms when they have been completed. The ability to rapidly label data is important because of the volume of data required for training machine learning algorithms. The GUI has the flexibility to allow addition of improved ML algorithms, methods for data visualization, and statistical computations. Statistical analyses via the GUI include areas of pits within a defined range of pit depths, correlations between Red-Green-Blue (RGB) or grayscale intensity and relative surface height, covariances between values associated with features, and feature histograms. The development of supervised machine learning algorithms, however, has been hindered by a lack of training data. The machine learning algorithms for crack identification are being refined but require improvements to the true positive rate for crack detection. This shortcoming is an artifact of the limited training data currently available, perhaps more so than the structure of the neural networks. At present, the best results are had from a consensus over an ensemble of randomly generated Deep Neural Network (DNN) or Convolutional Neural Network (CNN) algorithms. Although the consensus accuracy method has yielded optimum true positive and true negative rates in excess of 80%, additional validation testing is necessary. In addition to the suite of LCM data that was initially used, and which represents the majority of the work presented in this report, WAMS image data was also reviewed at a preliminary level. The review included a comparison between image resolution and dynamic range for each method. WAMS (ZON file) image data was found to have a pixel pitch of 3.69μm compared to 1 μm for the LCM (vk4 file) data, which implies a lower resolution for the WAMS images. Conversely, the ratio of dynamic range of the WAMS data to the LCM data was approximately 41:20 for height data, suggesting that information from WAMS should more accurately determine the depth of pits. At present, the significance of the greater dynamic range of the WAMS data relative to the LCM data has not yet been evaluated.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

NgramPPM: Compression Analytics without Compression

Arithmetic Coding (AC) using Prediction by Partial Matching (PPM) is a compression algorithm that can be used as a machine learning algorithm. This paper describes a new algorithm, NGram PPM. NGram PPM has all the predictive power of AC/PPM, but at a fraction of the computational cost. Unlike compression-based analytics, it is also amenable to a vector space interpretation, which creates the ability for integration with other traditional machine learning algorithms. AC/PPM is reviewed, including its application to machine learning. Then NGram PPM is described and test results are presented, comparing them to AC/PPM.

97 MATHEMATICS AND COMPUTING↗

Findings on Subtask 3.1 - Bakken Rich Gas Enhanced Oil Recovery Project

Total in-place oil for the Bakken petroleum system (BPS) (which includes the Bakken and Three Forks Formations) has been estimated to be 600 billion barrels (bbl). However, BPS wells have decline rates as high as 85% over the first 3 years of their lives, and primary recovery factors typically range from 3% to 10% of original oil in place. Given the low initial recovery rates, even small incremental productivity improvements could dramatically increase technically recoverable oil in the BPS. One potential solution is enhanced oil recovery (EOR) using gas injection, such as carbon dioxide (CO2) or hydrocarbon (HC) gases. While commonly used in conventional reservoirs, CO2 EOR in unconventional tight oil reservoirs has been limited to pilot tests. EOR using rich gas (mixture of methane, ethane, and propane) has also been employed in numerous pilots in several unconventional plays and has recently been successfully applied in the Eagle Ford play. If successful, large-scale gas-based EOR in the BPS could dramatically increase oil productivity and recovery factors and extend the life of the play for decades. While CO2 may be a technically suitable working fluid for EOR in the BPS, supplies are limited and costs for using CO2 in EOR pilots are prohibitively high. Meanwhile, produced gas flaring has presented challenges for BPS operators in North Dakota. Analysis conducted by the North Dakota Pipeline Authority indicates that the current gas-gathering infrastructure in North Dakota is insufficient to accommodate all of the associated gas that is produced from the BPS. The geographically isolated location of North Dakota relative to large natural gas markets, combined with suppressed natural gas prices, has made it economically challenging for industry to invest capital in expanding gas-gathering infrastructure in the state. These circumstances led to a research program conducted by the Energy & Environmental Research Center (EERC) in partnership with Liberty Resources Management Company LLC (LR) to examine the potential to use rich gas injection for EOR and mitigate flaring. A rich gas EOR pilot test was designed and executed by LR at its Stomping Horse development area in Williams County, North Dakota. From July 2018 through May 2019, a total of 160 million standard cubic feet (MMscf) of rich produced gas was injected into the BPS using five different wells in a sequential injection strategy. LR’s Leon–Gohrick drill spacing unit (DSU) was used as the test site. Regulatory oversight was provided by the North Dakota Industrial Commission (NDIC). Technical support was provided by the EERC through a series of laboratory, modeling, and field-based activities, and additional post-pilot research activities incorporated learnings from the test, developed new laboratory data, improved fracture modeling methods, and developed machine learning and big data analytics. The results from the Stomping Horse rich gas EOR pilot activities indicate that developing an effective, economical EOR approach for the BPS will require more field tests. Another key lesson learned from the Stomping Horse tests is that detailed pre- and posttest data on reservoir conditions and fluids production are essential. Robust reservoir characterization provides information that is crucial to creating realistic geomodels and conducting valid dynamic simulations of potential EOR scenarios. A detailed understanding of the completions and production history of offset wells is also necessary for valid test result interpretations. This knowledge is essential to designing the operational parameters of injectivity tests and interpreting the results. A conformance control strategy is also essential to success. Laboratory-based examinations of rich gas interactions with reservoir fluids and rocks were conducted, with an emphasis on determining the ability to mobilize oil in the tight reservoir rocks and shales of the BPS. Injection fluid composition was shown to have a positive impact on reducing reservoir oil minimum miscibility pressure (MMP), reducing interfacial tension (IFT), and altering wettability. IFT and contact angle measurements demonstrated that wettability can be altered in the presence of rich gas, suggesting the potential to improve oil recovery. Iterative modeling of surface infrastructure and reservoir performance using data generated by the various project activities was conducted. A geologic model of the Stomping Horse area was built; history-matched oil, gas, and water production was used in simulations of various EOR scenarios. Early programmatic modeling results were used to support LR’s design and operation of the EOR pilot and to provide insight regarding optimization of future commercial-scale BPS EOR design and operations. Post-pilot modeling focused on alternative methods of understanding complex fracture networks and accelerating simulation time. These led to improved simulation run times and provide excellent history-matching results. Several of these iterative models were used as the bases for developing algorithms into machine learning and big data analytics. History matching in reservoir simulation is time-consuming and computer processing-intensive. Machine learning algorithms were created, and an automated history-matching tool was developed. A large set of synthetic reservoir simulations were created to generate well responses (oil, gas, and water production, well bottomhole pressure [BHP], and tracer or propane breakthrough) for a set of EOR operating parameters that included offset well status (open or closed), injectate (rich gas or propane), injection rate, and injection well BHP. A user interface was developed to provide real-time visualization. Machine learning-based models were developed to provide rapid forecasting of well performance given a set of user-defined EOR operating parameters. These predictive models allow the user to modify the offset well status, injection rate, and injection well BHP and rapidly forecast future production performance. The combination of real-time visualization tools with real-time forecasting tools provides a framework for real-time control—operational changes that the EOR site operator can enact (e.g., changing gas injection rates) to affect the observed performance and potentially improve the EOR outcome. There is great reason to be optimistic about the future of EOR in the Bakken. The results of the laboratory studies suggest significant potential for high rates of oil mobilization using produced field gas injection under the right conditions. The results of the lab studies, combined with rigorous statistical analysis of well production data and associated modeling efforts, confirm the notion that fluid mobility within the reservoir is controlled by fractures. As more knowledge is gained about the nature and distribution of fracture networks in the Bakken, the industry will be in a better position to predict and, ultimately, influence fluid mobility. New field tests are necessary to develop a more complete understanding of those conditions. Thoughtful and creatively engineered field tests within a well-characterized geologic setting will yield the fundamental knowledge needed to take Bakken oil production to the next level. This subtask was cofunded through the EERC–U.S. Department of Energy Joint Program on Research and Development for Fossil Energy-Related Resources Cooperative Agreement No. DE-FE0024233. Nonfederal funding was provided by the North Dakota Industrial Commission’s Oil and Gas Research Program and Computer Modelling Group.

04 OIL SHALES AND TAR SANDS↗

Identifying Vehicle Signals in Continuous Seismic Data Using Unsupervised Machine-Learning Techniques

Seismic sensors deployed near roadways effectively capture ground vibrations generated by passing vehicles. Although both traditional and machine‐learning algorithms have been utilized for analyzing such signals, independent validation of detected vehicle events remains limited. We applied two unsupervised machine‐learning algorithms, uniform manifold approximation and projection for dimension reduction, and hierarchical density‐based spatial clustering of applications with noise, to continuous seismic data collected along a road on the main campus of Oak Ridge National Laboratory. The algorithms identified seven distinct cluster labels across the entire dataset. By comparing these cluster labels with precipitation records from a nearby weather station and image‐derived labels from a local camera system, we identified one cluster associated with rainfall and another with vehicle activity. Our algorithms identified a greater number of vehicle‐related labels compared to the camera‐derived labels because seismic data are unaffected by poor lighting conditions. The arrival times of the newly detected vehicle signals corresponded well with the road’s speed limit, supporting our findings. Our algorithm outperformed the short‐term average/long‐term average method and k‐means clustering. Our results suggest that seismic data, when analyzed with machine‐learning algorithms, can complement existing vehicle monitoring systems, particularly under challenging environmental conditions.

Chai, Chengping [Oak Ridge National Laboratory (OR↗

Differentiation and classification of bacterial endotoxins based on surface enhanced Raman scattering and advanced machine learning

Bacterial endotoxin, a major component of the Gram-negative bacterial outer membrane leaflet, is a lipopolysaccharide shed from bacteria during their growth and infection and can be utilized as a biomarker for bacterial detection. Here, the surface enhanced Raman scattering (SERS) spectra of eleven bacterial endotoxins with an average detection amount of 8.75 pg per measurement have been obtained based on silver nanorod array substrates, and the characteristic SERS peaks have been identified. With appropriate spectral pre-processing procedures, different classical machine learning algorithms, including support vector machine, k-nearest neighbor, random forest, etc., and a modified deep learning algorithm, RamanNet, have been applied to differentiate and classify these endotoxins. It has been found that most conventional machine learning algorithms can attain a differentiation accuracy of >99%, while RamanNet can achieve 100% accuracy. Such an approach has the potential for precise classification of endotoxins and could be used for rapid medical diagnoses and therapeutic decisions for pathogenic infections.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Romans v.0.2.0

SAND2022-1748 O Romans is a library for implementing compression-based machine learning algorithms and algorithms for Human Constrained Machine Learning. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Bauer, Travis↗

Ensemble Federated Machine Learning‐Based Cybersecurity Situational Awareness in Microgrid Network

Cyber-physical microgrids are vulnerable to stealthy cybersecurity threats that disguise their actions through the exploitation of system knowledge. Such actions can severely impacts microgrids deployed in defense bases, slowing the response time of military forces during national emergencies. Several machine-learning algorithms have been proposed to detect intrusions in the grid networks; however, these traditional machine-learning algorithms lack data privacy and are subject to several adversarial machine-learning threats. This paper proposes a novel federated machine learning (FML)-based three-model framework to detect and identify stealthy data-integrity attacks while ensuring data privacy in microgrid networks. The proposed architecture uses a variational mode decomposition technique to extract derived features from incoming measurement and control datasets. The extraction of these derived features allows FML models to learn minute variations in data patterns that allow them to perform significantly better than the models trained with generic datasets consisting of raw features. Our experimental results show the efficient performance of the proposed methodology against different types of data integrity attacks while considering primary and secondary controllers in microgrids. Further, the applied FML-integrated random forest ensemble algorithm outperforms the existing generic FML algorithms during noisy and noise-free datasets with prediction latencies of only 91–134 µs per sample within the 0.1 s sampling interval and requires communication bandwidth of around ∼8.25 KB/s at the control center and ∼2.7 KB/s per edge client for communication.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Polarized and unpolarized gluon PDFs: Generative machine learning applications for lattice QCD matrix elements at short distance and large momentum

Lattice quantum chromodynamics (QCD) calculations share a defining challenge by requiring a small finite range of spatial separation z between quark/gluon bilinears for controllable power corrections in the perturbative QCD factorization, and a large hadron boost p z for a successful determination of collinear parton distribution functions (PDFs). However, these two requirements make the determination of PDFs from lattice data very challenging. We present the application of generative machine learning algorithms to estimate the polarized and unpolarized gluon correlation functions utilizing short-distance data and extending the correlation up to z p z ≲ 14 , surpassing the current capabilities of lattice QCD calculations. We train physics-informed machine learning algorithms to learn from the short-distance correlation at z ≲ 0.36 fm and take the limit, p z → ∞ , thereby minimizing possible contamination from the higher-twist effects for a successful reconstruction of the polarized gluon PDF. We also expose the bias and problems with underestimating uncertainties associated with the use of model-dependent and overly constrained functional forms, such as x α ( 1 − x ) β and its variants to extract PDFs from the lattice data. We propose the use of generative machine learning algorithms to mitigate these issues and present our determination of the polarized and unpolarized gluon PDFs in the nucleon. Published by the American Physical Society 2025

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Hls4ml Synthesis Testing

HLS4ml (high level synthesis for machine learning) Is a Python package used to translate commonly used open-source machine learning models into HLS. This is useful in machine learning applications on FPGAs. Machine learning algorithms are only as fast as the hardware that they are used on, and some applications require high speed without sacrificing accuracy. In these situations, an FPGA is a good choice since it is faster than a CPU or a GPU, but programming an FPGA is difficult. This is where HLS4ml can be used to simplify the process, as a well-known learning model can be converted to HLS and more easily deployed onto an FPGA. There are many use cases for a machine learning algorithm running on an FPGA. For example, detectors in a particle accelerator cannot keep every event that they detect, and so a computer must decide which events to keep and which to discard. Using an FPGA with a machine learning algorithm would be a good way to keep as many events as possible.

Swanson, Caiden↗