Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “machine learning algorithms”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24

Master State Threat Identifier (masti)

The MSE cyber sensor offers a novel approach for scalable network monitoring using heterogeneous computing. Packet data is collected, analyzed, and stored by a stand alone cyber sensor. Machine learning algorithms are run to discover patterns in the communication without the need for deep packet inspection. Tuning of the detection algorithms happens automatically, without need for an engineer to manually create detection rules.

Russell, PierceL↗

Environmental Molecular Network (ENVnet) v1

Here, we present an approach that integrates mass difference based deconvolution with molecular networking to build a static reference network from all publicly available organic matter metabolomics datasets. This is accomplished using MS/MS deconvolution coupled with both recently reported (BUDDY) and novel machine learning algorithms to determine chemical formulas and perform MS/MS alignments (REM-BLINK).

Bowen, Benjamin↗

PyJMAK: An Open-Source Python Toolkit for Modeling Solid-State Metallurgical Phase Transformations

Accurate prediction of metallurgical phase transformations is an essential basis for autonomous optimization and rapid part qualification. Several methods can be used to estimate the evolution of phase fractions such as JMAK kinetics-based models, phase-field models, thermodynamic models, and data-driven machine learning models. Thermodynamic and phase-field-based methodologies solve multiphysics equations requiring numerous calibration parameters and significant computational resources. As a result, the computation domain is limited to a point or on order of micron-meters. The data-driven models rely on large datasets from experiments and simulations. While the JMAK model only provides information about phase fraction evolution, it can predict this evolution in near real-time using thermal history and thermodynamic data without restriction on the domain. JMAK models have been popularly used by researchers to model phase transformations occuring during additive manufacturing or over arbitrary temperature profiles. Commercial proprietary software such as Abaqus and Ansys or closed-source in-house implementations offer the ability to model JMAK based kinetics to predict phase transformation. However, these software packages are not open-source or freely available for use and development in conjunction with manufacturing machines, sensors, and machine learning algorithms. In addition, the use of the model is restricted by a license token. In contrast, given temperature profiles at multiple points in the domain, this Python-based PyJMAK model can compute phase evolution in parallel due to its stand-alone modular, voxel-based structure, and it can be executed on high-performance computing resources without any license restrictions.

Prabhune, Bhagya [Oak Ridge National Laboratory (O↗

East Asian Rainbands and Associated Circulation over the Tibetan Plateau Region

Abstract Rainbands that migrate northward from spring to summer are persistent features of the East Asian summer monsoon. This study employs a machine learning algorithm to identify individual East Asian rainbands from May to August in the 6-hourly ERA-Interim reanalysis product and captures rainband events during these months for the period 1979–2018. The median duration of rainband events at any location in East Asia is 12 h, and the centroids of these rainbands move northward continuously from approximately 28°N in late May to approximately 33°N in July, instead of making jumps between quasi-stationary periods. Whereas the length and overall area of the rainbands grow monotonically from May to June, the intensity of the rainfall within the rainband dips slightly in early June before it peaks in late June. We find that extratropical northerly winds on all pressure levels over East China are the most important anomalous flow accompanying the rainband events. The anomalous northerlies augment climatological background northerlies in bringing low moist static energy air and thus generate the front associated with the rainband. Persistent lower-tropospheric southerly winds bring in moisture that feeds the rainband and are enhanced a few days prior to rainband events, but they are not directly tied to the actual rainband formation. The background northerlies could originate as part of the Rossby waves resulting from the jet stream interaction with the Tibetan Plateau. The ageostrophic circulation in the jet entrance region peaks in May and weakens in June and July and does not prove to be critical to the formation of the rainbands.

54 ENVIRONMENTAL SCIENCES↗

The Role of Internal Variability in Springtime Arctic Amplification from 1980 to 2022

Arctic amplification (AA) refers to the enhanced warming of the Arctic relative to the global average due to rising greenhouse gases, measured as the ratio of Arctic-mean to global-mean surface air temperature (SAT) trends. From 1980 to 2022, annual-mean AA reached 4.2 (Arctic defined as north of 70°N). Climate models simulate AA but fail to reproduce its magnitude. Sweeney et al. attributed much of this model–observation discrepancy to internal variability. AA shows seasonality and so does the discrepancy. Spring (March–May) shows the largest gap: Observed AA is 4.2, while the multimodel mean is 2.7. This raises several questions: 1) What role does internal variability play in observed spring AA? 2) How does simulated spring AA compare to observations when internal variability is removed? 3) If internal variability is significant, what mechanisms drive it? To address these, we adapted the machine learning algorithm from Sweeney et al., training on simulated multidecadal spring SAT and sea level pressure (SLP) trend maps. Our results show that internal variability enhanced spring Arctic warming by 37% and reduced global warming by 10%. Removing internal variability reconciles the spring AA discrepancy. The estimated internal contribution to Arctic spring warming is supported by an independent dynamical adjustment approach. We identify an atmospheric circulation pattern in observations associated with this internal warming. Observed internal Siberian SAT and SLP trends follow the simulated SAT–SLP relationship but lie at the distribution’s extreme, suggesting models generally underestimate internal variability unless the observed configuration reflects a rare real-world realization.

Arctic↗

Machine Learning–Adjusted WRF Forecasts to Support Wind Energy Needs in Black Start Operations

Abstract The push for increased capacity of renewable sources of electricity has led to the growth of wind-power generation, with a need for accurate forecasts of winds at hub height. Forecasts for these levels were uncommon until recently, and that, combined with the nocturnal collapse of the well-mixed boundary layer and daytime growth of the boundary layer through the levels important for energy generation, has contributed to errors in numerical modeling of wind generation resources. The present study explores several machine learning algorithms to both forecast and correct standard WRF Model forecasts of winds and temperature at hub height within wind turbine plants over several different time periods that are critical for the anticipation of potential blackouts and aiding in black start operations on the power grid. It was found that mean square error for day-2 wind forecasts from the WRF Model can be improved by over 90% with the use of a multioutput neural network, and that 60-min forecasts of WRF error, which can then be used to adjust forecasts, can be made with an LSTM with great accuracy. Nowcasting of temperature and wind speed over a 10-min period using an LSTM produced very low error and especially skillful forecasts of maximum and minimum values over the turbine plant area.

17 WIND ENERGY↗

Dynamic Ride-Matching for Large-Scale Transportation Systems

Efficient dynamic ride-matching (DRM) in large-scale transportation systems is a key driver in transport simulations to yield answers to challenging problems. Although the DRM problem is simple to solve, it quickly becomes a computationally challenging problem in large-scale transportation system simulations. Therefore, this study thoroughly examines the DRM problem dynamics and proposes an optimization-based solution framework to solve the problem efficiently. To benefit from parallel computing and reduce computational times, the problem’s network is divided into clusters utilizing a commonly used unsupervised machine learning algorithm along with a linear programming model. Then, these sub-problems are solved using another linear program to finalize the ride-matching. At the clustering level, the framework allows users adjusting cluster sizes to balance the trade-off between the computational time savings and the solution quality deviation. A case study in the Chicago Metropolitan Area, U.S., illustrates that the framework can reduce the average computational time by 58% at the cost of increasing the average pick up time by 26% compared with a system optimum, that is, non-clustered, approach. Another case study in a relatively small city, Bloomington, Illinois, U.S., shows that the framework provides quite similar results to the system-optimum approach in approximately 62% less computational time.

33 ADVANCED PROPULSION SYSTEMS↗

Estimation of Arrivals on Green at Signalized Intersections Using Stop-Bar Video Detection

Across the world, traffic congestion is increasing with alarming rapidity. Traffic signal control effectiveness, in coordinated networks, is often investigated in relation to the type of vehicle arrivals at the signalized intersections. Recently, several transportation agencies have switched from traditional loop detectors to video detection. When video cameras are accompanied by computer vision, one can extract more information about traffic “dynamics” than by using traditional inductive loop detectors. Collecting arrival times of multiple vehicles after the first arrival at the stop-bar detector might be challenging when using inductive loop detectors (since after the first arrival, detector status is always occupied). However, emerging video detection systems allow tracking of each vehicle’s entrance time in the detection zone, departure time from the detection zone, and the type of vehicle. This information can be used to estimate vehicular arrival and departure times, which then can be fed into machine learning algorithms to estimate arrivals on green (AOG). However, such research ideas have not been documented so far. Thus, this paper presents an estimation model for AOG, which was developed using multigene genetic programming. A robust experimental dataset was collected from a highly calibrated and validated microsimulation model of an 11-intersection corridor in Chattanooga, TN. The results of the model’s performance analysis showed the high accuracy of the training-, testing-, and validation datasets. The practical benefit of this model is that it can be applied to estimate arrival types at intersections where only stop-bar video detection exists.

Engineering↗

Optimal dimensionality selection for independent component analysis of transcriptomic data

Independent component analysis is an unsupervised machine learning algorithm that separates a set of mixed signals into a set of statistically independent source signals. Applied to high-quality gene expression datasets, independent component analysis effectively reveals both the source signals of the transcriptome as co-regulated gene sets, and the activity levels of the underlying regulators across diverse experimental conditions. Two major variables that affect the final gene sets are the diversity of the expression profiles contained in the underlying data, and the user-defined number of independent components, or dimensionality, to compute. Availability of high-quality transcriptomic datasets has grown exponentially as high-throughput technologies have advanced; however, optimal dimensionality selection remains an open question. We computed independent components across a range of dimensionalities for four gene expression datasets with varying dimensions (both in terms of number of genes and number of samples). We computed the correlation between independent components across different dimensionalities to understand how the overall structure evolves as the number of user-defined components increases. We then measured how well the resulting gene clusters reflected known regulatory mechanisms, and developed a set of metrics to assess the accuracy of the decomposition at a given dimension. We found that over-decomposition results in many independent components dominated by a single gene, whereas under-decomposition results in independent components that poorly capture the known regulatory structure. From these results, we developed a new method, called OptICA, for finding the optimal dimensionality that controls for both over- and under-decomposition. Specifically, OptICA selects the highest dimension that produces a low number of components that are dominated by a single gene. We show that OptICA outperforms two previously proposed methods for selecting the number of independent components across four transcriptomic databases of varying sizes. OptICA avoids both over-decomposition and under-decomposition of transcriptomic datasets resulting in the best representation of the organism’s underlying transcriptional regulatory network.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Mapathons versus automated feature extraction: a comparative analysis for strengthening immunization microplanning

Background: Social instability and logistical factors like the displacement of vulnerable populations, the difficulty of accessing these populations, and the lack of geographic information for hard-to-reach areas continue to serve as barriers to global essential immunizations (EI). Microplanning, a population-based, healthcare intervention planning method has begun to leverage geographic information system (GIS) technology and geospatial methods to improve the remote identification and mapping of vulnerable populations to ensure inclusion in outreach and immunization services, when feasible. We compare two methods of accomplishing a remote inventory of building locations to assess their accuracy and similarity to currently employed microplan line-lists in the study area. Methods: The outputs of a crowd-sourced digitization effort, or mapathon, were compared to those of a machine-learning algorithm for digitization, referred to as automatic feature extraction (AFE). The following accuracy assessments were employed to determine the performance of each feature generation method: (1) an agreement analysis of the two methods assessed the occurrence of matches across the two outputs, where agreements were labeled as “befriended” and disagreements as “lonely”; (2) true and false positive percentages of each method were calculated in comparison to satellite imagery; (3) counts of features generated from both the mapathon and AFE were statistically compared to the number of features listed in the microplan line-list for the study area; and (4) population estimates for both feature generation method were determined for every structure identified assuming a total of three households per compound, with each household averaging two adults and 5 children. Results: The mapathon and AFE outputs detected 92,713 and 53,150 features, respectively. A higher proportion (30%) of AFE features were befriended compared with befriended mapathon points (28%). The AFE had a higher true positive rate (90.5%) of identifying structures than the mapathon (84.5%). The difference in the average number of features identified per area between the microplan and mapathon points was larger (t = 3.56) than the microplan and AFE (t = -2.09) (alpha = 0.05). Conclusions: Our findings indicate AFE outputs had higher agreement (i.e., befriended), slightly higher likelihood of correctly identifying a structure, and were more similar to the local microplan line-lists than the mapathon outputs. These findings suggest AFE may be more accurate for identifying structures in high-resolution satellite imagery than mapathons. However, they both had their advantages and the ideal method would utilize both methods in tandem.

59 BASIC BIOLOGICAL SCIENCES↗

SMC 2021 : Analyzing Resource Utilization and User Behavior on Titan Supercomputer

Resource utilization statistics of submitted jobs on a supercomputer can help us understand how users from various scientific domains use HPC platforms and better design a job scheduler. We explore to generate insight regarding workload distribution and usage pattern domains from job scheduler trace, GPU failure information, and project-specific information collected from Titan supercomputer. Furthermore, we want to know how the scheduler performance varies over time and how the users' scheduling behavior changes following a system failure. These observations have the potential to provide valuable insight, which is helpful to prepare for system failures. These practices will help us develop and apply novel machine learning algorithms in understanding system behavior, requirement, and better scheduling of HPC systems. There are two datasets, RUR and GPU. RUR: This dataset is the job scheduler traces collected from the Titan supercomputerfrom 01/01/2015 to 07/31/2019 (2015.csv - 2019.csv). These were collected usingResource Utilization Report (RUR), a Cray-developed resource-usage data collectionand reporting system. It contains the usage information of its critical resources (CPU,Memory, GPU, and I/O) of each running job on Titan during that period [2]. ProjectAreas: Every job is associated with a project ID. TheProjectAreas.csvdatasetprovides a mapping of the project ID to its domain science. GPU: There have been some hardware-related issues in the GPUs in Titan that caused some GPUs to fail, sometimes irrecoverably during some job runs. This dataset provides information regarding these failures during the execution of the submitted jobs. GPUs on Titan are uniquely identified by a serial number (SN), and they are installed in a location. A GPU can be installed in a location, then removed from that location following a failure, and then re-installed in a different location after fixing the problem. If the failure can't be recovered, the GPU might be removed entirely from Titan. There are two prominent types of failures that resulted in the removal of GPUs from Titan: Double Bit Error (DBE) and Out of the Bus (OTB). The dataset (gc_full.csv) has the following fields: 1. SN : Serial number of a GPU 2. location : The location where it is installed 3. insert : The time when it was inserted into that location 4. remove : The time when it was removed from that location 5. duration : Amount of time the GPU spent in this location 6. out : If the device was taken out entirely w/o a re-installment into a new location. 7. event : If the GPU was taken out entirely, the reason for its removal. To learn more about this dataset, please refer to the git repositoryhttps://github.com/olcf/TitanGPULifeand the related publication [1]. References [1] George Ostrouchov, Don Maxwell, Rizwan A Ashraf, Christian Engelmann, MallikarjunShankar, and James H Rogers. Gpu lifetimes on titan supercomputer: Survival analysisand reliability. InSC20: International Conference for High Performance Computing,Networking, Storage and Analysis, pages 1-14. IEEE, 2020. [2] Feiyi Wang, Sarp Oral, Satyabrata Sen, and Neena Imam. Learning from five-yearresource-utilization data of titan system. In2019 IEEE International Conference onCluster Computing (CLUSTER), pages 1-6. IEEE, 2019.

42 ENGINEERING↗

SMC 2021 Data Challenge: Analyzing Resource Utilization and User Behavior on Titan Supercomputer

Resource utilization statistics of submitted jobs on a supercomputer can help us understand how users from various scientific domains use HPC platforms and better design a job scheduler. We explore to generate insight regarding workload distribution and usage pattern domains from job scheduler trace, GPU failure information, and project-specific information collected from Titan supercomputer. Furthermore, we want to know how the scheduler performance varies over time and how the users' scheduling behavior changes following a system failure. These observations have the potential to provide valuable insight, which is helpful to prepare for system failures. These practices will help us develop and apply novel machine learning algorithms in understanding system behavior, requirement, and better scheduling of HPC systems. There are two datasets, RUR and GPU: RUR dataset is the job scheduler traces collected from the Titan supercomputer from 01/01/2015 to 07/31/2019 (2015.csv - 2019.csv). These were collected using resource Utilization Report (RUR), a Cray-developed resource-usage data collection and reporting system. It contains the usage information of its critical resources (CPU, Memory, GPU, and I/O) of each running job on Titan during that period (https://ieeexplore.ieee.org/abstract/document/8891001). It includes ProjectAreas as additional information, every job is associated with a project ID. TheProjectAreas.csv dataset provides a mapping of the project ID to its domain science. GPU dataset has information regarding GPU failure on Titan. There have been some hardware-related issues in the GPUs in Titan that caused some GPUs to fail, sometimes irrecoverably during some job runs. This dataset provides information regarding these failures during the execution of the submitted jobs. GPUs on Titan are uniquely identified by a serial number (SN), and they are installed in a location. A GPU can be installed in a location, then removed from that location following a failure, and then re-installed in a different location after fixing the problem. If the failure can't be recovered, the GPU might be removed entirely from Titan. There are two prominent types of failures that resulted in the removal of GPUs from Titan: Double Bit Error (DBE) and Out of the Bus (OTB). The dataset (gc_full.csv) has seven attributes, we provided a short description of these attributes in the ReadMe file. To learn more about this dataset, please refer to the git repository https://github.com/olcf/TitanGPULife and the related publication (https://ieeexplore.ieee.org/abstract/document/9355319).

42 ENGINEERING↗

Investigation of Delayed Neutron Sensitivities for Several ICSBEP Benchmarks using MCNP

The effective delayed neutron (β$_{eff}$) is a very important parameter for reactor and criticality applications. This parameter is equal to the difference in reactivity between delayed critical ($k_{eff}$ = 1, which requires both prompt and delayed neutrons to achieve criticality) and prompt critical ($k_p$ = 1, which requires only prompt neutrons to achieve criticality). This is often referred to as the delayed critical "window" (the region of criticality between delayed and prompt critical). β$_{eff}$ is a reactor kinetics parameters and depends on the nuclides in the system that undergo fission as well as the spectral characteristics of the system. Measurements of β$_{eff}$ have been performed for many criticality experiments. The EUCLID (Experiments Underpinned by Computational Learning for Improvements in nuclear Data) project at Los Alamos National Laboratory (LANL) aims to constrain nuclear data by using a suite of measurement types beyond $k_{eff}$. Our team has recently investigated the use of pulsed spheres for nuclear data validation. Several other measurement methods are also of interest, including β$_{eff}$ (investigated here) and reactivity coefficients (investigated in a separate work at this same meeting). One focus of our work is to determine if other methods are complimentary to the critical experiments already utilized for nuclear data validation. This is important because if a method has similar sensitivities then it will not be particularly useful for nuclear data validation as it will provide the same information as the critical experiments already being used. Here "similar" could refer to several characteristics, one being the shape of a sensitivity profile over energy. In the future, these methods will be utilized (with both existing and new experiments) in machine learning algorithms for nuclear validation, similar to what is currently done for criticality experiments. In order to use a measurement type for nuclear validation, it is necessary to obtain cross-section sensitivities for that parameter. This work looks at one approach to estimate β$_{eff}$ sensitivities by utilizing $k_{eff}$ sensitivities within Monte Carlo N-Particle ® Code Version 6.2. This is applied to several criticality benchmarks. Results are compared and the benefits and limitations of this approach are discussed.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Machine learning enabled lineshape analysis in optical two-dimensional coherent spectroscopy

Optical two-dimensional (2D) coherent spectroscopy excels in studying coupling and dynamics in complex systems. The dynamical information can be learned from lineshape analysis to extract the corresponding linewidth. However, it is usually challenging to fit a 2D spectrum, especially when the homogeneous and inhomogeneous linewidths are comparable. We implemented a machine learning algorithm to analyze 2D spectra to retrieve homogeneous and inhomogeneous linewidths. The algorithm was trained using simulated 2D spectra with known linewidth values. The trained algorithm can analyze both simulated (not used in training) and experimental spectra to extract the homogeneous and inhomogeneous linewidths. Finally, this approach can be potentially applied to 2D spectra with more sophisticated spectral features.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Distributed fiber sensor and machine learning data analytics for pipeline protection against extrinsic intrusions and intrinsic corrosions

This paper presents an integrated technical framework to protect pipelines against both malicious intrusions and piping degradation using a distributed fiber sensing technology and artificial intelligence. A distributed acoustic sensing (DAS) system based on phase-sensitive optical time-domain reflectometry (φ-OTDR) was used to detect acoustic wave propagation and scattering along pipeline structures consisting of straight piping and sharp bend elbow. Signal to noise ratio of the DAS system was enhanced by femtosecond induced artificial Rayleigh scattering centers. Data harnessed by the DAS system were analyzed by neural network-based machine learning algorithms. The system identified with over 85% accuracy in various external impact events, and over 94% accuracy for defect identification through supervised learning and 71% accuracy through unsupervised learning.

Peng, Zhaoqiang↗

UNNT: A novel Utility for comparing Neural Net and Tree-based models

The use of deep learning (DL) is steadily gaining traction in scientific challenges such as cancer research. Advances in enhanced data generation, machine learning algorithms, and compute infrastructure have led to an acceleration in the use of deep learning in various domains of cancer research such as drug response problems. In our study, we explored tree-based models to improve the accuracy of a single drug response model and demonstrate that tree-based models such as XGBoost (eXtreme Gradient Boosting) have advantages over deep learning models, such as a convolutional neural network (CNN), for single drug response problems. However, comparing models is not a trivial task. To make training and comparing CNNs and XGBoost more accessible to users, we developed an open-source library called UNNT (A novel Utility for comparing Neural Net and Tree-based models). The case studies, in this manuscript, focus on cancer drug response datasets however the application can be used on datasets from other domains, such as chemistry.

59 BASIC BIOLOGICAL SCIENCES↗

Smart Manufacturing Pathways for Industrial Decarbonization and Thermal Process Intensification

Rapid decarbonization is fast becoming the primary environmental and sustainability initiative for many economic sectors. Industry consumes more than 30 % of all primary energy in the United States and accounts for nearly 25 % of all greenhouse gas (GHG) emissions. More than 70 % of energy consumed by the industrial sector is related to thermal processes, which are also the largest contributors of carbon emissions, overwhelmingly due to the combustion of fossil fuels. Thermal process intensification (TPI) seeks to dramatically improve the energy performance of thermal systems through technology pillars focusing on alternative energy sources and processes, supplemental technologies, and waste heat management. The impacts of TPI have significant overlap with the goals of industrial decarbonization (ID) that seeks to phase out all GHG emissions from industrial activities. Emerging supplemental technologies such as smart manufacturing (SM) and the industrial internet of things (IoT) enable significant opportunities for the optimization of manufacturing processes. Combining strategies for TPI and ID with SM and IoT can open and enhance existing opportunities for saving time and energy via approaches such as tighter control of temperature zones, better adjustment of thermal systems for variations in production levels and feedstock properties, and increased process throughput. Data collected by smart processes will also enable new advanced solutions such as digital twins and machine learning algorithms to further improve thermal system savings. Herein, this paper examines the individual pathways of TPI, ID, and SM and how the combination of all three can accelerate energy and GHG reductions.

42 ENGINEERING↗

CalderaCast Web Interface

CalderaCast may be accessed as a web-based tool at the first link in the references section of this dataset. All of the necessary datasets to run the tool are built into the simulation software running behind the web interface. These input datasets are referenced by the additional links in the references section below. Many of those datasets are taken into machine-learning algorithms by the Caldera team and heavily processed to create internal datasets, which are then relied upon by the simulation to produce individual results. These internal datasets are not accessible and are not necessary for use of the CalderaCast tool.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗