Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data mining”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

My vehicle is a data mine

In this talk we explore how analysis of vehicle data provides information of individual vehicle behaviors, information of other vehicles in the flow of traffic, and insights into the behavior of drivers. Over the last two decades traditional passenger vehicles have been transformed from integrated two-port electrical nodes to cyber-physical systems of communicating computational nodes whose individual state and control variables are shared on a standard controller area network (CAN) bus. As driver assistance systems have crept into vehicles as safety features, driver behaviors can be observed through analysis of the data streams on the CAN bus as these new nodes communicate with one another. The properties of these data streams, as well as architectures and approaches to gather the data, are important to consider when drawing conclusions on the relevance of the data in making decisions at varying levels of the information hierarchy. We will demonstrate several technical challenges associated with these data collection processes, as well as preliminary results that demonstrate application relevance of the data to behavior, traffic, and systems domains.

42 ENGINEERING↗

Unsupervised atomic data mining via multi-kernel graph autoencoders for machine learning force fields

Constructing a chemically diverse dataset while avoiding sampling bias is critical to training efficient and generalizable force fields. However, in computational chemistry and materials science, many common dataset generation techniques are prone to oversampling regions of the potential energy surface. Furthermore, these regions can be difficult to identify and isolate from each other or may not align well with human intuition, making it challenging to systematically remove bias in the dataset. While traditional clustering and pruning (down-sampling) approaches can be useful for this, they can often lead to information loss or a failure to properly identify distinct regions of the potential energy surface due to difficulties associated with the high dimensionality of atomic descriptors. In this work, we introduce the Multi-kernel Edge Attention-based Graph Autoencoder (MEAGraph) model, an unsupervised approach for analyzing atomic datasets. MEAGraph combines multiple linear kernel transformations with attention-based message passing to capture geometric sensitivity and enable effective dataset pruning without relying on labels or extensive training. Demonstrated applications on niobium, tantalum, and iron datasets show that MEAGraph efficiently groups similar atomic environments, allowing for the use of basic pruning techniques for removing sampling bias. This approach provides an effective method for representation learning and clustering that can be used for data analysis, outlier detection, and dataset optimization.

Materials science↗

Computational Estimation by Scientific Data Mining with Classical Methods to Automate Learning Strategies of Scientists

Experimental results are often plotted as 2-dimensional graphical plots (aka graphs) in scientific domains depicting dependent versus independent variables to aid visual analysis of processes. Repeatedly performing laboratory experiments consumes significant time and resources, motivating the need for computational estimation. The goals are to estimate the graph obtained in an experiment given its input conditions, and to estimate the conditions that would lead to a desired graph. Existing estimation approaches often do not meet accuracy and efficiency needs of targeted applications. We develop a computational estimation approach called AutoDomainMine that integrates clustering and classification over complex scientific data in a framework so as to automate classical learning methods of scientists. Knowledge discovered thereby from a database of existing experiments serves as the basis for estimation. Challenges include preserving domain semantics in clustering, finding matching strategies in classification, striking a good balance between elaboration and conciseness while displaying estimation results based on needs of targeted users, and deriving objective measures to capture subjective user interests. These and other challenges are addressed in this work. The AutoDomainMine approach is used to build a computational estimation system, rigorously evaluated with real data in Materials Science. Our evaluation confirms that AutoDomainMine provides desired accuracy and efficiency in computational estimation. It is extendable to other science and engineering domains as proved by adaptation of its sub-processes within fields such as Bioinformatics and Nanotechnology.

Computer Science↗

A systematic analysis and data mining of opioid-related adverse events submitted to the FAERS database

The opioid epidemic has become a serious national crisis in the United States. An indepth systematic analysis of opioid-related adverse events (AEs) can clarify the risks presented by opioid exposure, as well as the individual risk profiles of specific opioid drugs and the potential relationships among the opioids. In this study, 92 opioids were identified from the list of all Food and Drug Administration (FDA)-approved drugs, annotated by RxNorm and were classified into 13 opioid groups: buprenorphine, codeine, dihydrocodeine, fentanyl, hydrocodone, hydromorphone, meperidine, methadone, morphine, oxycodone, oxymorphone, tapentadol, and tramadol. A total of 14,970,399 AE reports were retrieved and downloaded from the FDA Adverse Events Reporting System (FAERS) from 2004, Quarter 1 to 2020, Quarter 3. After data processing, Empirical Bayes Geometric Mean (EBGM) was then applied which identified 3317 pairs of potential risk signals within the 13 opioid groups. Based on these potential safety signals, a comparative analysis was pursued to provide a global overview of opioid-related AEs for all 13 groups of FDA-approved prescription opioids. The top 10 most reported AEs for each opioid class were then presented. Both network analysis and hierarchical clustering analysis were conducted to further explore the relationship between opioids. Results from the network analysis revealed a close association among fentanyl, oxycodone, hydrocodone, and hydromorphone, which shared more than 22 AEs. In addition, much less commonly reported AEs were shared among dihydrocodeine, meperidine, oxymorphone, and tapentadol. On the contrary, the hierarchical clustering analysis further categorized the 13 opioid classes into two groups by comparing the full profiles of presence/absence of AEs. The results of network analysis and hierarchical clustering analysis were not only consistent and cross-validated each other but also provided a better and deeper understanding of the associations and relationships between the 13 opioid groups with respect to their adverse effect profiles.

Research & Experimental Medicine↗

Data Mining for Faster, Interpretable Solutions to Inverse Problems:A Case Study Using Additive Manufacturing

Solving inverse problems, where we nd the input values that result in desired values of outputs, can be challenging. The solution process is often computationally expensive and it can be di cult to interpret the solution in high-dimensional input spaces. In this paper, we use a problem from additive manufacturing to address these two issues with the intent of making it easier to solve inverse problems and exploit their results. First, focusing on Gaussian process surrogates that are used to solve inverse problems, we describe how a simple modi cation to the idea of tapering can substantially speed up the surrogate without losing accuracy in prediction. Second, we demonstrate that Kohonen self-organizing maps can be used to visualize and interpret the solution to the inverse problem in the high-dimensional input space. For our data set, as not all input dimensions are equally important, we show that using weighted distances results in a better organized map that makes the relationships among the inputs obvious

97 MATHEMATICS AND COMPUTING↗

Data Mining – Image Analysis of Radiography for Zr Redistribution

The research effort described in this report represents a first attempt to investigate the radial redistribution of Zr in ternary fuel alloys U-xPu-10Zr (x = 0, 8, 19) irradiated in the in-reactor fuel experiments during the operation of the U.S. Department of Energy’s (DOE’s) Experimental Breeder Reactor II (EBR-II) at Idaho National Laboratory (INL) using methods of image processing on post-irradiation neutron radiographs. Approximately 130,000 metal fuel pins were irradiated in EBR-II during its 30 years of operation to develop and characterize existing and prospective fuels. For many of the metal fuel irradiation experiments, neutron radiography imaging was performed historically now allowing application of modern image analysis techniques to characterize fuel behavior, such as fuel swelling, fluff formation, and now, fuel alloy constituent redistribution. The redistribution of fuel components depends on the temperature field, radially, within the fuel. Specifically, Zr is expected to redistribute radially towards the center of the pin, as well as towards the outer zones. Currently, direct imaging of a pin cross-section through optical methods or scanning electron microscopy (SEM) is used to study the constituents’ redistribution, which is very time-consuming and can only be applied to a limited number of pins. An automated image processing technique allowing for the investigation of fuel radial redistribution zones would significantly accelerate data analysis. In general, if the fuel temperature is hot enough, three redistribution zones are expected, corresponding to the main fuel components (e.g., U, Pu, Zr). While more assessment will be performed in fiscal year (FY)-2023, it seems possible to differentiate the pins according to their fuel composition using image analysis techniques on neutron radiographs of metallic fuel pins based on the analysis to date.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Data mining and computational screening of Rashba-Dresselhaus splitting and optoelectronic properties in two-dimensional perovskite materials

Recent developments highlighting the promise of two-dimensional perovskites have vastly increased the compositional search space in the perovskite family. This presents a great opportunity for the realization of highly performant devices and practical challenges associated with the identification of candidate materials. High-fidelity computational screening offers great value in this regard. In this study, we carry out a multiscale computational workflow, generating a dataset of two-dimensional perovskites in the Dion-Jacobson and Ruddlesden-Popper phases. Our dataset comprises ten B-site cations, four halogens, and over 20 organic cations across over 2000 materials. We compute electronic properties, thermoelectric performance, and numerous geometric characteristics. Furthermore, we introduce a framework for the high-throughput computation of Rashba-Dresselhaus splitting. Finally, we use this dataset to train machine learning models for the accurate prediction of band gaps, candidate Rashba-Dresselhaus materials, and partial charges. The work presented herein can aid future investigations of two-dimensional perovskites with targeted applications in mind.

14 SOLAR ENERGY↗

Exploring two-dimensional van der Waals heavy-fermion material: Data mining theoretical approach

Abstract The discovery of two-dimensional (2D) van der Waals (vdW) materials often provides interesting playgrounds to explore novel phenomena. One of the missing components in 2D vdW materials is the intrinsic heavy-fermion systems, which can provide an additional degree of freedom to study quantum critical point (QCP), unconventional superconductivity, and emergent phenomena in vdW heterostructures. Here, we investigate 2D vdW heavy-fermion candidates through the database of experimentally known compounds based on dynamical mean-field theory calculation combined with density functional theory (DFT+DMFT). We have found that the Kondo resonance state of CeSiI does not change upon exfoliation and can be easily controlled by strain and surface doping. Our result indicates that CeSiI is an ideal 2D vdW heavy-fermion material and the quantum critical point can be identified by external perturbations.

36 MATERIALS SCIENCE↗

Data Mining of the Rocky Flats Library Archive for Plutonium Compatibility Studies

The objective of this project was to create a searchable database of plutonium compatibility studies from the Rocky Flats Archive. The Rocky Flats Plant was a manufacturing complex in Golden, Colorado that produced nuclear weapons. It primarily focused on producing plutonium pits, which were the cores of many nuclear implosion-type weapons. Because plutonium is an extremely reactive metal, avoiding the use of incompatible materials is imperative to avoid damaging the plutonium pit during its production. The Rocky Flats Plant had conducted extensive surveys of materials used in plutonium pit fabrication for their compatibility with plutonium.

36 MATERIALS SCIENCE↗