Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data mining”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Coastal Typologies: Methods for Improving Representation of Arctic Coastal Environments, Starting with Alaska's Northern Slope [Slides]

The objectives of the project included: Broaden QGIS (mapping) and Python (data analysis) skill set; Contribute to overall development of coastal typologies by conducting data mining and analysis; Expand scientific literacy; Narrow data sources to those most useful for project; Convert chosen data sets into image formats for analysis; and, Utilize python to analyze images.

58 GEOSCIENCES↗

Spatio-Temporal Surrogates for Interaction of a Jet with High Explosives: Part II - Clustering Extremely High-Dimensional Grid-Based Data

Building an accurate surrogate model for the spatio-temporal outputs of a computer simulation is a challenging task. A simple approach to improve the accuracy of the surrogate is to cluster the outputs based on similarity and build a separate surrogate model for each cluster. This clustering is relatively straightforward when the output at each time step is of moderate size. However, when the spatial domain is represented by a large number of grid points, numbering in the millions, the clustering of the data becomes more challenging. In this report, we consider output data from simulations of a jet interacting with high explosives. These data are available on spatial domains of different sizes, at grid points that vary in their spatial coordinates, and in a format that distributes the output across multiple files at each time step of the simulation. We first describe how we bring these data into a consistent format prior to clustering. Borrowing the idea of random projections from data mining, we reduce the dimension of our data by a factor of thousand, making it possible to use the iterative k-means method for clustering. We show how we can use the randomness of both the random projections, and the choice of initial centroids in k-means clustering, to determine the number of clusters in our data set. Our approach makes clustering of extremely high dimensional data tractable, generating meaningful cluster assignments for our problem, despite the approximation introduced in the random projections.

97 MATHEMATICS AND COMPUTING↗

Big Data Analysis and Technical Review of Regeneration for Carbon Capture Processes

Carbon capture remains an integral technology to mitigate pollution from one of the most prevalent greenhouse gases. CO 2 desorption/absorbent regeneration for both solid- and liquid-based systems is widely recognized as an energy-intensive and costly process operation. Consequently, tremendous work was devoted towards developing new absorbents and regeneration processes to promote their economic feasibility for extensive implementation. In this review, we broadly and deeply review more than 10,000 papers and extract the hidden trends of carbon capture and absorbents regeneration in the past few decades, using a novel data-mining analysis technique. We comprehensively analyzed an array of recent absorbent regeneration methods utilized in post-combustion, pre-combustion, carbon capture from industrial point sources, and direct air carbon capture, with an emphasis on sorbent and solvent-based techniques. In conclusion, advanced regeneration methods in these techniques were illustrated and discussed, followed by recommendations for further research efforts.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A highly wear resistant nanostructured bainitic steel with accelerated transformation kinetics

A coupled Calculation of Phase Diagrams (CALPHAD), machine learning, and data mining approach was used to design a new, highly wear-resistant nanostructured bainitic steel. Arc melting of the designed compositions, dilatometry, and advanced microscopy indicate that the designed steel had a nanoscale dual-phase structure of ferrite and austenite (approximately 50 nm) with kinetics 7x faster for the onset of bainite and 2x faster for complete transformation. Under dry sliding conditions using the current state-of-the-art AISI 52100 bearing steel as the counter sample, the designed steel little to no wear, indicating its potential for applications in high-wear service conditions.

36 MATERIALS SCIENCE↗

ECO: the Evidence and Conclusion Ontology, an update for 2022

The Evidence and Conclusion Ontology (ECO) is a community resource that provides an ontology of terms used to capture the type of evidence that supports biomedical annotations and assertions. Consistent capture of evidence information with ECO allows tracking of annotation provenance, establishment of quality control measures, and evidence-based data mining. ECO is in use by dozens of data repositories and resources with both specific and general areas of focus. ECO is continually being expanded and enhanced in response to user requests as well as our aim to adhere to community best-practices for ontology development. The ECO support team engages in multiple collaborations with other ontologies and annotating groups. Here we report on recent updates to the ECO ontology itself as well as associated resources that are available through this project. ECO project products are freely available for download from the project website (https://evidenceontology.org/) and GitHub (https://github.com/evidenceontology/evidenceontology). ECO is released into the public domain under a CC0 1.0 Universal license.

59 BASIC BIOLOGICAL SCIENCES↗

Critical Assessment of Electronic Structure Descriptors for Predicting Perovskite Catalytic Properties

The discovery and design of materials which can efficiently catalyze the oxygen reduction and evolution reactions at reduced temperatures is important for facilitating the widespread adoption of fuel cell and electrolyzer technologies. Numerous studies have produced correlations between catalytic properties, such as oxygen surface exchange or electrode area specific resistance (ASR), and properties of the catalyst material. However, correlations have historically been limited in scope (e.g., using only a few materials or at a single temperature) and it has been difficult to provide detailed assessments of their robustness. Here, in this study, we assess the ability of the O p-band center electronic structure descriptor, obtained from density functional theory (DFT) calculations, to correlate with oxygen surface exchange rates, diffusivities, and area specific resistances for a large database of perovskite oxide catalytic properties. By data mining the literature, we obtain 747 catalytic property value data points spanning 299 unique perovskite compositions from 313 studies. We assess linear correlations of each property with the O p-band center and find generally modest correlations that are qualitatively useful (prediction mean absolute errors of about 0.5 log units are typical), where the correlations are improved at higher temperatures (e.g., 800 °C vs. 500 °C) and significantly improve when considering fits to the subset of materials which have multiple independent measurements. These findings suggest that the spread of property data is significantly influenced by experimental uncertainty, and subsequent measurements of additional materials will likely improve the O p-band center correlations.

30 DIRECT ENERGY CONVERSION↗

Leveraging large language models to address data scarcity in machine learning for graphene synthesis

Machine learning in experimental materials science faces significant challenges due to the scarcity of data, which are costly and time-consuming to generate, particularly when relying on in-house experiments. Literature data mining offers a potential solution but introduces issues like mixed data quality, inconsistent formats, and non-uniform reporting of synthesis parameters, resulting in partially missing and heterogeneous features across the dataset. Here, we propose data imputation and feature engineering methods that employ pre-trained large language models (LLMs) to enhance machine learning performance on scarce, heterogeneous datasets, demonstrated on graphene CVD synthesis data and the ML-HydPARK hydrogen storage dataset. GPT models perform data imputation via tailored prompting and semantic normalization of inconsistently reported features through embeddings, for example, to harmonize the complex nomenclature of CVD substrates. Beyond yielding more diverse and richer feature representations than traditional methods such as K-nearest neighbors (KNN) and Multivariate Imputation by Chained Equations (MICE), LLM-based data imputation is evaluated against dataset characteristics and prompting strategies. We vary the level of autonomy granted to the LLM, from generic prompting that leverages pre-trained knowledge for autonomous data generation to data-informed prompting that constrains outputs using target-specific information, and demonstrate which level of autonomy yields superior imputation performance across datasets and feature types. The proposed data engineering methods markedly improve downstream performance; for example, in graphene layer number classification using a support vector machine (SVM), binary accuracy increases from 39% to 65% and ternary accuracy from 52% to 72%. Fine-tuning experiments on both datasets show that combining our proposed LLM-based data imputation and feature encoding methods with numerical machine learning predictors outperforms standalone fine-tuned LLM predictors in data-scarce settings. The proposed strategies emphasize data enhancement techniques rather than refining learning architectures or regularizing loss functions, offering a broadly applicable framework for improving machine learning performance on scarce, inhomogeneous datasets.

Chemical vapor deposition↗

Clinical knowledge extraction via sparse embedding regression (KESER) with multi-center large scale electronic health record data

The increasing availability of electronic health record (EHR) systems has created enormous potential for translational research. However, it is difficult to know all the relevant codes related to a phenotype due to the large number of codes available. Traditional data mining approaches often require the use of patient-level data, which hinders the ability to share data across institutions. In this project, we demonstrate that multi-center large-scale code embeddings can be used to efficiently identify relevant features related to a disease of interest. We constructed large-scale code embeddings for a wide range of codified concepts from EHRs from two large medical centers. We developed knowledge extraction via sparse embedding regression (KESER) for feature selection and integrative network analysis. We evaluated the quality of the code embeddings and assessed the performance of KESER in feature selection for eight diseases. Besides, we developed an integrated clinical knowledge map combining embedding data from both institutions. The features selected by KESER were comprehensive compared to lists of codified data generated by domain experts. Features identified via KESER resulted in comparable performance to those built upon features selected manually or with patient-level data. The knowledge map created using an integrative analysis identified disease-disease and disease-drug pairs more accurately compared to those identified using single institution data. Analysis of code embeddings via KESER can effectively reveal clinical knowledge and infer relatedness among codified concepts. KESER bypasses the need for patient-level data in individual analyses providing a significant advance in enabling multi-center studies using EHR data.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Unsupervised learning of representative local atomic arrangements in molecular dynamics data

Molecular dynamics (MD) simulations present a data-mining challenge, given that they can generate a considerable amount of data but often rely on limited or biased human interpretation to examine their information content. By not asking the right questions of MD data we may miss critical information hidden within it. Here we combine dimensionality reduction (UMAP) and unsupervised hierarchical clustering (HDBSCAN) to quantitatively characterize prevalent coordination environments of chemical species within MD data. By focusing on local coordination, we significantly reduce the amount of data to be analyzed by extracting all distinct molecular formulas within a given coordination sphere. We then efficiently combine UMAP and HDBSCAN with alignment or shape-matching algorithms to partition these formulas into structural isomer families indicating their relative populations. The method was employed to reveal details of cation coordination in electrolytes based on molecular liquids.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Critical review of functionalized silica sorbent strategies for selective extraction of rare earth elements from acid mine drainage

We report the ubiquitous and growing global reliance on rare earth elements (REEs) for modern technology and the need for reliable domestic sources underscore the rising trend in REE-related research. Adsorption-based methods for REE recovery from liquid waste sources are well-positioned to compete with those of solvent extraction, both because of their expected lower negative environmental impact and simpler process operations. Functionalized silica represents a rising category of low cost and stable sorbents for heavy metal and REE recovery. These materials have collectively achieved high capacity and/or high selective removal of REEs from ideal solutions and synthetic or real coal wastewater and other leachate source. These sorbents are competitive with conventional materials, such as ion exchange resins, activated carbon; and novel polymeric materials like ion-imprinted particles and metal organic frameworks (MOFs). This critical review first presents a data mining analysis for rare earth element recovery publications indexed in Web of science, highlighting changes in REE recovery research foci and confirming the sharply growing interest in functionalized silica sorbents. A detailed examination of sorbent formulation and operation strategies to selectively separate heavy (HREE), middle (MREE), and light (LREE) REEs from the aqueous sources is presented. Selectivity values for sorbents were largely calculated from available figure data and gauged the success of the associated strategies, primarily: (1) silane-grafted ligands, (2) impregnated ligands, and (3) bottom-up ligand/silica hybrids. These were often accompanied by successful co-strategies, especially bite angle control, site saturation, and selective REE elution. Recognizing the need to remove competing fouling metals to achieve purified REE “baskets,” we highlight techniques for eliminating these species from acid mine drainage (AMD) and suggest a novel adsorption-based process for purified REE extraction that could be adapted to different water systems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Open Chemistry, JupyterLab, REST, and quantum chemistry

Quantum chemistry must evolve if it wants to fully leverage the benefits of the internet age, where the worldwide web offers a vast tapestry of tools that enable users to communicate and interact with complex data at the speed and convenience of a button press. The Open Chemistry project has developed an open-source framework that offers an end-to-end solution for producing, sharing, and visualizing quantum chemical data interactively on the web using an array of modern tools and approaches. These tools build on some of the best open-source community projects such as Jupyter for interactive online notebooks, coupled with 3D accelerated visualization, state-of-the-art computational chemistry codes including NWChem and Psi4, and emerging machine learning and data mining tools such as ChemML and ANI. They offer flexible formats to import and export data, along with approaches to compare computational and experimental data.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Asi Nuclear Energy Sensors Data Portal Chatbot And Data Structuring Tool

The Idaho National Laboratory (INL) is advancing the development of an AI-powered chatbot and data structuring tool specifically designed to accelerate data mining processes for sensor-related information and seamlessly integrate the results into the ASI Sensors Data Portal (https://nes.energy.gov/). By doing so, the software aims to enhance the accessibility, usability, and organization of sensor data for nuclear energy applications. The software initial phase focuses on retrieving comprehensive datasets, prioritizing the past five years of publicly available information from the Office of Scientific and Technical Information (OSTI). These datasets will be meticulously processed to ensure compatibility, employing cleaning and preprocessing steps to eliminate irrelevant, incomplete, or corrupted information, thus establishing a robust foundation for subsequent AI use. The data will serve as the backbone for training an AI model and chatbot, which will act as an interactive tool enabling users to ask complex, context-specific questions and receive accurate, validated answers derived from constrained literature. In parallel, the project incorporates a data structuring process supported by AI to organize sensor information from multiple sources into a standardized format. This structured data will include detailed sensor specifications, such as measurement range, applications, accuracy, and operating conditions, generated and documented with AI. These specifications will be systematically integrated into the sensor portal. To maintain the highest levels of accuracy and relevance, all AI-generated outputs will be reviewed and validated by subject matter experts (SMEs), with additional fields or parameters added as needed. Future stages of the project aim to expand the dataset beyond OSTI to include other sources and potentially incorporate unclassified controlled information (UCI) with restricted access protocols to address security and confidentiality requirements.

Mapes, NormanJ. [Idaho National Laboratory (INL), ↗

Compiler and Runtime Approaches to Enable Large-Scale Irregular Programs. Final report, July 2013 - July 2019

While regular algorithms, characterized by operations on dense matrices and arrays, have long been the mainstay of scientific, high-performance computing, irregular algorithms, which feature unpredictable accesses to pointer-based data structures, are becoming increasingly common in high performance computing, arising in graph analysis, data mining and visualization, among other domains. Unfortunately, the defining characteristics of irregular applications, their dynamic, unpredictable, data-dependent access patterns and data layouts, make achieving high performance on large scale systems difficult. Scaling applications to peta- and exa-scale requires carefully controlling communication and data movement and placement, an inherently difficult task when access patterns and data layouts are unpredictable! Most irregular applications that attain high performance must be painstakingly hand-written and hand-tuned, with few common principles or paradigms uniting various implementations and easing future development. Despite the increasing importance of irregular applications, there is little programmer knowledge, and even less compiler ability, devoted to optimizing them. This project aims to solve these problems. By allowing programmers to write irregular applications in high level forms, with at most a few annotations highlighting key structural properties, programmers can focus on developing their algorithms and methods. The compiler and run-time system can take on the tedious task of optimizing the application for execution at large scales, and can automatically provide efficient implementations. This will provide portability and ease maintenance for existing irregular applications, but, more importantly, open up whole new domains of computational science to large-scale, high-performance simulation codes.

97 MATHEMATICS AND COMPUTING↗

A chemistry-informed hybrid machine learning approach to predict metal adsorption onto mineral surfaces

Historically, surface complexation model (SCM) constants and distribution coefficients (K d ) have been employed to quantify mineral-based retardation effects controlling the fate of metals in subsurface geologic systems. Our recent SCM development workflow, based on the Lawrence Livermore National Laboratory Surface Complexation/Ion Exchange (L-SCIE) database, illustrated a community FAIR data approach to SCM development by predicting uranium(VI)-quartz adsorption for a large number of literature-mined data. Here, we present an alternative hybrid machine learning (ML) approach that shows promise in achieving equivalent high-quality predictions compared to traditional surface complexation models. At its core, the hybrid random forest (RF) ML approach is motivated by the proliferation of incongruent SCMs in the literature that limit their applicability in reactive transport models. Our hybrid ML approach implements PHREEQC-based aqueous speciation calculations; values from these simulations are automatically used as input features for a random forest (RF) algorithm to quantify adsorption and avoid SCM modeling constraints entirely. Named the LLNL Speciation Updated Random Forest (L-SURF) model, this hybrid approach is shown to have applicability to U(VI) sorption cases driven by both ion-exchange and surface complexation, as is shown for quartz and montmorillonite cases. The approach can be applied to reactive transport modeling and may provide an alternative to the costly development of self-consistent SCM reaction databases.

38 RADIATION CHEMISTRY, RADIOCHEMISTRY, AND NUCLEA↗

A Sparse Tensor Benchmark Suite for CPUs and GPUs

Tensor computations present significant performance chal- lenges that impact a wide spectrum of applications ranging from machine learning, healthcare analytics, social network analysis, data mining to quantum chemistry and signal processing. Efforts to improve the perfor- mance of tensor computations include exploring data layout, execution scheduling, and parallelism in common tensor kernels. This work presents a benchmark suite for arbitrary-order sparse tensor kernels using state- of-the-art tensor formats: coordinate (COO) and hierarchical coordinate (HiCOO) on CPUs and GPUs. It presents a set of reference tensor kernel implementations that are compatible with real-world tensors and power law tensors extended from synthetic graph generation techniques. We also propose Roofline performance models for these kernels to provide insights of computer platforms from sparse tensor view. This benchmark suite along with the synthetic tensor generator is publicly available.

Li, Jiajia↗

Text-mined dataset of gold nanoparticle synthesis procedures, morphologies, and size entities

Abstract Gold nanoparticles are highly desired for a range of technological applications due to their tunable properties, which are dictated by the size and shape of the constituent particles. Many heuristic methods for controlling the morphological characteristics of gold nanoparticles are well known. However, the underlying mechanisms controlling their size and shape remain poorly understood, partly due to the immense range of possible combinations of synthesis parameters. Data-driven methods can offer insight to help guide understanding of these underlying mechanisms, so long as sufficient synthesis data are available. To facilitate data mining in this direction, we have constructed and made publicly available a dataset of codified gold nanoparticle synthesis protocols and outcomes extracted directly from the nanoparticle materials science literature using natural language processing and text-mining techniques. This dataset contains 5,154 data records, each representing a single gold nanoparticle synthesis article, filtered from a database of 4,973,165 publications. Each record contains codified synthesis protocols and extracted morphological information from a total of 7,608 experimental and 12,519 characterization paragraphs.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Multi-kernel Edge Attention Graph Autoencoder

MEAGraph (Multi-kernel Edge Attention Graph Autoencoder) is a graph-based autoencoder model designed for unsupervised data mining for datasets used in machine learning potentials. It provides accurate clustering for atomic environment identification, unsupervised and unlabeled data pruning for dataset construction.

Sun, Hong↗