Engineering PapersSearch

SEARCH · Engineering Papers

Results for “data mining system”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Polynomial chaos expansions on principal geodesic Grassmannian submanifolds for surrogate modeling and uncertainty quantification

In this work we introduce a manifold learning-based surrogate modeling framework for uncertainty quantification in high-dimensional stochastic systems. Our first goal is to perform data mining on the available simulation data to identify a set of low-dimensional (latent) descriptors that efficiently parameterize the response of the high-dimensional computational model. To this end, we employ Principal Geodesic Analysis on the Grassmann manifold of the response to identify a set of disjoint principal geodesic submanifolds, of possibly different dimension, that captures the variation in the data. Since operations on the Grassmann require the data to be concentrated, we propose an adaptive algorithm based on Riemannian K-means and the minimization of the sample Fréchet variance on the Grassmann manifold to identify “local” principal geodesic submanifolds that represent different system behavior across the parameter space. Polynomial chaos expansion is then used to construct a mapping between the random input parameters and the projection of the response on these local principal geodesic submanifolds. Here, the method is demonstrated on four test cases, a toy-example that involves points on a hypersphere, a Lotka-Volterra dynamical system, a continuous-flow stirred-tank chemical reactor system, and a two-dimensional Rayleigh-Bénard convection problem.

42 ENGINEERING

Large language models for batteries

Large Language Models (LLMs) are advanced artificial intelligence systems capable of solving diverse tasks using language, reasoning, and external tools. Despite their growing deployment in academia and industry, their potential remains underexplored in battery research. This review presents a comprehensive overview of existing and emerging applications of LLMs in batterie field, addressing two critical questions: What can LLMs offer to support battery-related tasks, and how to develop more effective models for this purpose. We begin by outlining the principles of LLMs and criteria for selecting appropriate models and tools for battery research and development. We then explore their roles in text-mining, data interpretation, and the development of intelligent battery systems. In parallel, we discuss technical challenges, such as data standardizing and sharing, model evaluation, and tool integration. Lastly, we propose future research directions with short-, medium-, and long-term goals and highlight more broad perspectives for connecting experts and cross-disciplinary collaborations.

SoC

Foil Bearing Supported Compressor-Expander

R&D Dynamics Corporation (RDD) is developing a fuel cell system compressor-expander designed for heavy-duty vehicle applications. The system emphasizes high reliability, versatility, long service life, and cost-effective mass production. Its design integrates a high-speed centrifugal compressor/expander, oil-free foil air/gas bearings, a direct current permanent magnet (DCPM) motor, and an inverter utilizing silicon carbide (SiC) switches. Key advantages include high efficiency, compact size and reduced weight, enhanced reliability, and elimination of oil contamination risks. The technology targets multiple sectors such as long-haul semi-trucks, mining and construction equipment, power generation, data centers, marine systems, and other heavy machinery. Evaluation units are ready for shipment, and manufacturing capabilities are currently being developed. Contract Number: DE-EE0009617 (2021).

30 DIRECT ENERGY CONVERSION

LIB Design Module for Grid Energy System Application

We will employ a machine learning approach with intelligent data mining and database construction to analyze enormous data repositories for identifying and extracting geographic-dependent cell design specifications from publicly accessible grid-scale energy storage usage databases in an automated way at scale.

Liu, Dianying [Pacific Northwest National Laborato

System Engineers and Decisions: It?s All about Knowledge

In order to guarantee that a system meets adequate levels of reliability and availability, system performances are continuously monitored and analyzed thanks to the technological advancements driving the Industry 4.0 revolution. An Industry 4.0 approach is typically based on advanced statistical, big data mining, machine learning, and internet-of-things methods designed to detect anomalies in the behavior of system, detect the most likely failure modes, and provide indications to system engineers on when maintenance activities should be performed before system performance are deemed unacceptable (which can be generated by diagnostic and prognostic methods). However, these analyses, which are designed to automatize and increase the efficacy of the system maintenance program, require large amount of data which can come in various forms: numeric, textual, images, sounds etc. Such data constitutes the historic knowledge benchmark to track system performances and support system engineer decisions. Here we claim that data is not sufficient to support this kind of analyses when applied to systems characterized by complex architectures and behaviors. Robust system engineer decisions require the ability to understand the system operational context that lies behind the observed data elements. In this respect, system models are in fact necessary to “put data in context” and capture relationships between data elements. Industry 4.0 methods require in fact contextual knowledge as a basis upon which hypotheses can be generated and assumptions tested. In our view, for complex systems, model-based system engineering (MBSE) models can afford this contextual knowledge, as they are typically used to describe systems architecture and dynamic behaviors. System knowledge is here intended as the blending of collected data and system architecture which takes the form of a “knowledge graph”. A knowledge graph is a database which consists of a large set of nodes (in our case an entity can be either a data or an MBSE element) which are linked to each other. The types of nodes and links follow a pre-defined topology, sometimes also refers as an ontology, that is designed to fit the actual decisions that needs to be performed. We show here how a knowledge graph can be defined to support system engineer maintenance decisions and how the same graph can be built based on system MBSE models and pre-processed data from numeric (through anomaly detections and diagnostic methods) and textual elements (through technical language processing TLP).

97 - MATHEMATICS AND COMPUTING

Application of PRIM for understanding patterns in carbon dioxide model-observation differences

Reducing uncertainties in regional carbon balances requires a better understanding of CO 2 transport in synoptic weather systems. Here, we apply the Patient Rule Induction Method (PRIM), a data-mining method to identify high-density regions for a target-class within an input parameter space, to airborne observations of potential temperature, wind speed, water vapor mixing ratio, and CO 2 dry mol fraction gathered during the Atmospheric Carbon and Transport (ACT)-America Summer 2016 and Winter 2017 campaigns. ACT observations were targeted at expert-designated cases of fair weather and near-frontal warm and cold sector air at atmospheric boundary-layer, lower-, and higher free tropospheric levels (ABL, LFT, and HFT, respectively). We investigate atmospheric characteristics of these pre-defined cases and associated CO 2 model-observation-differences in the mesoscale WRF-Chem model. PRIM results separate winter- and summertime observations as well as observations from ABL, LFT, and HFT with enrichment factors of 4.0–20.5 inside the PRIM box compared to the entire dataset but cannot distinguish between near-frontal warm and cold sector observations in the higher free troposphere. Analyzing of the parameter space constrained by PRIM, we find that large magnitude model observation differences preferentially associated with times when atmospheric conditions are less typical. This association suggests that PRIM could provide a useful tool for isolating atmospheric conditions with large-magnitude and non-Gaussian CO 2 -residuals for targeted transport model evaluation and to potentially improve inversion results during synoptically active periods.

Gerken, Tobias [James Madison Univ., Harrisonburg,

Autonomous Synthesis and Inverse Design of Electrochromic Polymers with High Efficiency and Accuracy

Here, the design and synthesis of functional polymers, aimed at targeted properties through specific structures, have long been challenged by their complex and often nonlinear structure–property relationships. Key processes, including knowledge accumulation for predictive design and experimental refinement and validation, are traditionally labor-insensitive and time-consuming, making it difficult to balance accuracy and efficiency. Here, we introduce an accelerated, autonomous system for the on-demand synthesis of electronic polymers that achieves the desired electrochromic functionality with high accuracy and efficiency. Our approach leverages large language model-assisted data mining, a physics-informed copolymer machine learning model, and an AI-driven autonomous robotic workflow in the Polybot lab. Within 72 h, Polybot autonomously synthesized electrochromic polymers (ECPs) with targeted, previously-unreported color values, including green polymers with specific absorption profiles, precisely fine-tuning copolymer structures with a 5% step size in comonomer composition within a three-monomer system. A publicly accessible ECP informatics database has also been created to foster knowledge exchange.

AI-driven Robotic Lab

Rapid Adaptation of Chemical Named Entity Recognition Using Few-Shot Learning and LLM Distillation

Named entity recognition (NER) has been widely used in chemical text mining for the automatic identification and extraction of chemical entities. However, existing chemical NER systems primarily focus on scenarios with abundant training data, requiring significant human effort on annotations. This poses challenges for applications in the chemical field, such as catalysis, where many advancements have traditionally relied on trial-and-error investigations and incremental adjustment of variables. This hinders catalysis science and technology progress in addressing emerging energy and environmental crises. In this work, we propose a few-shot NER model that can quickly adapt to extract new types of chemical entities by using only a limited number of annotated examples. Our model employs a metric-learning approach to transfer entity similarity knowledge from high-resource chemical domains (with abundant annotations) to enable effective entity recognition in low-resource specialized domains (limited annotation). We validate the effectiveness of our model on a few-shot chemical NER benchmark built based on six existing chemical NER data sets. Experiments show that the proposed few-shot NER model can achieve reasonable performance with only 5 examples per entity type and shows consistent improvement as the number of examples increases. Furthermore, we demonstrate how the proposed model can be trained with large language model (LLM) annotated data, opening a new pathway for rapid adaptation of NER systems. Furthermore, our approach leverages the knowledge broadness of large language models for chemistry while distilling this knowledge into a lightweight model suitable for efficient and in-house use.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

NETL Coal Energy Atlas: A Collection of Coal/Energy Related Maps

The NETL Coal Energy Atlas contains a comprehensive collection of coal and energy-related maps and graphics curated by the National Energy Technology Laboratory (NETL) Systems Analysis group. It serves as a living document providing an overview of the U.S. coal and energy sectors. The volume is structurally organized into six key thematic areas. Ultimately, the atlas functions as a modular baseline for data integration, allowing researchers to drill down into specific regional locations or customize geographic base layers for advanced systems analysis.

bituminous coal

Auger@TA: In-situ Cross-Calibration of the World's Largest Cosmic Ray Observatories

The Pierre Auger Observatory (Auger) and the Telescope Array (TA) are the world's two largest ultra-high-energy cosmic ray (UHECR) observatories. They operate in the Southern and Northern hemispheres, respectively, at similar latitudes but with distinct surface detector (SD) designs. A significant challenge in studying UHECR physics across the full sky is the apparent discrepancy in flux measurements between the two experiments. This discrepancy could arise from astrophysical differences and/or systematic effects related to their detector designs and sensitivities to extensive air shower components. To address this, the Auger@TA working group aims to cross-calibrate the two observatories with a self-triggering micro-Auger array within the TA array. This micro-array consists of eight Auger Surface Detector (SD) stations equipped with Water Cherenkov Detectors (WCDs) and AugerPrime Surface Scintillator Detectors. Seven SD stations, configured with a centered-1-PMT design, are arranged in a hexagonal pattern with one station in the center, with 1.5 km spacing, mirroring the Auger layout. The eighth station, which features a standard 3-PMT Auger station, is located in conjunction with a TA detector at the center of the hexagon, forming a triplet for high-statistics and low-uncertainty cross-calibration. A custom communication system that uses readily available components enables seamless communication between stations and remote access to each station through a central computer. The micro-array is now fully deployed, and initial data-taking is about to start. This presentation will detail the instrumentation, communication systems, central data acquisition system, expected performance of the micro-array, and preliminary results as appropriate.

Mocellin, Adriel G.B. [Colorado School of Mines]

Tethys Water Demand Data

U.S. water demand varies sharply by sector and region as land use, population, weather patterns, and economic activity co-evolve. High-resolution water demand data is required to capture these dynamics, support integrated energy-water-land modeling, and local-to-regional water scarcity assessments. This dataset contains gridded (1/8 degree), monthly, multi-sector water demand dataset for the contiguous United States (CONUS) covering 1980-2100 across eight future scenarios of human-Earth system change. The dataset covers irrigation, thermoelectric, municipal (public-supply and domestic), livestock, manufacturing, and mining demands, separately for withdrawals and consumption, and includes per-cell renewable vs. non-renewable water source attributions. The dataset is validated against the latest USGS 2010-2020 water-use data for the three largest water demand sectors (Domestic, Electricity, and Irrigation), with correlations ranging from 0.73-0.95 at the HUC6 scale. The two datasets largely agree on an aggregate basis with per-sector bias falling within +/-7%, but they disagree on the spatial allocation of water with individual HUC6 basins having normalized RMSE from 68-171% and median absolute percent difference from 37-86%. This dataset advances prior global products by combining state-resolved sectoral demands from GCAM-USA, future power-plant siting from the CERF model, and scenario-consistent high-resolution climate and population forcing data across the eight scenarios.

GCAM-USA

Improving Self-Driving Labs: Quantifying System-Level Experiment Repeatability and Broadening Instrument-Level Compatibility

Modular Autonomous Research System (MARS) is a self-driving laboratory (SDL) which performs wet-lab science with peptide-lanthanide combinations in an automated and, ultimately, an autonomous manner to aid in soil analysis for domestic lithium mining. Autonomous experimentation involves automated experimentation, experiment planning, and active learning. MARS consists of a 6-axis robotic arm (UR5e) on a linear rail, pipette robots (Opentrons 2), and microplate readers. These components transport, operate on, and collect data with chemical solutions in standard labware. For effective autonomy, MARS must perform system-level labware operations repeatably, plan experiments autonomously, and be portable between research-domains. Repeatability is evaluated by labware placement precision, such that future operations can properly locate labware, as well as the elapsed time, so that low variance mean estimates of experiment duration can inform high-level researcher decision making. Autonomous experiment planning is the next step to decouple experimentation from human management; however, there is a conflict between the ideal system-level experiment goals and the constraints imposed by instruments’ limitations. Sub-domain portability is a long-term goal to extend MARS’ research beyond the chemistry of peptide-lanthanide binding to other sub-domains without having to invest significant overhead to system retrofitting. To address these goals, we manually trained the robotic arm labware placement and modelled statistical failurerate and uncertainty Additionally, we benchmarked the duration and variance of each experiment sub-operation as a heuristic for research decision making. Next, we use a parameterized geometric program (PGP) approach to design experiments that optimize system-level objectives and satisfy instrument-level constraints. Lastly, we proposed a Python framework to maximize MARS’ extensibility to other scientific sub-domains through a JSON-based experiment specification.

36 MATERIALS SCIENCE

Beneficial Use of Harvested Ponded Fly Ash and Landfilled FGD Materials for High-Volume Surface Mine Reclamation

The overall motivation of this project was to demonstrate at laboratory, bench-scale, and full-scale demonstration levels that (a) coal ash surface impoundments can go through closure by removal as per USEPA and state regulations so that the material can be used as is (other than draining free water using CCRs piles) in high-volume beneficial applications, (b) FGD material from closed out FGD facilities can be excavated and recompacted for coal mine reclamation, and (c) harvested CCRs can be beneficially utilized (providing a net environmental gain) in large-volumes for reclamation at abandoned coal mine sites across the US, especially in the Eastern and Midwest coal mining regions. The objectives of this project were to: 1) promote the safe and cost-effective closure by removal of coal ash impoundments, 2) harvest landfilled FGD, and 3) promote the high-volume beneficial use of these harvested CCRs in the reclamation of abandoned surface coal mine sites across the eastern and midwestern coal mining regions of the United States. The major tasks carried out for this project are summarized below: 1) Conesville Full-Scale Demonstration Project: About 2 million tons of harvested CCR materials from the closure by removal of an inactive fly ash pond and an adjacent old FGD landfill were used for the full-scale demonstration project to fully reclaim a nearby partially completed abandoned surface coal mine. Site monitoring for the project duration was carried out and results are discussed. 2) Laboratory Testing: Geotechnical and environmental testing of harvested ponded fly ash and landfilled FGD material at the former Conesville power plant were carried out. Completing the laboratory testing allowed for QA/QC for the full-scale site construction and informed the formulation of the risk analysis. 3) Risk Analysis: We developed a reliable computational model for fate and transport. We used these models and the rich set of monitored data for the Conesville site to analyze risks to human health and ecological risks associated with high-volume surface mine reclamation using harvested CCRs. 4) GIS Siting Study: A Geographic Information System (GIS) study was carried out for three states in the Eastern coal mining region and two states in the Midwest coal region. This effort provided site specific GIS information for five states and allowed us to establish protocols that other states can follow in implementing their own state specific GIS study.

01 COAL, LIGNITE, AND PEAT

Mauka Energy FEVER Tool DOE SBIR Phase 1 Final Scientific/Technical Report

This report is on the Forestry Electric Vehicle Energy Routing (FEVER) Tool, a novel software system developed to support heavy-duty electric vehicle (EV) operations in remote, forested, and mountainous regions. The Phase I project aimed to demonstrate the feasibility of modeling EV energy consumption using terrain elevation, road conditions, and route features specific to forestry logistics. The tool combines geographic information systems (GIS), electric motor physics, and vehicle-specific data to calculate feasible, energy-efficient routes. Collaborations with Oregon State University’s Research Forests and Titan Freight Systems enabled collection and validation of GPS and elevation-based trip data. The FEVER Tool offers substantial opportunities for the efficient management of medium- and heavy-duty electric vehicles in sectors like forestry, agriculture, mining, defense and waste management—areas which are beginning to adopt HDEVs. The project demonstrated technical feasibility and lays the groundwork for commercial development and deployment in other industries and environmental conditions in Phase II.

29 ENERGY PLANNING, POLICY, AND ECONOMY

Toward Drilling the Perfect Geothermal Well: An International Research Coordination Network for Geothermal Drilling Optimization Supported by Deep Machine Learning and Cloud Based Data Aggregation

The EDGE project, supported by the U.S. Department of Energy Geothermal Technologies Office under award DE-EE0008793, established a data-driven framework for improving the efficiency, cost-effectiveness, and reliability of geothermal well drilling. The project focused on developing scalable data infrastructure, advanced machine learning and probabilistic models, and integrated analytics tools to support continuous drilling optimization. A central objective was to reduce geothermal drilling costs by up to seventy percent while minimizing the risk of well failure through predictive diagnostics and adaptive planning. Over the project period, a comprehensive data repository was designed and deployed, incorporating records from over one hundred geothermal wells across varied geological settings. This repository supported both structured and unstructured data and adhered to FAIR data principles, enabling provenance tracking, quality control, and standardized metadata. The project introduced automated ingestion pipelines and a cloud-hosted platform that facilitated access to raw, processed, and derived datasets. This infrastructure served as the foundation for model development and analysis. Machine learning workflows were developed to predict key drilling metrics including rate of penetration, non-productive time, and total drilling costs. Self-organizing maps and dimensionality reduction methods were used to uncover operational patterns and outliers, while supervised learning algorithms such as random forests and deep neural networks were applied to forecast performance outcomes. The models were validated on heterogeneous datasets from both U.S. and Icelandic fields, demonstrating variable but significant predictive accuracy. The results indicated that finer temporal resolution, inclusion of lithological data, and consistency in operational annotations could substantially improve model performance. The project also implemented process mining techniques to reconstruct state-transition models from drilling event logs. These models enabled the identification of deviations from optimal workflows and provided insights into recurring failure modes. Analysis of non-productive time highlighted the impact of equipment failures, geological challenges, and human factors, offering opportunities for targeted mitigation strategies. The EDGE Dashboard was developed as a web-based expert system integrating data visualization, model outputs, and user-driven queries. It provided an accessible interface for operators to explore historical data, evaluate predicted outcomes, and compare drilling scenarios. Initial feedback from project partners suggested that the dashboard could serve as a foundation for more advanced advisory and optimization tools. Overall, the EDGE project demonstrated the feasibility and value of applying modern data science techniques to geothermal drilling. It delivered a set of interoperable tools and models that can support more efficient, lower-risk well development. The findings point toward a viable path for transitioning from advisory analytics to semi-autonomous drilling systems, contingent on continued collaboration, expanded datasets, and field validation. The project results have immediate relevance for drilling operations, data management practices, and future geothermal R&D efforts aimed at achieving reliable, cost-competitive geothermal energy at scale.

15 GEOTHERMAL ENERGY

North America’s Potential for an Environmentally Sustainable Nickel, Manganese, and Cobalt Battery Value Chain

The Detroit Big Three General Motors (GMs), Ford, and Stellantis predict that electric vehicle (EV) sales will comprise 40–50% of the annual vehicle sales by 2030. Among the key components of LIBs, the LiNixMnyCo1−x−yO2 cathode, which comprises nickel, manganese, and cobalt (NMC) in various stoichiometric ratios, is widely used in EV batteries. This review reveals NMC cathodes from laboratory research. Furthermore, this study examines the environmental effect of NMC cathode production for EV batteries (including coating technologies), encompassing aspects such as energy consumption, water usage, and air emissions. Although gaps persist in NMC cathode environmental assessments (NMC111, NMC532, NMC622, and NMC811), limited life cycle assessments “(LCA)” have been conducted. Most available data originate from Asia (primarily China), accounting for 85% of the production of EV LIB cathode materials. The concept of battery passports for data collection on LIB components has been proposed to facilitate material traceability as a system for ensuring a sustainable supply chain for critical minerals. The automotive industry’s shift to electrification necessitates a sustainable supply chain from mine to vehicle end-of-life. As the critical mineral supply moves from Asia to North America, environmentally friendly industrial methods must be studied to provide this supply chain direction.

25 ENERGY STORAGE