Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Feature engineering”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Phase Picking Beyond Local Distances: Where Waveform Filtering Still Matters for Deep Learning Models

Waveform filtering is a standard step in traditional seismic phase picking but often receives little attention in deep learning workflows, where models are typically trained on raw or minimally processed waveforms. Although this strategy performs well for local events, we show that performance can degrade substantially at regional distances. To address this limitation, we introduce two ways to incorporate multiband-filtered waveforms into deep learning phase pickers. The stacking approach concatenates filtered inputs along the channel dimension, while the branching approach processes each frequency band through a dedicated network branch before feature fusion. Both approaches can substantially improve performance across epicentral distances of 0° to 20°, but their effectiveness depends strongly on the selected frequency bands. Tests with multiple filter banks show that filter-bank design should be treated as part of model optimization rather than as a fixed preprocessing choice. Grad-CAM analysis of the branching model indicates that band importance varies among waveform samples and across training realizations, with only a weak overall preference for the 0.25 to 0.5 Hz band. These results show that no single filter band is consistently optimal and demonstrate that explicit feature engineering remains valuable for robust deep learning-based seismic phase picking.

58 GEOSCIENCES

A Semi-supervised Hybrid Machine Learning Framework for the Qualification of Resistance Spot Welds

• Industries requiring high structural integrity, including automotive, aerospace, and construction, place considerable significance on weld quality classification. • The inspection normally involves human expertise through predefined quality metrics that are subjective, error-prone, and time-intensive • The challenge to classification model development is the scarcity of labeled data and imbalanced distributions in the data that are labeled. • This work develops a new hybrid methodology that achieves clustering using KMeans++ together with supervised classification to overcome these challenges. • The ensemble-based classifiers were identified as optimal, with accuracy enhancements of up to 8% using the pseudo-labeled dataset. • The work provides practical insight into feature engineering and machine learning integration in industrial quality assurance applications.

Rogers, Jeremy K. [Savannah River National Laborat

Unsupervised Detection of SOC Spoofing in OCPP 2.0.1 EV Charging Communication Protocol Using One-Class SVM

The electric vehicles (EVs) market keeps growing globally; thus, it is critical to secure the EV charging communication protocols in order to guarantee reliable and fair charging operations among the customers. The Open Charge Point Protocol (OCPP) 2.0.1 supports the communication between the Electric Vehicle Supply Equipment (EVSE) and Charging Station Management Systems (CSMSs); therefore, it becomes vulnerable to several types of attacks, which aim to jeopardize smart charging, billing, and energy management. Specifically, OCPP 2.0.1 allows the self-reporting of the State of Charge (SOC) values, which makes it vulnerable to spoofing-based cyberattacks, which target manipulating the scheduling priorities, distorting the load forecasts, and extending the charging sessions in an unfair manner. In this paper, we try to address this type of attack by providing a comprehensive analysis of the SOC spoofing attacks and introducing a novel unsupervised detection framework based on the One-Class Support Vector Machine (OCSVM) algorithm. Specifically, two types of attack scenarios are analyzed (i.e., priority manipulation and session extension) by deriving engineered features that capture the nonlinear relationships under normal charging behavior. Detailed simulation-based results are derived by utilizing the DESL-EPFL Level 3 EV charging dataset. Our results demonstrate high F1-score and recall in identifying spoofed SOC values and that the proposed OCSVM model demonstrates superior performance compared to alternative clustering and deep-learning based detectors.

EV charging

Recent Progress on Surface Water Quality Models Utilizing Machine Learning Techniques

Surface waterbodies are heavily exposed to pollutants caused by natural disasters and human activities. Empowering sensor technologies in water quality monitoring, sufficient measurements have become available to develop machine learning (ML) models. Numerous ML models have quickly been adopted to predict water quality indicators in various surface waterbodies. This paper reviews 78 recent articles from 2022 to October 2024, categorizing water quality models utilizing ML into three groups: Point-to-Point (P2P), which estimates the current target value based on other measurements at the same time point; Sequence-to-Point (S2P), which utilizes previous time series data to predict the target value at one time point ahead; and Sequence-to-Sequence (S2S), which uses previous time series data to forecast sequential target values in the future. The ML models used in each group are classified and compared according to water quality indicators, data availability, and model performance. Widely used strategies for improving performance, including feature engineering, hyperparameter tuning, and transfer learning, are recognized and described to enhance model effectiveness. The interpretability limitations of ML applications are discussed. This review provides a perspective on emerging ML for surface water quality models.

machine learning (ML)

Tandem Predictions for HPC Jobs: Preprint

At the core of the predictive analytics applied to High Performance Computing (HPC), the most prominent tasks are the prediction of job runtimes and the prediction of job queue times, both of which have the potential for informing HPC users during their every-day decision making. Accurate runtime predictions can help users better choose so-called wallclock times at job submission, decreasing the odds of their jobs waiting in queues longer than necessary. The accurate and timely queue time predictions offered for the available partitions can inform the favorable selection of partitions for running jobs. This potential is well understood as we see in the abundance of research studies that propose solutions for these tasks, including the work published in the last several years. These tasks are seemingly receptive to the Machine Learning (ML) solutions, considering that there is no shortage of training data where HPC centers over time run millions and millions of jobs. However, we study the existing research literature, as well as look for examples in the toolchains supported on the exemplar HPC facilities, and, surprisingly, do not find any practical solutions that are ready to be adopted. We interpret this as a manifestation of the shortage of UX/UI efforts that support HPC analytics and also as a sign that the research has not come to the consensus on solving these tasks. In this study, we aim to shed new light on the long-running task of job queue time prediction by exploring the utility of runtime predictions in improving prediction accuracy and, actually, predicting these two metrics together, in tandem. In other words, we show how runtime predictions become valuable input in the queue time modeling. We challenge the existing approaches to feature engineering for the queue time prediction and describe promising results we obtained for a large dataset of HPC jobs from a supercomputer at the National Renewable Energy Laboratory.

97 MATHEMATICS AND COMPUTING

Achieving Unprecedented CO 2 Utilization InCO 2 Concrete™: System Design, Product Development and Process Demonstration

Anthropogenic sources of carbon dioxide are generated from a number of sources, but the key among these are ordinary Portland cement (OPC) production and combustion of fossil fuels. Cement production is the largest global CO 2 source from the mineral decomposition of carbonates. This is due to the clinkering process whereby limestone (mainly consisting of CaCO 3 ) is decomposed into CaO and CO 2 , and combined with silica rich clays at high temperatures to form clinkers (i.e. the four key minerals that comprise cement). The high temperature range of 1400 – 1550°C required for this process accounts for up to 60% of the generated CO 2 from cement production. Combination of the limestone decomposition and thermal requirements of the clinkering process causes cement production to contribute 8-9% of annual global CO 2 emissions. Combustion of fossil fuels (coal, oil and gas) was shown to contribute a much larger portion of global CO 2 emissions. As of 2018, combustion of fossil fuels accounted for 65% of global CO 2 , where 41% was derived from stationary sources for electricity and heat generation and the other 24% was related to transport. To reduce these contributions, key steps forward in CO 2 utilization technologies are required. Therefore, a CO 2 mineralization technology (CO 2 mineralization concrete) to reduce the OPC content in concrete, while utilizing flue gas emissions from fossil fuel combustion has been developed to address both areas simultaneously. This Reversa™ technology utilizes low-carbon cementation agents produced by in situ CO 2 mineralization (“mineral carbonation reactions”) to offer a promising alternative to OPC. CO 2 mineralization relies upon the reaction of dissolved CO 2 with inorganic alkaline reactants to precipitate mineral carbonates (e.g., CaCO 3 ), which bind proximate particles and achieve cementation. Herein, a concrete green body, which is composed of a mixture of binder, water, and mineral aggregates, is exposed to CO 2 borne in industrial flue gas streams. This manner of CO 2 mineralization allows the production of construction components that feature equivalent engineering attributes as their OPC-based counterparts while featuring a much smaller embodied carbon intensity (eCI). The purpose of this project is to demonstrate the feasibility of the Reversa process evolving from a TRL-3 technology at the bench-scale up to TRL-6 technology at the pilot-scale. The reliability of the Reversa technology was tested to prove the effective production of three standard industrial concrete products selected during the course of the project. The results detailed herein will demonstrate the evolution of this technology to the industrial scale. The culmination of this work resulted in 9 production runs completed at the National Carbon Capture Center (NCCC), Wilsonville, AL, using natural gas (NG) flue gas as the CO 2 source. Over the course of the production runs at NCCC, the CO 2 utilization as a function of time, 24-h CO 2 uptake, electricity usage, and 28-d net area compressive strength recorded for each run. Collection of this data will be used to determine the success of the demonstration goals: (1) achieving in excess of 0.2gCO 2 /g reactant , (2) achieving greater than 50% reduction in global warming potential compared to standard produced units, and (3) ensuring compliance of carbonated concrete with industry standard specifications.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

CFD modeling of non-catalytic, partial-oxidation engine reformer for flare mitigation

Flaring associated natural gas is commonly employed in the oil and gas industry to reduce methane (CH 4 ) emissions but generates carbon dioxide (CO 2 ) and harmful pollutants, significantly contributing to air pollution and posing risks to public health. To mitigate this impact, M2X Energy Inc. has developed a small-scale, modular gas-to-methanol system. This system features an engine reformer that performs fuel-rich partial oxidation of wellhead gas to produce syngas—a mixture of carbon monoxide (CO) and hydrogen (H 2 )—followed by a downstream reactor for methanol synthesis. This study focused on computational fluid dynamics (CFD) modeling of the engine reformer to simulate partial oxidation chemistry, predict the rich-burn operating limit, and assess syngas quality, ultimately aiding in design and operational optimization. The CFD model, developed within a Reynolds-Averaged Navier-Stokes (RANS) turbulence framework, incorporated sub-models for turbulent combustion, a chemical mechanism with polycyclic aromatic hydrocarbon (PAH) pathways, and soot emissions to accurately capture the fuel-rich, turbulent jet ignition and combustion processes. Model validation against experimental data showed good agreement across pre- and main-chamber pressures, apparent heat release rates, and exhaust gas concentrations of key species (H 2 , CO, CO 2 , CH 4 ) for varying intake equivalence ratios. Here, the model identified a rich-burn operating limit near a fuel-air equivalence ratio of 2.35, consistent with experimental observations. Furthermore, syngas quality analysis revealed that extending the rich-burn limit through engine reformer optimization could enhance syngas production, contributing to higher methanol synthesis efficiency.

Computational Fluid Dynamics

Feature issue introduction: laser driven inertial confinement fusion and bridging the gaps to inertial fusion energy systems

Major fusion research milestones have been achieved using laser driven inertial confinement fusion (ICF) in recent years, and these successes have ignited tremendous enthusiasm for inertial fusion energy (IFE). However, the complexity and difficulty of obtaining fusion ignition with a laser driver in a research setting are often underappreciated, as are the gaps to high driver efficiency, high repetition rates, and laser and target durability requirements needs for IFE. On the academic side, several new research laser systems have been constructed over the past few years, enabling researchers to probe the limits of ICF physics and engineering. This feature issue highlights the challenges and capabilities of laser research and development targeted towards advancing IFE.

Physics - Plasma physics

A systematic review of machine learning in groundwater monitoring

With increasing concerns about water scarcity, groundwater has become crucial since this resource provides most of the freshwater needs. However, various human and natural activities often contaminate the groundwater, making it unsuitable for use. Over the years, scientists and engineers have used many methods to predict and track groundwater contamination as part of environmental monitoring. Consequently, there is an urgent need for improved methods, particularly in the face of increasing contamination. Machine learning has sometimes been used to monitor groundwater, air quality, and climate. Traditional methods must be improved due to the complexity and large amount of environmental data. This includes using hybrid models that combine traditional and new techniques. Despite the use of machine learning in many scientific areas, there is a lack of comprehensive reviews focusing on its use in environmental monitoring, especially groundwater monitoring. We aim to fill this gap by exploring machine-learning applications in groundwater monitoring. We discuss relevant methods, their limitations, and future potential. We summarize research on automating data processing and model training using groundwater sensor data. Our research underscores the transformative potential of machine learning to revolutionize long-term groundwater monitoring and contamination detection, providing valuable insights for future research and practical applications.

AI/ML

Systematic feature design for cycle life prediction of lithium-ion batteries during formation

Optimization of the formation step in lithium-ion battery manufacturing is challenging due to limited physical understanding of solid-electrolyte interphase formation and the long testing time (∼100 days) for cells to reach the end of life. We propose a systematic feature-design framework that requires minimal domain knowledge for accurate cycle life prediction during formation. By only using two simple Q (V) features designed from our framework, extracted from formation data without any additional diagnostic cycles, we achieved an average of 9.87% error for cycle life prediction. Here, the physics-based investigation guided by the two designed features shows that the voltage ranges identified by our framework capture the effects of formation temperature and microscopic-particle resistance heterogeneity. By designing highly predictive, robust, and interpretable features, our approach can accelerate industrial battery formation research, leveraging the interplay between data-driven feature design and mechanistic understanding.

25 ENERGY STORAGE

Comparison of Expert Vocabulary Usage Patterns Between Mental Health and Nonmental Health Clinicians When Diagnosing Pediatric Anxiety Disorders

Objective: To compare the utilization patterns of expert vocabulary (EVo) in diagnosing pediatric anxiety between mental health and non-mental health clinical notes from electronic health records to understand the role of Evo in informing classification and decision-making in anxiety diagnoses. Study design: We conducted a retrospective study using a cohort less than age 25 from Cincinnati Children's Hospital including 897 685 patients with 61 586 446 notes. We analyzed EVo, collected from mental health clinicians, in both mental and nonmental health notes. We compared classification accuracy using EVo-based patient-level embedding from all clinical notes, mental-health notes, and nonmental health notes for 2 tasks: 1) pre-vs postdiagnosis anxiety patients, and 2) prediagnosis anxiety vs nonanxiety patients. Results: EVo usage was highest in prediagnosis anxiety, lower in nonanxiety, and lowest in post-diagnosis. Classification models using EVo features from all, mental-health, and non-mental health notes showed similar F1 scores for prediagnosis anxiety (0.70 ± 0.2 for 2 categories). For anxiety vs nonanxiety classification, all clinical and nonmental health notes had better F1 scores than mental-health notes (above 0.90 for 3 categories). There was a notable difference in class-wise performance across both tasks. Conclusions: There are significant differences in anxiety EVo use between mental health and nonmental health clinicians. Despite less anxiety-specific terminology, non-mental health notes still captured key aspects of patient presentations, emphasizing the importance of including all clinicians' notes in analysis. EVo's utility for anxiety classification is most effective in prediagnostic phases, suggesting the need for a dedicated diagnostic lexicon and further study before incorporating EVo into classification models.

feature engineering

Monitoring Plan for the Idaho National Laboratory Remote Handled Low Level Waste Disposal Facility

This monitoring plan for Idaho National Laboratory’s Remote-Handled Low Level Waste Disposal Facility was developed to meet the requirements for monitoring low-level waste disposal facilities according to the U.S. Department of Energy (DOE) Order 435.1, “Radioactive Waste Management,” and the guidance provided in the associated technical standard “Disposal Authorization Statement and Tank Closure Documentation” (DOE-STD-5002-2017). The purpose of this monitoring plan is to document a monitoring strategy that includes (1) compliance monitoring activities to demonstrate compliance with regulatory standards/limits and (2) performance monitoring to build confidence the facility is performing as demonstrated in the facility performance assessment (PA) (DOE-ID 2018a), composite analysis (CA) (DOE ID 2012), and CA addendum (DOE-ID 2018b). The de minimus impact to the aquifer predicted by the PA suggests that aquifer compliance monitoring should be augmented with performance monitoring of the drainage course materials and sedimentary interbeds in the vadose zone beneath the facility to provide a more effective means of identifying performance deviations. The monitoring approach delineated in this document was informed by the systems evaluation of natural and engineered facility features presented in the PA, an assessment of aquifer baseline conditions (INL 2017d), the dose analysis conducted in support of the PA and CA, and monitoring data collected during the first four years of facility operations (baseline monitoring phase) (INL 2023b). This plan provides monitoring locations, sampling frequencies, and sampling methods; recommendations for data evaluation; and a description of the monitoring plan implementation. Collected data will be used to demonstrate facility compliance and to identify conditions that are not consistent with the key assumptions made by the PA and CA.

12 - MGMT OF RADIOACTIVE AND NON-RADIOACTIVE WASTE

Hydropower Infrastructure - LAkes, Reservoirs, and RIvers (HILARRI), v4

HILARRI is a database of links between major datasets of operational hydropower dams and powerplants, and inland water bodies. These connections are critical for conducting large-scale analysis of hydropower infrastructure and their associated natural and engineered water systems. Features include: – Dams from the National Inventory of Dams (2025) and the Global Reservoir and Dam Database (GRanD v1.3) – Hydropower plants from the Existing Hydropower Assets dataset (EHA 2025) – Power plants that are listed in the 2025 U.S. Hydropower Development Pipeline Data or were listed in previous versions of the dataset These hydropower infrastructure features are linked to several major datasets that provide hydrologic and hydraulic information relevant for analysis of hydropower systems that includes the integral water resources. That information comes from: – Products from the National Hydrography Dataset (NHD) – NHDPlusV2 Medium Resolution river network flowlines, – NHD waterbodies (limited to lakes and reservoirs), – NHD Watershed Boundary Dataset (HUC12-level for the Conterminous United States (CONUS)) – NHD High Resolution waterbodies – HydroLAKES water bodies (lakes and reservoirs) – LAGOS-US lakes and reservoirs – EPA National Lakes Assessment (2007, 2012, 2017, and 2022) – The Reservoir Sedimentation Database (RESSED) – EPA SuRGE sampling locations Unique identifiers are used to facilitate joining to the original full datasets. For example, characteristics of NHD flowlines such as estimated average flow rate can be joined from the NHDPlusV2 dataset to a dam or power plant listed in HILARRI based on the ID field, “COMID”, that is common to both datasets. HILARRI only includes basic information about identifiers, location, and data quality or usage notes. It does not contain the attributes or time series data associated with these sites. The HILARRI dataset incorporates information from several datasets to facilitate more effective and accurate analysis of hydropower infrastructure and their associated waterbodies. For example, dams were checked against the most recent American Rivers Dam Removal Database to identify and flag facilities that may no longer exist. Additionally, dams that are listed multiple times in the NID are identified and flagged to avoid double-counting when analyzing and summarizing information. Other quality flags include certainty of operational hydropower (i.e., if one or more datasets indicates hydropower at a particular location), whether an associated water body is accurate or composed of multiple polygons, or whether there is a known issue with reported characteristics in one of the underlying datasets. These additional data flags are designed to increase confidence in data usage for individual to large-scale analyses.

Hansen, Carly [ORNL] (ORCID:0000000193280838)

Using Explainable Artificial Intelligence to Predict Perovskite Solar Cell Electrical Metastability from Operando Photoluminescence Images in Accelerated Stress Testing

Metal halide perovskite (MHP) solar cells exhibit a metastable response to bias governed by coupled ionic–electronic processes, complicating the conventional reciprocity relation between luminescence intensity and device open-circuit voltage (V oc ). This limits the use of luminescence as a diagnostic for device screening or accelerated stress testing, motivating new approaches that can interpret photoluminescence (PL) signals under nonequilibrium conditions. From the artificial intelligence perspective, we develop an explainable deep learning framework that integrates convolutional neural networks (CNN), long short-term memory (LSTM) layers, and an attention mechanism to learn spatiotemporal features from operando photoluminescence PL image sequences. The model achieves a mean absolute error of ±0.027 V in predicting open-circuit voltage transients and reduces extreme-tail errors by up to 78% compared to physics-based reciprocity calculations. Gradient-weighted Class Activation Mapping (Grad-CAM) provides interpretability by highlighting physically meaningful regions such as electrode edges and emergent defect features. From the engineering application perspective, this framework enables accurate, contactless prediction of device V oc and identification of degradation-relevant features during accelerated aging of perovskite solar cells. This approach demonstrates how explainable AI can enhance operando diagnostics and reliability analysis in photovoltaic devices under nonequilibrium conditions.

14 SOLAR ENERGY

RMCProfile7 : reverse Monte Carlo for multiphase systems

This work introduces a completely rewritten version of the programRMCProfile(version 7), big-box, reverse Monte Carlo modelling software for analysis of total scattering data. The major new feature ofRMCProfile7is the ability to refine multiple phases simultaneously, which is relevant for many current research areas such as energy materials, catalysis and engineering. Other new features include improved support for molecular potentials and rigid-body refinements, as well as multiple different data sets. An empirical resolution correction and calculation of the pair distribution function as a back-Fourier transform are now also available.RMCProfile7is freely available for download at https://rmcprofile.ornl.gov/.

Chemistry

An Innovative High Throughput Genome Releaser for Rapid and Efficient PCR Screening

High-throughput PCR screening is vital in synthetic biology and metabolic engineering as it allows researchers to rapidly analyze and detect numerous targeted genetic mutation in the genome. Current challenges for high-throughput PCR screening in synthetic biology include efficiently preparing genomic DNA, optimizing protocols for diverse sample types, managing contamination risks, and effectively analyzing the large volumes of data generated while ensuring consistent and accurate results. In this study, we present the development of a High Throughput Genome Releaser (HTGR), an innovative device addressing common challenges in screening PCR. This genome DNA releaser is designed based on a squash method for rapid, cost-effective, and efficient DNA release, optimized for subsequent PCR reactions. After experimenting with various synthetic materials, we selected a plastic that closely replicates the smooth surface and compression properties of microscope slides, ensuring reliable performance. We engineered a device featuring a 96-Well Plate and a shear applicator, operable both manually and automatically, and compatible with standard liquid-handling robot platform. This compatibility enhances ease of use in high-throughput PCR workflows. Additionally, we developed software to support its automatic functions. Our results demonstrated that the specially engineered 96-Well Plate and HTGR can effectively squash fungal spores , which release enough genome DNA for PCR screening. The genome releaser facilitates the preparation of PCR-amplifiable genomic DNA substrate from 96 samples within minutes, eliminates the need for extraction buffers, and is adaptable to a wide range of microorganisms and cells, which could significantly advance biomanufacturing processes.

Yuan, Guoliang [BATTELLE (PACIFIC NW LAB)]

The spherical tokamak advanced reactor (STAR) fusion power plant design

Scientific and technical advancements have been made that improve fusion’s prospects to provide a new energy source, showing enhanced plasma confinement conditions with plasma temperatures reaching or exceeding 100 million degrees. Overshadowing this progress is the challenge involved in developing an economically viable fusion power plant design. Many proposed next-step DEMO and pilot plant designs are extensions of existing physics-focused experimental devices defined to understand and control plasma operations to achieve and sustain a fusion reaction. Transitioning scientific and technical advancements into a functional power plant requires a dedicated focus on architectural designs that integrate diverse technologies, while optimizing physics conditions, with a focus on economic viability. This holistic approach is essential in turning the promise of fusion energy into a reality. The Spherical Tokamak Advanced Reactor (STAR) is a fusion power plant conceptual design with the architectural focus that strives to balance physics, engineering, and cost considerations. In conclusion, it has been set up to introduce relevant physics, engineering and concept features that an intermediate pilot plant might follow, with the goal of meeting system performances and economic requirements that lead to a commercially competitive fusion power plant.

Blanket segmentation