Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “preprocessed data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Model scripts associated with “Revisiting controls on hyporheic respiration with knowledge-guided machine learning at continental scale”

NOTE: The manuscript associated with this data package is currently in review. The data/scripts may be revised based on reviewer feedback. Upon manuscript acceptance, this data package will be updated with the final scripts and additional metadata. This data package is associated with the publication “Revisiting controls on hyporheic respiration with knowledge-guided machine learning at continental scale” submitted to Environmental Science & Technology (Zheng et al. 2026). The project combines mechanistic process modeling with knowledge-guided machine learning (KGML) to evaluate how organic matter chemistry, microbial biomass, and physical substrate accessibility regulate realized respiration rates across river corridors. All data used in this paper have been previously published and can be accessed at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1729719 (Goldman et al., 2020). This data package contains 3 R-markdown (Rmd) preprocessing scripts for the previously published data and subsequent modelling workflows. The full workflow with input and output data can be found in the associated GitHub repository at https://github.com/jianqiuz/KGML-WHONDRS.

Biogeochemistry↗

Using soil library hyperspectral reflectance and machine learning to predict soil organic carbon: Assessing potential of airborne and spaceborne optical soil sensing

Soil organic carbon (SOC) is a key variable to determine soil functioning, ecosystem services, and global carbon cycles. Spectroscopy, particularly optical hyperspectral reflectance coupled with machine learning, can provide rapid, efficient, and cost-effective quantification of SOC. However, how to exploit soil hyperspectral reflectance to predict SOC concentration, and the potential performance of airborne and satellite data for predicting surface SOC at large scales remain relatively underknown. Here, this study utilized a continental-scale soil laboratory spectral library (37,540 full-pedon 350–2500 nm reflectance spectra with SOC concentration of 0–780 g·kg –1 across the US) to thoroughly evaluate seven machine learning algorithms including Partial-Least Squares Regression (PLSR), Random Forest (RF), K-Nearest Neighbors (KNN), Ridge, Artificial Neural Networks (ANN), Convolutional Neural Networks (CNN), and Long Short-Term Memory (LSTM) along with four preprocessed spectra, i.e. original, vector normalization, continuum removal, and first-order derivative, to quantify SOC concentration. Furthermore, by using the coupled soil-vegetation-atmosphere radiative transfer model, we simulated twelve airborne and spaceborne hyper/multi-spectral remote sensing data from surface bare soil laboratory spectra to evaluate their potential for estimating SOC concentration of surface bare soils. Results show that LSTM achieved best predictive performance of quantifying SOC concentration for the whole data sets (R 2 = 0.96, RMSE = 30.81 g·kg –1 ), mineral soils (SOC ≤ 120 g·kg –1 , R 2 = 0.71, RMSE = 10.60 g·kg –1 ), and organic soils (SOC > 120 g·kg –1 , R 2 = 0.78, RMSE = 62.31 g·kg –1 ). Spectral data preprocessing, particularly the first-order derivative, improved the performance of PLSR, RF, Ridge, KNN, and ANN, but not LSTM or CNN. We found that the SOC models of mineral and organic soils should be distinguished given their distinct spectral signatures. Finally, we identified that the shortwave infrared is vital for airborne and spaceborne hyperspectral sensors to monitor surface SOC. This study highlights the high accuracy of LSTM with hyperspectral/multispectral data to mitigate a certain level of noise (soil moisture <0.4 m 3 ·m –3 , green leaf area < 0.3 m 2 ·m –2 , plant residue <0.4 m 2 ·m –2 ) for quantifying surface SOC concentration. Forthcoming satellite hyperspectral missions like Surface Biology and Geology (SBG) have a high potential for future global soil carbon monitoring, while high-resolution satellite multispectral fusion data can be an alternative.

54 ENVIRONMENTAL SCIENCES↗

Fault Detection Utilizing Convolution Neural Network on Timeseries Synchrophasor Data From Phasor Measurement Units

An end-to-end supervised learning method is proposed for fault detection in the electric grid using Big Data from multiple Phasor Measurement Units (PMUs). The approach consists of preprocessing steps aimed at reducing data noise and dimensionality, followed by utilization of six classification models considered for detecting faults. Three of the models were variants of Convolutional Neural Network (CNN) architectures that consider a single type of measurement (voltage, current or frequency) at all PMUs or all types together also at all PMUs. CNN based models were compared to traditional methods of Logistic Regression (LR), Multi-layer Perceptron (MLP) and Support Vector Machine (SVM). Evaluation was conducted on two-year data measured by PMUs at 37 locations in a large electric grid. Here, the response variable for classification were extracted from the grid-wide outage event log. Experiments show that CNN-based models outperformed traditional methods on one year out-of-sample outage detection over the entire grid.

42 ENGINEERING↗

Salton Sea Geodatabase for Geothermal Exploration Artificial Intelligence

These files contain the geodatabases related to Salton Sea Geothermal Field. It includes all input and output files used with the Geothermal Exploration Artificial Intelligence. Input and output files are sorted into three categories: raw data, pre-processed data, and analysis (post-processed data). In each of these categories there are six additional types of raster catalogs which are titled Radar, SWIR, Thermal, Geophysics, Geology, and Wells. The files are used with the Geothermal Exploration Artificial Intelligence for the Salton Sea Geothermal Site to identify indicators of blind geothermal systems. The included zip file is a geodatabase to be used with ArcGIS and the tar file is an inclusive database that encompasses the inputs and outputs for the Salton Sea Geothermal Site.

15 GEOTHERMAL ENERGY↗

Desert Peak Geodatabase for Geothermal Exploration Artificial Intelligence

These files contain the geodatabases related to the Desert Peak Geothermal Field. It includes all input and output files used in the project. The files include data categories of raw data, pre-processed data, and analysis (post-processed data). In each of these categories there are six additional types of raster catalogs including Radar, SWIR, Thermal, Geophysics, Geology, and Wells. The files for the Desert Peak Geothermal Site are used with the Geothermal Exploration Artificial Intelligence to identify indicators of blind geothermal systems. The included zip file is a geodatabase to be used with ArcGIS and the tar file is an inclusive database that encompasses the inputs and outputs for the Desert Peak Geothermal Field.

15 GEOTHERMAL ENERGY↗

Brady Geodatabase for Geothermal Exploration Artificial Intelligence

These files contain the geodatabases related to Brady's Geothermal Field. It includes all input and output files for the Geothermal Exploration Artificial Intelligence. Input and output files are sorted into three categories: raw data, pre-processed data, and analysis (post-processed data). In each of these categories there are six additional types of raster catalogs which are titled Radar, SWIR, Thermal, Geophysics, Geology, and Wells. These inputs and outputs were used with the Geothermal Exploration Artificial Intelligence to identify indicators of blind geothermal systems at the Brady Hot Springs Geothermal Site. The included zip file is a geodatabase to be used with ArcGIS and the tar file is an inclusive database that encompasses the inputs and outputs for the Brady Hot Springs Geothermal Site.

15 GEOTHERMAL ENERGY↗

Survey of Time Shift Detection Algorithms for Measured PV Data

In this research, three variations of time shift detection algorithms were tested for their ability to detect time shift issues (including daylight savings time and random time shifts) in measured PV data sets. Two algorithms from the Python PVAnalytics package were assessed, and one algorithm from the Solar-Data-Tools package was assessed. Each algorithm's ability to accurately detect and measure time shifts was assessed.

automated preprocessing↗

A Data-Driven Framework for Predicting the Sorting and Screening Performance of an Integrated Biomass Feedstock Preprocessing System

The characteristics of mechanically sorted and screened lignocellulosic biomass, such as the mass contents of corn stover anatomical fractions (leaves, husks, stalks, cobs, etc.), can be used to calculate the intermediate feedstock quality attributes “yield” and “purity” that indicate the conversion efficiency of biocrude. No prior study has investigated the correlations from the characteristics of raw biomass and preprocessing unit operation parameters to those intermediate feedstock quality attributes. This work presents a data-driven framework for assessing and predicting the intermediate feedstock quality attributes in an integrated biomass feedstock preprocessing system. Our study used corn stover as a typical type of herbaceous biomass because of its abundance in the U.S. It began with data acquisition of moisture content, particle size distribution, and anatomical fractions of the materials after each unit operation in the system. The objective of this preprocessing system is to minimize husks and leaves and maximizing cobs and stalks by mechanically separating the materials into three streams via disc screen and air separator. Prototype neural network models were then developed to evaluate the feasibility of predicting process outcomes based on measurable parameters. It is found that incorporating physical constraints into these prediction models significantly enhances the accuracy of the predicted yield and purity against the ground truth data. The experimental data and model predictions indicate that decreasing throughput increases purity, while higher throughput results in lower purity. Finally, an optimization problem was introduced to search optimal combinations of feed material properties and preprocessing unit operation parameters, as the intermediate feedstock quality attributes – yield and purity, appeared to be competing factors. The study also suggests the continual need to improve the data-driven framework’s predictability by incorporating more accurate physical models to describe the dynamics in the preprocessing units such as the air separator.

09 - BIOMASS FUELS↗

Day-Ahead Probabilistic Forecasting of Net-Load and Demand Response Potentials with High Penetration of Behind-the-Meter Solar-plus-Storage

The goal of this project is to develop advanced methods for day-ahead net-load forecasting, by leveraging the state-of-the-art machine learning techniques. The developed models produce both point and probabilistic forecasts for a variety of use cases, and are versatile to work with different types of data sets. The innovation lies in the novel design of the architectures, leveraging the most recent advances in machine learning that have not been explored in power systems, accompanied by techniques in the broader artificial intelligence fields such as fuzzy systems. This project has achieved the following accomplishments: (1) preprocessing of over 10 data sets covering varying geographical regions, time horizons, and system levels, which form a robust foundation for training and evaluating forecasting models across a wide range of realistic grid scenarios; (2) development of an interactive web app that enables exploratory analysis of load and generation data, and supports better understanding of data trends, anomalies, and correlations, facilitating model development and stakeholder engagement; (3) implementation of over 10 benchmark models for point and probabilistic forecasting, which include a mix of conventional machine learning methods and state-of-the-art deep learning approaches, providing a comprehensive baseline for performance comparison and validation of the proposed models; (4) development of a fuzzy system based gradient boosting model, tailored for small (less than 3 years) data sets, which achieves a mean absolute percentage error (MAPE) of 4% for point forecasting and a 20% improvement in average pinball loss for probabilistic forecasting; (5) development of a Transformer (a state-of-the-art deep learning architecture) based neural network model, tailored for large (3 years or more) data sets, which achieves a MAPE of 2% for point forecasting and a 20% improvement in average pinball loss for probabilistic forecasting; (6) development of a methodology for quantifying DR potential, and extensions of the previous models for multi-target forecasting of net load and DR potential, which achieve a MAPE of 10% for DR potential.

24 POWER TRANSMISSION AND DISTRIBUTION↗

I-GCN: A Graph Convolutional Network Accelerator with Runtime Locality Enhancement through Islandization

In this paper, we propose a novel hardware accelerator for GCN inference called I-GCN that significantly improves data locality and reduces unnecessary computation through a new online graph restructuring algorithm we refer to as islandization. The proposed algorithm finds clusters of nodes with strong internal but weak external connections. The islandization process yields two major benefits. First, by processing islands rather than individual nodes, there is better on-chip data reuse and fewer off-chip memory accesses. Second, there is less redundant computation as aggregation for common/shared neighbors in an island can be reused. The parallel search, identification, and leverage of graph islands are all handled purely in hardware at runtime working in an incremental pipelined manner. This is done without any preprocessing of the graph data or adjustment of the GCN model structure.

Geng, Tong↗

Data Arrays for Microearthquake (MEQ) Monitoring using Deep Learning for the Newberry EGS Sites

The 'Machine Learning Approaches to Predicting Induced Seismicity and Imaging Geothermal Reservoir Properties' project looks to apply machine learning (ML) methods to Microearthquake (MEQ) data for imaging geothermal reservoir properties and forecasting seismic events, in order to advance geothermal exploration and safe geothermal energy production. As part of the project, this submission provides data arrays for 149 microearthquakes between the year 2012 and 2013 at the Newberry EGS Site for use with the Deep Learning Algorithm that has been developed. The data provided includes raw waveform data, location data, normalized waveform data, and processed waveform data. Penn State Geothermal Team has shared the following files from the project: - 149 microearthquakes (MEQs) between 2012 and 2013 at Newberry EGS sites, 'Normalized Waveform Inputs.npz' are normalized waveforms. - labels of 149 MEQs: Processed Waveform Inputs.npz - location labels of 149 MEQs: Location Data.npz Note: .npz is the python file format by NumPy that provides storage of array data.

15 GEOTHERMAL ENERGY↗

In situ melt pool measurements for laser powder bed fusion using multi sensing and correlation analysis

Laser powder bed fusion is a promising technology for local deposition and microstructure control, but it suffers from defects such as delamination and porosity due to the lack of understanding of melt pool dynamics. To study the fundamental behavior of the melt pool, both geometric and thermal sensing with high spatial and temporal resolutions are necessary. This work applies and integrates three advanced sensing technologies: synchrotron X-ray imaging, high-speed IR camera, and high-spatial-resolution IR camera to characterize the evolution of the melt pool shape, keyhole, vapor plume, and thermal evolution in Ti–6Al–4V and 410 stainless steel spot melt cases. Aside from presenting the sensing capability, this paper develops an effective algorithm for high-speed X-ray imaging data to identify melt pool geometries accurately. Preprocessing methods are also implemented for the IR data to estimate the emissivity value and extrapolate the saturated pixels. Quantifications on boundary velocities, melt pool dimensions, thermal gradients, and cooling rates are performed, enabling future comprehensive melt pool dynamics and microstructure analysis. The study discovers a strong correlation between the thermal and X-ray data, demonstrating the feasibility of using relatively cheap IR cameras to predict features that currently can only be captured using costly synchrotron X-ray imaging. Such correlation can be used for future thermal-based melt pool control and model validation.

47 OTHER INSTRUMENTATION↗

Machine Learning (ML) Classifier to Assist Metadata Creation

The Atmospheric Radiation Measurement (ARM) Data Center is responsible for the timely collection, archival, and curation of science data products. These products are freely available through an online data repository. Metadata creation is paramount for scientific users to find and access over seven petabytes of atmospheric science data. The hierarchical metadata structure allows users to search for information at both broad and narrow levels. This project aims to leverage 30 years’ worth of manually created metadata to enable machine predictions of broad-term classifications from narrow-term descriptions. These classification predictions would assist metadata coordinators with their term selections. This paper discusses the cleaning and preprocessing of the training data, the pipeline developed to determine the best model for this task, and the creation of an API metadata classifier for ARM measurement metadata. Our results show that the Linear Support Vector Classification (LinearSVC) algorithm, along with the Term Frequency – Inverse Document Frequency (TF-IDF) vectorizer, is well-suited for our multi-class classification task. Lengthier input training data led to better results, and artificial balancing was unnecessary for this particular use case. This predictive classifier enhances efficiency in metadata creation, as well as supports greater consistency and accuracy in metadata tagging.

Collier, Hannah [ORNL] (ORCID:0000000341284292)↗

Knowledge Oriented Graph Unified Transformer (KOGUT) v0.1

KOGUT — Knowledge Oriented Graph Unified Transformer KOGUT implements the Relational Graph Transformer (RelGT) architecture for knowledge graph link prediction in biological domains, with a primary focus on microbial growth media prediction. While the original RelGT (arXiv:2505.10960) targets relational tables, time series, and multi-table databases, KOGUT adapts this architecture for heterogeneous biological knowledge graphs, providing first-in-class AI predictive models for microbial cultivation. Key Adaptations Beyond Original RelGT: - Knowledge Graph Focus: Applied to biological KGs with semantic node types (taxa, chemicals, media, phenotypes, environments) versus generic relational database tables, trained on the KG-Microbe knowledge graph (1.3M entities, 2.9M edges, 24 relation types). - Multimodal Node Encoding: Integrates node labels, categories, descriptions, and synonyms from KG metadata through learned embedding layers—adapting relational column features to graph node attributes with textual semantics. - Extended K-Hop Subgraph Strategy: Optimized neighborhood sampling (3-hop default, configurable up to 200 nodes) tuned for sparse biological networks, building on the original local-global attention framework with biological relation preservation. - Biolink Predicate Preservation: Type-specific transformations for 24 biological edge semantics (occurs_in, consumes, produces, has_phenotype, subclass_of) beyond standard relational foreign keys, enabling multi-relation link prediction. - Inductive Learning Support: Enables zero-shot predictions for novel taxa through feature-based embeddings (temperature, oxygen requirements, gram stain, cell shape), extending the original transductive relational benchmark scope to uncultured microorganisms. CheapSOTA Performance Optimizations (This Distribution): - VQ-EMA Centroid Attention: Vector quantization with exponential moving average for improved global context modeling (+5-10% MRR improvement). - HDF5 Precomputed Data Loading: One-time preprocessing of k-hop subgraphs to eliminate redundant graph traversals (2-5× training speedup). - Distributed Data Parallel Training: Multi-GPU support for scaling to larger knowledge graphs (tested on 4× NVIDIA A100 GPUs at NERSC Perlmutter). - Mixed Precision Training: Automatic mixed precision (AMP) for memory efficiency and faster training. Advantages Over Standard Knowledge Graph Embedding Models: Combines RelGT's proven multi-element tokenization (features, type, hop, structure) with graph-native biological representations, enabling interpretable link prediction across heterogeneous entities that standard embedding models (TransE, RotatE, ComplEx) and table-based transformers cannot directly model. Achieves near-perfect performance on microbial growth media prediction (MRR: 0.9966, Precision@1: 0.9932, Hit@10: 1.0000) while maintaining explainability through attention-based reasoning over biological pathways. Training Data: - KG-Microbe merged knowledge graph: 1,379,337 nodes, 2,960,472 edges - 24 biological relation types including taxonomic hierarchies, metabolic interactions, phenotype associations, and environmental relationships - Primary prediction task: Growth media suitability for microbial taxa (biolink:occurs_in, 50K edges) - Multi-relation capability: Predicts links for any of the 24 relation types, including chemical consumption/production, phenotype associations, and taxonomic classification Citation: Original RelGT Architecture: Dwivedi et al., "Relational Graph Transformer", arXiv:2505.10960, 2025 KOGUT Implementation: Knowledge Oriented Graph Unified Transformer for Microbial Growth Media Prediction Developed at Lawrence Berkeley National Laboratory (LBNL) Trained on NERSC Perlmutter supercomputer

Joachimiak, Marcin [Lawrence Berkeley National Lab↗

An Envelope Time Synchronous Averaging for Wind Turbine Gearbox Fault Diagnosis

Vibration-based condition monitoring techniques are widely used for diagnosing faults in rotating machines. These techniques are implemented in the time domain, the frequency domain, or both. However, the composite and noisy nature of the raw data collected requires a preprocessing stage such as filtering and decomposition using in-depth processing techniques. Moreover, these methods require good frequency resolution and involve examining a broad frequency range to discern both healthy and faulty cases. In this work, we introduce a simple and fast diagnostic scheme for wind turbine gear teeth wear based on time domain analysis. The proposed method is based on the local minima interpolation of a filtered version of the vibration signal following time synchronous averaging (TSA) technique. Given tachometer signal, the TSA of the vibration data is performed using MTALAB software. Then, local minima of the filtered signal are interpolated using the Piecewise Cubic Hermite Interpolating Polynomial (PCHIP) function. The variance of the interpolated curve built a gear fault index. The derived fault index resulting of the proposed technique allows a substantial distinction between the healthy and faulty cases. Its efficiency is validated using 10 real-world datasets of vibration stemmed from a wind turbine planetary gearbox. The proposed method boasts a low computation time and ease of interpretation, specifically beneficial for gearbox fault diagnosis purposes.

fault diagnosis↗

Transfer learning-based soybean LAI estimations by integrating PROSAIL, UAV, and PlanetScope imagery

Accurate Leaf Area Index (LAI) estimations at the soybean plot scale is achievable using high-resolution Unmanned Aerial Vehicle (UAV) imagery and field measurement samples. However, the limited coverage of UAV flights restricts large-scale remote sensing monitoring in expansive soybean fields. This study leverages the broad coverage and 3-m resolution of PlanetScope satellite imagery to extend LAI prediction from UAV to satellite scales through transfer learning, using UAV-scale LAI estimates as a benchmark to validate cross-scale consistency. To address this challenge, this study proposed the LAI-TransNet, a two-stage transfer learning framework designed for precise and scalable soybean LAI prediction across large areas, demonstrating its effectiveness in cross-scale monitoring. In Stage 1, a UAV-scale benchmark is established using PROSAIL-simulated UAV reflectance data (UAV-Sim) and field-measured soybean LAI. Traditional machine learning, deep learning, and transfer learning models are trained on a hybrid UAV-Sim and field-measured dataset (UAV-Sim_Measured), with the transfer learning model CNN-TL, fine-tuned using pre-trained weights derived from UAV-Sim, achieving the highest accuracy (R 2 = 0.81, RMSE = 0.64 m 2 /m 2 , rRMSE = 11.5 %). In Stage 2, LAI-TransNet is developed by fine-tuning the CNN-TL model on PlanetScope simulated data (PS-Sim), preprocessed via cross-domain mapping to align UAV and satellite spectral features. Real PlanetScope imagery is corrected for reflectance consistency with reference to UAV imagery spectral profiles. LAI-TransNet outperforms other deep learning models trained directly on PS-Sim (R 2 = 0.69 vs. 0.60–0.63), ensuring robust cross-scale consistency. In conclusion, by bridging UAV and satellite scales, LAI-TransNet enables large-scale soybean LAI monitoring, enhancing precision agriculture management through improved monitoring with the PlanetScope imagery.

Leaf area index (LAI)↗

MyCrunchGPT: A LLM Assisted Framework for Scientific Machine Learning

Scientific machine learning (SciML) has advanced recently across many different areas in computational science and engineering. Here, the objective is to integrate data and physics seamlessly without the need of employing elaborate and computationally taxing data assimilation schemes. However, preprocessing, problem formulation, code generation, postprocessing, and analysis are still time- consuming and may prevent SciML from wide applicability in industrial applications and in digital twin frameworks. Here, we integrate the various stages of SciML under the umbrella of ChatGPT, to formulate MyCrunchGPT, which plays the role of a conductor orchestrating the entire workflow of SciML based on simple prompts by the user. Specifically, we present two examples that demonstrate the potential use of MyCrunchGPT in optimizing airfoils in aerodynamics, and in obtaining flow fields in various geometries in interactive mode, with emphasis on the validation stage. To demonstrate the flow of the MyCrunchGPT, and create an infrastructure that can facilitate a broader vision, we built a web app based guided user interface, that includes options for a comprehensive summary report. The overall objective is to extend MyCrunchGPT to handle diverse problems in computational mechanics, design, optimization and controls, and general scientific computing tasks involved in SciML, hence using it as a research assistant tool but also as an educational tool. While here the examples focus on fluid mechanics, future versions will target solid mechanics and materials science, geophysics, systems biology, and bioinformatics.

97 MATHEMATICS AND COMPUTING↗

Tools for Assessing Performance: FY23-Q1 Report [Slides]

LANL was tasked with working with NREL to implement QUIC within the TAP API by the end of Q1, FY2023. If QUICURB cannot be implemented in a way that it can be run outside of the QUIC platform, there will be no further R&D funding for QUIC development. QUIC includes a graphical user interface that imports and preprocesses the various input data streams (3D building databases, ambient wind, vegetation, etc.) and writes them in a format that the QUICURB Fortran executable requires. In order to interface with the TAP API, a Python script was developed that would replace the functionality previously only available within QUICGUI. LANL worked with NREL to test the Python script to ensure that it was operation and fulfilled the requirements.

58 GEOSCIENCES↗