Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “preprocessed data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

PhotonIDs: ML-Powered Photon Identification System for Dark Count Elimination

Reliable single photon detection is the foundation for practical quantum communication and networking. However, today's superconducting nanowire single photon detector(SNSPD) inherently fails to distinguish between genuine photon events and dark counts, leading to degraded fidelity in long-distance quantum communication. In this work, we introduce PhotonIDs, a machine learning-powered photon identification system that is the first end-to-end solution for real-time discrimination between photons and dark count based on full SNSPD readout signal waveform analysis. PhotonIDs ~demonstrates: 1) an FPGA-based high-speed data acquisition platform that selectively captures the full waveform of signal only while filtering out the background data in real time; 2) an efficient signal preprocessing pipeline, and a novel pseudo-position metric that is derived from the physical temporal-spatial features of each detected event; 3) a hybrid machine learning model with near 98% accuracy achieved on photon/dark count classification. Additionally, proposed PhotonIDs ~ is evaluated on the dark count elimination performance with two real-world case studies: (1) 20 km quantum link, and (2) Erbium ion-based photon emission system. Our result demonstrates that PhotonIDs ~could improve more than 31.2 times of signal-noise-ratio~(SNR) on dark count elimination. PhotonIDs ~ marks a step forward in noise-resilient quantum communication infrastructure.

Linne, Karl C. [Chicago U.] (ORCID:000900091870358↗

Techno-Economic Analysis of decentralized preprocessing systems for fast pyrolysis biorefineries with blended feedstocks in the southeastern United States

This study evaluated the economic feasibility of fast pyrolysis biorefineries fed with blended pine residues and switchgrass in the Southeastern U.S. with different supply chain design. Previous techno-economic analyses (TEA) have focused on either blended biomass or decentralized preprocessing without investigating the impacts of varied process parameters, technology options, and real-world biomass distribution. This study fills the literature gap by modeling scenarios for different biomass blending ratios, biorefinery and preprocessing site (so-called depot) capacities, and alternative preprocessing technologies. High-resolution, real-world geospatial data were analyzed using Geographic Information Systems to facilitate supply chain design and TEA. For a decentralized system, the minimum fuel selling price (MFSP) of biofuel was 3.92–4.33 per gallon gasoline equivalent (GGE), while the MFSP for the centralized biorefinery at the same capacities ranged between 3.75–4.02/GGE. Implementing a high moisture pelleting process depot rather than a conventional pelleting process lowered the MFSP by 0.03–0.17/GGE. Scenario analysis indicated decreased MFSP with increasing biorefinery capacities but not necessarily with increasing depot size. Medium-size depots (500 OMDT/day) achieved the lowest MFSP. Here, this analysis identified the optimal blending ratios for two preprocessing technologies at varied depot sizes. Counterintuitively, increasing the proportion of higher cost switchgrass reduced the MFSP for large biorefineries (>5000 ODMT/day), but increased the MFSP for small biorefineries (1000–2500 ODMT/day). Although the decentralized systems have a higher MFSP based on current analysis, it has other potential benefits such as mitigated supply chain risks and improved feedstock quality that are difficult to be quantified in this TEA.

09 BIOMASS FUELS↗

Air Classification of Forest Residue for Tissue and Ash Separation Efficiency

The goal of this Case Study was to evaluate the performance of air classification of logging residues toward meeting conversion CMAs for carbon and ash contents, as compared to the static status quo Base Case system in which the residues are first dried and then ground in a hammer mill with a 6 mm screen and fines less than 1.18 mm are removed. Also considered were moisture and ash impacts on throughput and Overall Operating Effectiveness (OOE), as well as delivered feedstock cost and minimum fuel selling price (MFSP). Laboratory data on the impacts of fan speed and moisture content on the separation efficiency of soil ash, needles and bark from white wood were received from FCIC Subtask 5.2: Preprocessing, High Temperature Conversion Preprocessing (Jordan Klinger and Tiasha Bhattacharjee, INL). Average throughput and energy consumption data were obtained from the Bioenergy Feedstock National User Facility (BFNUF) (Neal Yancey, INL) for the same air classifier. These data were utilized to develop the necessary response surface equations to perform throughput analysis using discrete event simulation. Feedstock-Conversion Interface Consortium. Because the Base Case status quo system utilizes drying prior to grinding, we modeled the Case Study with drying prior to air classification and subsequent grinding of the separated white wood to isolate the individual quality and cost impacts of air classification relative to the Base Case system.

CMA↗

RAP: Resource-aware Automated GPU Sharing for Multi-GPU Recommendation Model Training and Input Preprocessing

Ensuring high-quality recommendations for newly onboarded users requires the continuous retraining of Deep Learning Recommendation Models (DLRMs) with freshly generated data. To serve the online DLRM retraining, existing solutions use hundreds of CPU computing nodes designated for input preprocessing, causing significant power consumption that surpasses even the power usage of GPU trainers. To this end, we propose RAP, an end-to-end DLRM training framework that supports Resource-aware Automated GPU sharing for DLRM input Preprocessing and Training. The core idea of RAP is to accurately capture the remaining GPU computing resources during DLRM training for input preprocessing, achieving superior training efficiency without requiring additional resources. Specifically, RAP utilizes a co-running cost model to efficiently assess the costs of various input preprocessing operations, and it implements a resource-aware horizontal fusion technique that adaptively merges smaller kernels according to GPU availability, circumventing any interference with DLRM training. In addition, RAP leverages a heuristic searching algorithm that jointly optimizes both the input preprocessing graph mapping and the co-running schedule to maximize the end-to-end DLRM training throughput. The comprehensive evaluation shows that RAP achieves 78.3× speedup on average over CPU-based DLRM input preprocessing frameworks. In addition, the end-to-end training throughput of RAP is only 2.04% lower than the ideal case, which has no input preprocessing overhead.

Wang, Zheng↗

GMT: A deep learning approach to generalized multivariate translation for scientific data analysis and visualization

In scientific visualization, despite the significant advances of deep learning for data generation, researchers have not thoroughly investigated the issue of data translation. We present a new deep learning approach called generalized multivariate translation (GMT) for multivariate time-varying data analysis and visualization. Like V2V, GMT assumes a preprocessing step that selects suitable variables for translation. However, unlike V2V, which only handles one-to-one variable translation during training and inference, GMT enables one-to-many and many-to-many variable translation in the same framework. We leverage the recent StarGAN design from multi-domain image-to-image translation to achieve this generalization capability. We experiment with different loss functions and injection strategies to explore the best choices and leverage pre-training for performance improvement. We compare GMT with other state-of-the-art methods (i.e., Pix2Pix, V2V, StarGAN). Furthermore, the results demonstrate the overall advantage of GMT in translation quality and generalization ability.

97 MATHEMATICS AND COMPUTING↗

Integrating Analytical Solutions and U-Net Model for Predicting Groundwater Contaminant Plumes in Pump-and-Treat Systems

Pump-and-treat (P&T) is a common technique for groundwater remediation involving the extraction and treatment of contaminated water above ground. Optimizing the design and operation of the P&T well network is essential for maximizing the system’s effectiveness and efficiency. However, this optimization often necessitates many model evaluations, leading to computationally demanding tasks. This study introduces a novel approach that integrates analytical solutions for groundwater dynamics with the U-Net (Ronneberger et al., 2015) deep learning framework to predict groundwater contaminant plume migration under dynamic pumping conditions. By incorporating the Thiem equation (Thiem, 1906) into the input preprocessing, the U-Net model transforms sparse well data into a continuous spatial field that captures the hydraulic impacts of pumping activities. This integration enables the model to leverage both deep learning capabilities and classical physics-based groundwater theories, enhancing prediction accuracy and computational efficiency. These advancements can facilitate rapid, large-scale evaluations of P&T optimization simulations, allowing for timely and effective decision-making in well placement and system management. We demonstrate the model's robust performance across both simplified transient 2D models and a more complex 3D heterogeneous site model at the 200 West P&T facility at the Hanford Site. The U-Net-based model offers substantial computational advantages, reducing simulation times significantly compared to full physics-based models and providing a powerful tool for rapid site evaluation and P&T system optimization, such as evaluating alternative P&T well network designs. Our findings highlight the potential of advanced machine learning models to significantly enhance the efficiency and sustainability of groundwater remediation efforts, offering a novel application of U-Net architecture in environmental science.

Pump-and-treat↗

Preprocessing for Unintended Conducted Emissions Classification with ResNet

Characterization of Unintended Conducted Emissions (UCE) from electronic devices is important when diagnosing electromagnetic interference, performing nonintrusive load monitoring (NILM) of power systems, and monitoring electronic device health, among other applications. Prior work has demonstrated that UCE analysis can serve as a diagnostic tool for energy efficiency investigations and detailed load analysis. While explaining the feature selection of deep networks with certainty is often not fully comprehensive, or in other applications, quite lacking, additional tools/methods for further corroboration and confirmation can help further the understanding of the researcher. This is true especially in the subject application of the study in this paper. Often the focus of such efforts is the selected features themselves, and there is not as much understanding gained about the noise in the collected data. If selected feature and noise characteristics are known, it can be used to further shape the design of the deep network or associated preprocessing. This is additionally difficult when the available data are limited, as in the case which the authors investigated in this study. Here, the authors present a novel work (which is a proposed complementary portion of the overall solution to the deep network classification explainability problem for this application) by applying a systematic progression of preprocessing and a deep neural network (ResNet architecture) to classify UCE data obtained via current transformers. By using a methodical application of preprocessing techniques prior to a deep classifier, hypotheses can be produced concerning what features the deep network deems important relative to what it perceives as noise. For instance, it is hypothesized in this particular study as a result of execution of the proposed method and periodic inspection of the classifier output that the UCE spectral features are relatively close to each other or to the interferers, as systematically reducing the beta parameter of the Kaiser window produced progressively better classification performance, but only to a point, as going below the Beta of eight produced decreased classifier performance, as well as the hypothesis that further spectral feature resolution was not as important to the classifier as rejection of the leakage from a spectrally distant interference. This can be very important in unpredictable low-FNR applications, where knowing the difference between features and noise is difficult. As a side-benefit, much was learned regarding the best preprocessing to use with the selected deep network for the UCE collected from these low power consumer devices obtained via current transformers. Baseline rectangular windowed FFT preprocessing provided a 62% classification increase versus using raw samples. After performing a more optimal preprocessing, more than 90% classification accuracy was achieved across 18 low-power consumer devices for scenarios in which the in-band features-to-noise ratio (FNR) was very poor.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Adaptively coupled phase retrieval in multi-peak Bragg coherent diffraction imaging

Recent advances in Bragg coherent diffraction imaging (BCDI) experimental techniques permit routine measurement of multiple Bragg peaks from a single crystalline grain. The resulting images contain the full lattice distortion vector field which can be differentiated to provide lattice strain and rotation. With the advent of fourth-generation synchrotron light sources, such multi-peak datasets are produced at high rates, facilitating the need for rapid phase retrieval of the multiple peaks and subsequent image analysis. Here we describe and demonstrate a new implementation of a coupled phase retrieval technique for multi-peak BCDI which simultaneously treats each Bragg peak of the dataset and produces a three-dimensional image of the crystal's morphology and lattice distortion field. In addition, this method uses the redundant information contained in the various Bragg diffraction patterns to detect and suppress spurious signal appearing on the detector in a subset of the measurements. Compared with manual data editing, adaptive coupling produces a more consistent phase profile in reciprocal space and sharper surfaces in direct space, with no significant difference in computational cost. These improvements reduce the need for manual preprocessing and enable robust high-throughput analysis of multi-peak BCDI data, supporting near-real-time strain microscopy at modern synchrotron facilities.

36 MATERIALS SCIENCE↗

Common Column Identification for Table Similarity Detection in Electrified Transportation Data Lakes

Electrified transportation often requires researchers and operators to interact with datasets from a wide range of sources and disciplines, such as transportation, power systems, public health, policies, and regulations. These datasets vary in quality and format, making it difficult to understand, preprocess, and identify key columns representing real-world entities or values for indexing and joining, which can negatively impact downstream analysis and operation. Existing solutions are limited, requiring extensive manual customization or data expertise to utilize. In this article, we propose a multi-layered approach to automatically identify key columns to expedite preprocessing and aid in analysis of electrified transportation data. Our method leverages a dynamic ontology to identify common fields and an information theory-based strategy for edge cases that are difficult to generalize. Evaluations on a number of datasets from data.gov and kaggle.com show improved performance of our methods over several baseline techniques, and our ablation analyses illustrate the efficacy of individual components of our method. Our case studies also demonstrate that our methods have the potential to improve analysis of electrified transportation data and aid in automatic integration of such datasets.

33 ADVANCED PROPULSION SYSTEMS↗

A Centralized AI Lakehouse Framework for Brain Tumor MRI Classification and Segmentation, University KPI Forecasting, and Water Potability Prediction

In many university and healthcare projects, models are built for very different data types such as tables, institutional time series, and medical images, but they are deployed as separate applications. In this work, that separation made testing and maintenance difficult because each module had its own pipeline and runtime requirements. This paper presents an integrated AI lakehouse-style implementation that runs three model pipelines inside one containerized backend. For medical imaging, we used MRI datasets from IEEE DataPort: a four-class classification set with 7012 images (5708 train/1304 test) and a segmentation set with 3063 image–mask pairs. The classification model (ResNet50 transfer learning) is evaluated using a proper train–validation–test protocol across multiple splits (80/10/10, 70/10/20, 60/10/30, and 10/30/60), achieving a test accuracy of 99.00% under the standard 80/10/10 split. Additionally, a patient-level evaluation is conducted using an external glioma dataset to provide a more realistic assessment without data leakage. The segmentation model (DeepLabV3-ResNet50) achieved 83.09% validation mIoU and 88.79% Dice score. For university KPI forecasting, we used annual IPEDS and NSF HERD data from 2010 to 2023 for three universities (BSU, EOU, and UAB). To examine the effect of preprocessing on forecasting performance, two case studies are conducted. In the first case, linear interpolation is applied to generate semester-level data. In the second case, the original annual data is used directly without interpolation. Random Forest regression and ARIMA models are evaluated using MAE, RMSE, MAPE, and R 2 . The results showed that interpolation improved apparent forecasting performance due to smoothing, while evaluation on the original annual data provided a more realistic assessment of model behavior. To further validate the framework on a larger dataset, an additional case study is conducted using a student dropout dataset. For water potability, we trained and compared multiple tabular classifiers on a large dataset (1,048,575 samples). A Random Forest model (100 trees, max depth 10) achieved 85.86% test accuracy and high recall for unsafe samples (0.8447). All modules are served via FastAPI and deployed together using Docker, with workflow automation routing requests to the correct endpoint. System-level benchmarking indicates that the backend maintains stable throughput and latency under concurrent requests.

97 MATHEMATICS AND COMPUTING↗

Particle Scale Impacts on Deconstruction Energy of Pine Residues

The goal of this Case Study was to quantify the impacts of variable moisture and ash on hammer mill throughput and energy consumption and on generation of fines that are not able to be fed to conversion, as compared to a status quo Base Case system. Also considered was convertible carbon content (minimum carbon specification) and maximum ash content and the delivered feedstock cost impacts of not being able to feed residue not meeting both specifications to the conversion reactor. Laboratory data on the impacts of input particle size and moisture content on the exit particle size were received from FCIC Subtask 5.2: Preprocessing, High Temperature Conversion Preprocessing from their single particle impact population balance modeling study (Tiasha Bhattacharjee, INL). Additional throughput and energy consumption data were obtained from FCIC Subtask 5.2 (Jordan Klinger, INL) for the same grinder with a 6 mm screen in place. These data were utilized to develop the necessary response surface equations to perform throughput analysis using discrete event simulation. Because the ash contents in the separated fines had not been analyzed in the laboratory at the time of the model runs, we chose to assume that the ash distributed proportionally with total mass into the overs and unders in the disk screen following grinding.

energy consumption↗

0BGRaman: Graph Network based Simulator for Forecasting Molecular Polarizability

This report presents the work performed under the GRaman project, sponsored by the PCSD LDRD Seed program. The project aimed at accelerating ab initio molecular dynamics simulation using Graph Networks. The Graph Network framework is a ML framework that has been successfully employed to simulate the dynamics of several physical systems: including water splashing in a container and flags moving with the wind. In this effort, we performed a data collection campaign for 3 different molecules of interest. We have built tools for preprocessing the trajectories obtained by simulating Raman Spectroscopy with NWChem and translating them into a suitable format for training. We have developed a training algorithm to train the Graph Network based simulators based on our data and developed a simulator that produces trajectories in the same NWChem format. While the tool has improved with each iteration of development and subsequent experiments, the current state of the tool does not allow to directly incorporate the technology within the NWChem framework because the trajectories produced by the tool are not yet accurate enough. However, the technology has proved to have good potential and it is certainly worth further research and development.

97 MATHEMATICS AND COMPUTING↗

Impact of Anatomical Fractionation of Corn Stover on Hammer Mill Throughput and Energy Consumption

The goal of this Case Study was to quantify the impacts of variable moisture and ash on hammer mill throughput and energy consumption and on loss of very wet stover that causes failures in the first stage grinder and that are not able to be fed to conversion, as compared to a status quo Base Case system. Also considered was convertible carbohydrate content (minimum total carbohydrate specification), maximum ash content and the delivered feedstock cost impacts of not being able to feed stover that did not meet the total carbohydrate specification to the conversion reactor. Laboratory data on the impacts of moisture content and tissue fraction on throughput and energy consumption in a stage 2 hammer mill were received from FCIC Subtask 5.1: Preprocessing, Corn Stover Preprocessing (Neal Yancey and Sergio Hernandez, INL). Additional air classifier throughput, energy consumption and separation efficiency data were obtained from FCIC Subtask 5.1 (Neal Yancey, INL) for the new air classifier, which has three exit streams (lights, middle and heavies). These data were utilized to develop the necessary response surface equations to perform throughput analysis using discrete event simulation. Because the ash contents and particle sizes had not been analyzed in the laboratory at the time of the model runs, we assumed that the ash distributed proportionally with total mass into the lights and heavies in an air classifier having two exit streams (lights and heavies) and that the lights fraction from the air classifier was not removed.

ash content↗

Unsupervised classification for region of interest in X-ray ptychography

X-ray ptychography offers high-resolution imaging of large areas at a high computational cost due to the large volume of data provided. To address the cost issue, we propose a physics-informed unsupervised classification algorithm that is performed prior to reconstruction and removes data outside the region of interest (RoI) based on the multimodal features present in the diffraction patterns. The preprocessing time for the proposed method is inconsequential in contrast to the resource-intensive reconstruction process, leading to an impressive reduction in the data workload to a mere 20% of the initial dataset. This capability consequently reduces computational time dramatically while preserving reconstruction quality. Through further segmentation of the diffraction patterns, our proposed approach can also detect features that are smaller than beam size and correctly classify them as within the RoI.

97 MATHEMATICS AND COMPUTING↗

COBRA:COMPUTED-TOMOGRAPHY BASED RANDOM-FIELD APPROXIMATION

SF-25-115 COBRA (COmputed-tomography Based Random-field Approximation) is a Python application for generating statistically equivalent random fields from CT-scan imagery. It leverages Karhunen–Loève expansions to model microstructural variability, enabling users to: Preprocess CT scans (filtering and Gaussian transformation); Fit covariance kernels fromempirical data; Solve eigenproblems to obtain KL modes; Sample random fields onsistent with fitted statistics; Postprocess samples back into the physical domain.

Hu, Tianchen↗

REDI – Readiness Engine for Data Integration

The Readiness Engine for Data Integration (REDI) is an open-source framework for automating, standardizing, and assessing the process of preparing scientific data for AI training. REDI implements a five-stage pipeline (ingest, preprocess, transform, structure, output) with per-stage provenance instrumentation via Flowcept, domain-aware transformation logic (PII anonymization, regridding, graph encoding, and more), and built-in readiness assessment and validation modes. REDI has been evaluated across climate, proteomics, materials science, and nuclear fusion datasets, demonstrating near-ideal parallel scaling to 100 nodes on OLCF's Frontier system. REDI is deployable as an agent-callable skill in coding environments such as Claude Code and OpenAI Codex, and is complemented by SetGo for FAIR compliance and catalog publication.

Brewer, Wesley [Oak Ridge National Laboratory (ORN↗

Precision Plant Biomass Characterization in Agriculture: Harnessing Machine Learning and Hyperspectral Imaging [Slides]

Efficient Biomass Separation Object detection of anatomical parts (Cob, Stalk, Husk) in IR images enables precise separation, improving preprocessing (e.g., drying, grinding) for biofuel production. Detailed Biomass Characterization with Hyperspectral Data Hyperspectral imaging captures spectral signatures of biomass, allowing for the identification of specific traits like moisture content, lignin levels, and nutrient composition, leading to optimized treatments for each biomass part. Enhanced Feedstock Quality By leveraging hyperspectral data, feedstock can be processed based on its chemical composition, improving conversion efficiency and biofuel yield. Automation for Large-Scale Operations Automated object detection and hyperspectral data analysis reduce manual labor, ensuring accurate sorting and faster processing, making large-scale biofuel production more efficient. Maximized Biomass Utilization Accurate identification of biomass properties minimizes waste and ensures that each part is processed according to its highest biofuel potential.

09 BIOMASS FUELS↗

BCLink User Documentation

BCLink is a shared programming library providing general functionality to read temporally- and spatially-varying boundary condition data (contained in Exodus files) into MDG codes (ParaDyn and Diablo). This document is intended for prospective users of BCLink, presenting a summary of the available features of the library, including details regarding the input syntax with accompanying examples. The library’s capabilities are demonstrated through example problems run in ParaDyn, and using the BCRemap command line utility – a stand-alone preprocessing tool which leverages the native functionality provided by the BCLink library to remap surface data between dissimilar meshes.

42 ENGINEERING↗