Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “preprocessed data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Creating a Training Dataset for Semantic Segmentation of Canal Networks for Irrigation Modernization

Canal infrastructure has provided critical irrigation water to the western United States for over a century. To continue providing vital water resources to the semi-arid West, irrigation systems must undergo maintenance and modernization. Many canal companies are resource-constrained, and because funding opportunities often require detailed knowledge of existing infrastructure, they can struggle to secure financial capital. We address this problem by creating training data for a semantic segmentation deep learning model to map canal networks throughout the western United States. To create a diverse and robust training dataset, we labelled 1-m NAIP imagery with the locations of no canals, wet canals, and dry/vegetated canals. Since creating these datasets is time consuming, we first developed a preprocessing methodology to identify canals within our four study areas. We used NAIP imagery and provided canal centerline data to buffer, standardize, and cluster the imagery, automating the labeling process as much as possible. However, this still required manual cleaning and manual classification of canal type. Challenges arose when canals were interrupted (e.g., road culverts or piped sections) or when nearby features shared similar characteristics (e.g., irrigated fields, trees, and shadows). Combining automated preprocessing with manual refinement produced four detailed canal masks to be used in the semantic segmentation model developed by Richard Tapia.

13 - HYDRO ENERGY↗

A Consensus Equilibrium Approach for 3-D Land Seismic Shots Recovery

Physical and budget constraints often result in inadequate sampling for accurate subsurface imaging. Preprocessing approaches, such as missing trace interpolation, are typically employed to enhance seismic data in such cases. The compressed sensing (CS) framework has been applied for modeling missing seismic data, which is estimated by sparsity-based computational algorithms. While existing work mainly focuses on recovering missing traces resulting from receiver subsampling, source subsampling has greater economical advantages, as sources are more expensive than receivers. Moreover, stronger image models different from sparsity have not been explored for source recovery. This work presents a consensus equilibrium (CE) approach to recover missing seismic shots, which enables to incorporate various regularization operators modeling different data priors. Here, simulation results from a real 3-D land seismic dataset demonstrate that the CE approach provides more accurate estimations of the linear and hyperbolic events in the recovered shots, compared with pure sparsity-based reconstructions.

58 GEOSCIENCES↗

Comparative Techno-Economic Analysis of Available Feedstocks for High-Temperature Conversion: Whole Tree Thinnings and Mature Pine Residues

The TEA and LCA impacts of utilizing low-cost feedstocks in a high-temperature conversion process are of great interest. Here, we investigate the conversion cost impacts of two underutilized feedstocks from the commercial pine industry; 13-year-old whole trees, representing trees removed for the purpose of pre-commercial thinning, and 23-year-old pine residues, representing a waste stream produced from the deconstruction of mature trees for other purposes. Experimental fast pyrolysis data for each feedstock was used in tandem with results from supply and preprocessing analyses in order to evaluate the field-to-fuel economics. A low difference in MFSP was found between the conversion costs for the two feedstocks, with residues demonstrating a net benefit of $0.27/GGE compared to the whole tree thinnings, driven primarily by feedstock supply costs. This suggests that both whole tree thinnings and pine residues may be viable feedstock options for CFP conversion. Life cycle inventories were also generated for each case, enabling a field-to-fuel quantification of the cost and carbon cycle associated with each feedstock.

09 BIOMASS FUELS↗

Asi Nuclear Energy Sensors Data Portal Chatbot And Data Structuring Tool

The Idaho National Laboratory (INL) is advancing the development of an AI-powered chatbot and data structuring tool specifically designed to accelerate data mining processes for sensor-related information and seamlessly integrate the results into the ASI Sensors Data Portal (https://nes.energy.gov/). By doing so, the software aims to enhance the accessibility, usability, and organization of sensor data for nuclear energy applications. The software initial phase focuses on retrieving comprehensive datasets, prioritizing the past five years of publicly available information from the Office of Scientific and Technical Information (OSTI). These datasets will be meticulously processed to ensure compatibility, employing cleaning and preprocessing steps to eliminate irrelevant, incomplete, or corrupted information, thus establishing a robust foundation for subsequent AI use. The data will serve as the backbone for training an AI model and chatbot, which will act as an interactive tool enabling users to ask complex, context-specific questions and receive accurate, validated answers derived from constrained literature. In parallel, the project incorporates a data structuring process supported by AI to organize sensor information from multiple sources into a standardized format. This structured data will include detailed sensor specifications, such as measurement range, applications, accuracy, and operating conditions, generated and documented with AI. These specifications will be systematically integrated into the sensor portal. To maintain the highest levels of accuracy and relevance, all AI-generated outputs will be reviewed and validated by subject matter experts (SMEs), with additional fields or parameters added as needed. Future stages of the project aim to expand the dataset beyond OSTI to include other sources and potentially incorporate unclassified controlled information (UCI) with restricted access protocols to address security and confidentiality requirements.

Mapes, NormanJ. [Idaho National Laboratory (INL), ↗

Aggregation Methods for Quantifying PTM and Structural Changes in Bottom-Up Proteomics

Bottom-up proteomic workflows rely on sequential preprocessing steps, commonly including peptide-to-protein aggregation (“roll-up”), to enhance data reliability and interpretability. While roll-up is effective for protein-centered analyses, it may be suboptimal for applications focused on post-translational modifications (PTMs) or protein structural changes, such as limited proteolysis–mass spectrometry (LiP-MS). Here, we investigate how different roll-up strategies influence site-level quantification in PTM differential analysis. Moreover, we introduce a novel site-centric roll-up approach tailored for LiP-MS, which quantifies proteolytic fragments rather than solely tryptic peptides. We benchmark these methods through simulation studies, comparing their sensitivity and specificity in detecting structural and PTM-driven changes. We found that the median and mean roll-up methods outperform the sum method in both PTM and LiP proteomics, and site-level quantification in LiP outperforms peptide-level quantification. Our findings offer the first systematic, data-driven guidance for selecting roll-up techniques in site-level proteomic analyses, with implications for both PTM-focused and structural proteomics studies.

aggregation↗

The Double-edged Sword of Data-driven Super-Resolution: Adversarial Super-resolution Models

Data-driven super-resolution (SR) methods are often integrated into imaging pipelines as preprocessing steps to improve downstream tasks such as classification and detection. However, these SR models introduce a previously unexplored attack surface into imaging pipelines. In this paper, we present AdvSR, a framework demonstrating that adversarial behavior can be embedded directly into SR model weights during training, requiring no access to inputs at inference time. Unlike prior attacks that perturb inputs or rely on backdoor triggers, AdvSR operates entirely at the model level. By jointly optimizing for reconstruction quality and targeted adversarial outcomes, AdvSR produces models that appear benign under standard image quality metrics while inducing downstream misclassification. We evaluate AdvSR on three SR architectures (SRCNN, EDSR, SwinIR) paired with a YOLOv11 classifier and demonstrate that AdvSR models can achieve high attack success rates with minimal quality degradation. These findings highlight a new model-level threat for imaging pipelines, with implications for how practitioners source and validate models in safety-critical applications.

Sullivan, Haley [ORNL] (ORCID:0000000274069217)↗

Life-cycle analysis of sustainable aviation fuel production through catalytic hydrothermolysis

Catalytic hydrothermolysis (CH) is a sustainable aviation fuel (SAF) pathway that has been recently approved for use in aircraft fuel production. In alignment with broader sustainable aviation goals, SAF production through CH requires a quantitative assessment of carbon intensity (CI) impacts. In this study, a current-day life-cycle analysis (LCA) was performed on SAF produced via CH to determine the CI. Various oily feedstocks were considered, including vegetable oils (soybean, carinata, camelina and canola) and low-burden oils and greases (corn oil, yellow grease and brown grease). Life-cycle inventory data were collected on all processes within the CH LCA boundary: feedstock cultivation and/or collection, preprocessing, hydrothermal cleanup and CH, biocrude refining, fuel transportation and end use through combustion. Baseline results show that the CH-produced SAF can be generated with CI reductions ranging from 48 to 82% compared with conventional jet fuel. Modest improvements to CI can be achieved through incremental changes to the brown grease CH process, such as relaxing the dewatering specification and implementing renewable natural gas and electricity, which could decrease the CI from 22.9 to 7.9 g CO 2 e/MJ. Total CH fuel production potential was also assessed on the basis of current or near-future feedstock availability and CI. The total biofuel production potential of CH (SAF and renewable fuel co-products) in the US sums to approximately 3487 million gallons per year, with 97% of these volumes having a CI below 50% of that for petroleum jet fuel. The study shows that from an LCA perspective, CH offers a viable SAF pathway that is comparable with existing SAF pathways like hydroprocessed esters and fatty acids.

09 BIOMASS FUELS↗

Coordinate-Based Seismic Interpolation in Irregular Land Survey: A Deep Internal Learning Approach

Physical and budget constraints often result in irregular sampling, which complicates accurate subsurface imaging. Preprocessing approaches, such as missing trace or shot interpolation, are typically employed to enhance seismic data in such cases. Recently, deep learning has been used to address the trace interpolation problem at the expense of large amounts of training data to adequately represent typical seismic events. Nonetheless, most research in this area has focused on trace reconstruction, with little attention having been devoted to shot interpolation. Furthermore, existing methods assume regularly spaced receivers/sources failing in approximating seismic data from real (irregular) surveys. This work presents a novel shot gather interpolation approach which uses a continuous coordinate-based representation of the acquired seismic wavefield parameterized by a neural network. The proposed unsupervised approach, which we call coordinate-based seismic interpolation (CoBSI), enables the prediction of specific seismic characteristics in irregular land surveys without using external data during neural network training. Importantly, experimental results on real and synthetic 3-D data validate the ability of the proposed method to estimate continuous smooth seismic events in the time-space and frequency-wavenumber domains, improving sparsity or low-rank-based interpolation methods.

58 GEOSCIENCES↗

Capturing Travel Mode Adoption in Designing On-Demand Multimodal Transit Systems

This paper studies how to integrate rider mode preferences into the design of on-demand multimodal transit systems (ODMTSs). It is motivated by a common worry in transit agencies that an ODMTS may be poorly designed if the latent demand, that is, new riders adopting the system, is not captured. This paper proposes a bilevel optimization model to address this challenge, in which the leader problem determines the ODMTS design, and the follower problems identify the most cost efficient and convenient route for riders under the chosen design. The leader model contains a choice model for every potential rider that determines whether the rider adopts the ODMTS given her proposed route. To solve the bilevel optimization model, the paper proposes an exact decomposition method that includes Benders optimal cuts and no-good cuts to ensure the consistency of the rider choices in the leader and follower problems. Moreover, to improve computational efficiency, the paper proposes upper and lower bounds on trip durations for the follower problems, valid inequalities that strengthen the no-good cuts, and approaches to reduce the problem size with problem-specific preprocessing techniques. The proposed method is validated using an extensive computational study on a real data set from the Ann Arbor Area Transportation Authority, the transit agency for the broader Ann Arbor and Ypsilanti region in Michigan. The study considers the impact of a number of factors, including the price of on-demand shuttles, the number of hubs, and access to transit systems criteria. The designed ODMTSs feature high adoption rates and significantly shorter trip durations compared with the existing transit system and highlight the benefits of ensuring access for low-income riders. Finally, the computational study demonstrates the efficiency of the decomposition method for the case study and the benefits of computational enhancements that improve the baseline method by several orders of magnitude. Funding: This research was partly supported by National Science Foundation [Leap HI Proposal NSF-1854684] and the Department of Energy [Research Award 7F-30154].

Operations Research & Management Science↗

Data readiness pipeline patterns for scientific AI at scale: Insights from climate, fusion, life sciences, and materials

This article examines how data readiness for AI principles apply to large scientific datasets used to train foundation models. We analyze archetypal workflows across four representative domains—climate, nuclear fusion, life sciences, and materials—to identify common preprocessing patterns and domain‐specific constraints. We introduce a two‐dimensional readiness model that combines canonical preprocessing patterns with a five‐level operational readiness scale, both tailored to high‐performance computing (HPC) environments. This construct helps outline key challenges in transforming large‐scale scientific data into formats suitable for scalable AI training. Together, these dimensions form a conceptual maturity matrix that characterizes scientific data readiness and guides infrastructure development toward standardized, cross‐domain support for scalable and reproducible AI for science. Finally, we evaluate this maturity matrix in the context of case studies including ClimaX (climate), AFLOW (materials), OpenFold (proteomics), and DIII‐D fusion disruption‐prediction workflows, from which we distill lessons learned and provide recommendations to guide practitioners in developing robust AI‐readiness pipelines. Finally, we discuss remaining cross‐cutting challenges that persist across scientific domains.

97 MATHEMATICS AND COMPUTING↗

Distribution of Bound and Free Water in Anatomical Fractions of Pine Residues and Corn Stover as a Function of Biological Degradation

Biomass quality is influenced by water’s abundance, distribution, and status in relation to other chemical species within the polymer matrix. Water interacts with polymers that make up the cell walls, and these interactions govern the physical and chemical changes that occur during the storage and preprocessing of biomass feedstocks. Time-domain nuclear magnetic resonance (TD-NMR) was employed to explore variations in the physical constraints of water within the lignocellulosic microstructure in distinct anatomical fractions of biomass and as a function of biological degradation. The Carr–Purcell–Meiboom–Gill sequence, when combined with knowledge of the chemical composition and physical structure of pine residues and corn stover anatomical fractions, gives an accurate measurement of the bound and free water. In this work, the impacts of storage and biological degradation were investigated to elucidate changes in the status and distribution of water within distinct plant tissues. We also investigate how degradation during storage affects water interactions in different pine residues (e.g., bark, branch, and needle) and corn stover (e.g., cob, leaf, and stalk) anatomical fractions using transverse relaxation times (T2). As demonstrated herein, TD-NMR provides quantitative data on lignocellulosic biomass–water interactions within anatomical fractions, which can further aid in the investigation of preprocessing effects on feedstock quality. Our findings suggest that biological heating enhances biomass–water interactions at the cellular and macromolecular scale. In addition, analysis of three-dimensional scanning electron microscopy reconstructions indicates that surface roughness wavelengths align with microscale roughness, suggesting that pine forestry residue and corn stover particles have primarily hydrophobic exterior surfaces. Furthermore, this study offers multiscale insights into understanding the microstructure, wettability, and chemical environment that dictate diffusion, enzyme access, and recalcitrance of lignocellulosic biomass.

09 BIOMASS FUELS↗

Absorption Correction for Reliable Pair Distribution Functions from Low Energy X-ray Sources

This paper explores the development and testing of a simple absorption correction model for processing powder X-ray diffraction data from Debye−Scherrer geometry laboratory X-ray experiments. This may be used as a preprocessing step before using PDFGETX3 to obtain reliable pair distribution functions (PDFs). Various experimental and theoretical methods for estimating μR were explored, and the most appropriate μR values for correction were identified for different capillary diameters and X-ray beam sizes. We identify operational ranges of μR where a reasonable signal-to-noise ratio is possible after correction. A user-friendly software package, DIFFPY.LABPDFPROC, is presented that can help estimate μR and perform absorption corrections with a rapid calculation for efficient processing.

Absorption↗

Exploring Ion Mobility Mass Spectrometry Data File Conversions to Leverage Existing Tools and Enable New Workflows

Ion mobility (IM) is often combined with LC-MS experiments to provide an additional dimension of separation for complex sample analysis. While highly complex samples are better characterized by the full dimensionality of LC-IM-MS experiments to uncover new information, downstream data analysis workflows are often not equipped to properly mine the additional IM dimension. For many samples the data acquisition benefits of including IM separations are all that is necessary to uncover sample information and the full dimensionality of the data is not required for data analysis. Post-acquisition reduction and adaptation of the dimensions of LC-IM-MS and IM-MS experiments into an LC-MS format opens the possibility to use a plethora of existing software tools. In this work, we developed data file conversion tools to reduce the complexity of IM data analysis. Three data file transformations are introduced in the PNNL PreProcessor software: 1) mapping the IM axis to the LC axis for IM-MS data, 2) converting the drift time vs. m/z space to CCS/z vs m/z space, and 3) transforming All Ions IM/MS mobility aligned fragmentation data to a standard LC-MS DDA data file format. Finally, these new data file conversions are demonstrated with corresponding lipidomics and proteomics workflows that leverage existing LC-MS data analysis software to highlight the benefits of the data transformations.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

FPGA Acceleration of GCN in Light of the Symmetry of Graph Adjacency Matrix

Graph Convolutional Neural Networks (GCNs) are widely used to process large-scale graph data. Different from deep neural networks (DNNs), GCNs are sparse, irregular, and unstructured, posing unique challenges to hardware acceleration with regular processing elements (PEs). In particular, the adjacency matrix of a GCN is extremely sparse, leading to frequent but irregular memory access, low spatial/temporal data locality and poor data reuse. Furthermore, a realistic graph usually consists of unstructured data (e.g., unbalanced distributions), creating significantly different processing times and imbalanced workload for each node in GCN acceleration. To overcome these challenges, we propose an end-to-end hardware-software co-design to accelerate GCNs on resource-constrained FPGAs with the features including: (1) A custom dataflow that leverages symmetry along the diagonal of the adjacency matrix to accelerate feature aggregation for undirected graphs. We utilize either the upper or the lower triangular matrix of the adjacency matrix to perform aggregation in GCN to improve data reuse. (2) Unified compute cores for both aggregation and transform phases, with full support to the symmetry-based dataflow. These cores can be dynamically reconfigured to the systolic mode for transformation or as individual accumulators for aggregation in GCN processing. (3) Preprocessing of the graph in software to rearrange the edges and features to match the custom dataflow. This step improves the regularity in memory access and data reuse in the aggregation phase. Moreover, we quantize the GCN precision from FP32 to INT8 to reduce the memory footprint without losing the inference accuracy. We implement our accelerator design in Intel Stratix10 MX FPGA board with HBM2, and demonstrate 1.3x-110.5x improvement in end-to-end GCN latency as compared to the state-of the-art FPGA implementations, on the graph datasets of Cora, Pubmed, Citeseer and Reddit.

Nair, Gopikrishnan R.↗

Data Readiness for Scientific AI at Scale

This paper examines how Data Readiness for AI (DRAI) principles apply to leadership-scale scientific datasets used to train foundation models. We analyze archetypal workflows across four representative domains—climate, nuclear fusion, bio/health, and materials—to identify common preprocessing patterns and domain-specific constraints. We introduce a two-dimensional readiness framework that combines canonical preprocessing patterns with a five-level operational readiness scale, both tailored to high-performance computing (HPC) environments. This framework helps outline key challenges in transforming large-scale scientific data into formats suitable for scalable AI training. Together, these dimensions form a conceptual maturity matrix that characterizes scientific data readiness and guides infrastructure development toward standardized, cross-domain support for scalable and reproducible AI for science.

Brewer, Wes [ORNL] (ORCID:0000000236393956)↗

Subject-specific modeling framework for particle deposition using computational fluid dynamics

Quantifying particle deposition and dose in the respiratory tract requires a physiologically realistic representation and reproducible computational workflows. However, existing modeling frameworks, such as the International Commission on Radiological Protection (ICRP) compartmental models and the Multiple Path Particle Dosimetry (MPPD) tool, lack detailed deposition profiles and subject-specific capabilities. The combination of advances in computer vision algorithms applied to the respiratory tract and Computational Fluid and Particle Dynamics (CFPD) allows high-fidelity simulations of particle behavior in anatomically accurate geometries derived from individual CT scans. The segmentation, preprocessing, and file preparation task for a CFPD simulation was often time-consuming, and no prior studies to-date have yet presented a fully automated framework. This work presents a fully automated workflow to obtain individualized particle deposition profiles in the human respiratory tract. The pipeline starts with segmenting upper and lower airway geometries using morphological and deep learning-based methods, generating three-dimensional (3D) models from CT imaging data. Next, a series of algorithms are presented to quality check and prepare the 3D geometry for a CFD or CFPD simulation. The preprocessing step includes correcting geometric artifacts, enforcing a physically consistent mesh, and automatically identifying and capping multiple outlets, which is required for CFD/CFPD simulations. These processed models are then input into open-source (OpenFOAM) or commercial (StarCCM+) CFD solvers, where flow and transient particle transport equations — including turbulence and particle–wall interactions are solved under realistic breathing conditions. Finally, the resulting particle deposition profiles can be integrated with Monte Carlo radiation transport codes and state-of-the-art computational phantoms to assess organ-specific absorbed doses in scenarios of radioactive aerosol inhalation. The presented work streamlines respiratory tract segmentation, preprocessing for CFD/CFPD simulations, and integration with dose assessment workflows, reducing manual intervention and improving access to high-fidelity, subject-specific modeling. The high precision in predicted particle deposition and dose distributions can improve personalized treatment strategies in respiratory medicine and refine dose estimates for radiation protection.

AI↗

Data Science Techniques, Assumptions, and Challenges in Alloy Clustering and Property Prediction

Data analytics methods have been increasingly applied to understanding materials chemistry, processing due to the manufacturing approach, and uni-axial and cyclic property relationships in the highly complex space of alloy design. There are several benefits to applying data analytics to this space, including the ability to manage non-linearities in the responses of the alloy attributes and the resulting mechanical properties. However, key difficulties in applying and understanding the results of data analytics include the often lack of reported assumptions and data processing steps necessary to improve interpretation and reproducibility in derived results. In this work, the methods used to generate clustering and correlation analyses for experimental 9% Cr ferritic-martensitic steel data were investigated and the resulting implications for mechanical property predictions were assessed. This work uses principal component analysis, partitioning around medoids, t-SNE, and k-means clustering to investigate trends in composition, processing and microstructure information with creep and tensile properties, building on work done previously using a smaller version of the same dataset. The initial assumptions, preprocessing steps and methods are investigated and outlined in order to depict the fine level of detail required to convey the steps taken to process data and produce analytical results. Here, the variations in the resulting analyses are explored due to the influence of new and more varied data.

36 MATERIALS SCIENCE↗

Diffractive optical computing in free space

Abstract Structured optical materials create new computing paradigms using photons, with transformative impact on various fields, including machine learning, computer vision, imaging, telecommunications, and sensing. This Perspective sheds light on the potential of free-space optical systems based on engineered surfaces for advancing optical computing. Manipulating light in unprecedented ways, emerging structured surfaces enable all-optical implementation of various mathematical functions and machine learning tasks. Diffractive networks, in particular, bring deep-learning principles into the design and operation of free-space optical systems to create new functionalities. Metasurfaces consisting of deeply subwavelength units are achieving exotic optical responses that provide independent control over different properties of light and can bring major advances in computational throughput and data-transfer bandwidth of free-space optical processors. Unlike integrated photonics-based optoelectronic systems that demand preprocessed inputs, free-space optical processors have direct access to all the optical degrees of freedom that carry information about an input scene/object without needing digital recovery or preprocessing of information. To realize the full potential of free-space optical computing architectures, diffractive surfaces and metasurfaces need to advance symbiotically and co-evolve in their designs, 3D fabrication/integration, cascadability, and computing accuracy to serve the needs of next-generation machine vision, computational imaging, mathematical computing, and telecommunication technologies.

36 MATERIALS SCIENCE↗