Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “preprocessed”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

Adaptive continuity-preserving simplification of street networks

Street network data is widely used to study human-based activities and urban structure. Often, these data are geared towards transportation applications, which require highly granular, directed graphs that capture the complex relationships of potential traffic patterns. While this level of network detail is critical for certain fine-grained mobility models, it represents a hindrance for studies concerned with the morphology of the street network. For the latter case, street network simplification — the process of converting a highly granular input network into its most simple morphological form — is a necessary, but highly tedious preprocessing step, especially when conducted manually. In this manuscript, we develop and present a novel adaptive algorithm for simplifying street networks that is both fully automated and able to mimic results obtained through a manual simplification routine. The algorithm — available in the neatnet Python package — outperforms current state-of-the-art procedures when comparing those methods to manually, human-simplified data, while preserving network continuity.

Python↗

Pavement condition and climatic data in southeast Texas: A dataset for evaluating flood impacts on pavement performance

Effective pavement maintenance is essential for economic stability, optimal network performance, and roadway safety. Achieving this requires thorough evaluation of pavement conditions, including structural integrity, surface roughness, and distress characteristics. Pavement performance indicators play a critical role in influencing vehicle safety and ride quality. Recent advances have emphasized the use of data-driven modeling to anticipate pavement behavior, with the goal of optimizing resource allocation and refining Maintenance and Rehabilitation (M&R) strategies through accurate condition assessment. A foundational requirement for these modeling efforts is the availability of standardized, high-quality datasets that can support robust and reproducible infrastructure analysis. This data article presents a comprehensive dataset assembled to facilitate pavement performance prediction, with a geographic focus on Southeast Texas, particularly the flood-vulnerable area of Beaumont. The dataset encompasses pavement and traffic attributes, meteorological records, flood simulation outputs, ground deformation measurements, and topographic indices, enabling detailed examination of both load-associated and non-load-associated degradation mechanisms. Data preprocessing was performed using ArcGIS Pro, Microsoft Excel, and Python to ensure consistency and usability in data-driven modeling applications, including machine learning workflows. Key contributions of this dataset include its utility in analyzing the climatic and environmental factors affecting pavement conditions, identifying critical predictive features, and enabling in-depth correlation analysis across diverse variables. By filling existing gaps in input variable selection resources, this dataset supports the development of predictive tools for estimating future maintenance demand and enhancing the resilience of pavement networks in flood-impacted areas. The resource highlights the importance of standardized datasets for advancing pavement management practices and provides a robust foundation for ongoing infrastructure performance modeling.

42 ENGINEERING↗

A portable application framework for energy management and information systems (EMIS) solutions using Brick semantic schema

This paper introduces a portable framework for developing, scaling and maintaining energy management and information systems (EMIS) applications using an ontology-based approach. Key contributions include an interoperable layer based on Brick schema, the formalization of application constraints pertaining metadata and data requirements, and a field demonstration. The framework allows for querying metadata models, fetching data, preprocessing, and analyzing data, thereby offering a modular and flexible workflow for application development. Its effectiveness is demonstrated through a case study involving the development and implementation of a data-driven anomaly detection tool for the photovoltaic systems installed at the Politecnico di Torino, Italy. During eight months of testing, the framework was used to tackle practical challenges including: (i) developing a machine learning-based anomaly detection pipeline, (ii) replacing data-driven models during operation, (iii) optimizing model deployment and retraining, (iv) handling critical changes in variable naming conventions and sensor availability (v) extending the pipeline from one system to additional ones.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Optimization of harvest and logistics for multiple lignocellulosic biomass feedstocks in the northeastern United States

A mixed-integer linear programming model was developed to optimize the multiple biomass feedstock supply chains, including feedstock establishment, harvest, storage, transportation, and preprocessing. The model was applied for analyses of multiple biomass feedstocks at county level for 13 states in the northeastern United States. In the base case with a demand of 180,000 dry Mg/year of biomass, the delivered costs ranged from $\$67.90$ to $\$86.97$ per dry Mg with an average of $\$79.58$ /dry Mg. The biomass delivered costs by county were from $\$67.90$ to 150.81 per dry Mg across the northeastern U.S. Considered the entire study area, the delivered cost averaged $\$85.30$ /dry Mg for forest residues, $\$84.47$ /dry Mg for hybrid willow, $\$99.68$ for switchgrass and $\$97.87$ per dry Mg for Miscanthus. Seventy seven out of 387 counties could be able to deliver biomass at $84 per dry Mg or less a target set by US DOE by 2022. A sensitivity analysis was also conducted to evaluate the effects of feedstock availability, feedstock price, moisture content, procurement radius, and facility demand on the delivered cost. Lastly, our results showed that procurement radius, facility capacity, and forest residue availability were the most sensitive factors affecting the biomass delivered costs.

09 BIOMASS FUELS↗

Exploratory analysis and performance prediction of big data transfer in High-performance Networks

Big data transfer in large-scale scientific and business applications is increasingly carried out over connections with guaranteed bandwidth provisioned in High-performance Networks (HPNs) via advance bandwidth reservation. Provisioning agents need to carefully schedule data transfer requests, compute network paths, and allocate appropriate bandwidths. Such reserved bandwidths, if not fully utilized, could be simply wasted due to the exclusive access during the approved time window, and cause extra overhead and complexity for resource management. This calls for accurate performance prediction to reserve bandwidths that match actual needs and avoid over-provisioning. We employ machine learning algorithms to predict big data transfer performance based on extensive performance measurements collected in the past several years from data transfer tests using different protocols and toolkits between various end sites on several real-life physical or emulated testbeds. We first analyze the performance patterns in response to a comprehensive list of parameters in end-host systems, network connections, and data transfer applications, which motivate the use of machine learning and also help us identify the effects of latent factors. We then propose threshold- and clustering-based methods to eliminate negative effects of latent factors in data preprocessing and build a robust performance predictor based on customized domain-oriented loss functions. The performance of the proposed methods is verified by extensive experiments using SVR and RFR as well as theoretical analysis of the general performance bound.

97 MATHEMATICS AND COMPUTING↗

Spatial and temporal characterization of municipal solid waste based on resource recovery pathways

This study presents a two-year, quarterly assessment of MSW across four source sectors (residential, schools, restaurants, and grocery stores) from sixteen sites across five U.S. states. MSW was manually sorted into 27 categories and aggregated into pathway fractions: high-moisture (HM) organics, low-moisture (LM) organics, recyclable (RC) materials, and residuals for disposal. Organics represented 89 % of the MSW stream. The largest fraction was HM organics consisting of food waste (31 %) and yard waste (3 %), with large coefficient of variations (CV), 79 and 278 %, respectively, reflecting high seasonal and site variability that varied significantly (p < 0.01) across sampling periods. The HM fraction showed properties favorable for anaerobic digestion, with moisture content ranging from 56 to 95 % and volatile solids ranges of 86-95 %. In contrast, the LM and RC fractions remained more stable (plastics CV = 41 %; paper CV = 53 %) with heating values up to 26.9 MJ/kg across sources, reflecting suitability for gasification. Microstructural analysis revealed less porosity in residential waste sampled at the landfill, which can influence preprocessing efficiency and microbial accessibility. Pathway informed allocations showed that 35 % of MSW is suitable for anaerobic digestion, 36 % for gasification, and 18 % for recycling, leaving 11 % requiring landfill disposal. These results provide quantitative evidence to determine feedstock allocation, waste-to-energy system design, and the development of data-driven sustainability and resource recovery strategies within a circular bioeconomy.

09 BIOMASS FUELS↗

In-depth analysis on parallel processing patterns for high-performance Dataframes

The Data Science domain has expanded monumentally in both research and industry communities during the past decade, predominantly owing to the Big Data revolution. Artificial Intelligence (AI) and Machine Learning (ML) are bringing more complexities to data engineering applications, which are now integrated into data processing pipelines to process terabytes of data. Typically, a significant amount of time is spent on data preprocessing in these pipelines, and hence improving its efficiency directly impacts the overall pipeline performance. The community has recently embraced the concept of Dataframes as the de-facto data structure for data representation and manipulation. However, the most widely used serial Dataframes today (R, pandas) experience performance limitations while working on even moderately large data sets. We believe that there is plenty of room for improvement by taking a look at this problem from a high-performance computing point of view. In a prior publication, we presented a set of parallel processing patterns for distributed dataframe operators and the reference runtime implementation, Cylon. In this paper, we are expanding on the initial concept by introducing a cost model for evaluating the said patterns. Furthermore, we evaluate the performance of Cylon on the ORNL Summit supercomputer.

97 MATHEMATICS AND COMPUTING↗

Expanded analysis of machine learning models for nuclear transient identification using TPOT

Industries around the world are becoming more and more data driven. The nuclear field is no exception with several different applications being proposed. One popular area of research is the use of machine learning in transient detection. This paper seeks to build upon a previous study which made use of the AutoML package TPOT to train traditional machine learning models to classify transient events occurring with a reactor. Synthetic data was once again collected using a GPWR reactor simulator. Data on 12 different events was collected using 15 different initial conditions. Here, a dataset consisting of over 100,000 data points was compiled and used to train 7 different machine learning models using a pre-defined TPOT dictionary with 12 different preprocessing techniques. Three of the trained models were able to produce validation results in the 90s with the expanded dataset. Once the models were trained, it was possible to look into where during the simulation, misclassifications occurred. Using these three models, analysis was done to determine if TPOT could be used to train models that were effective if important features were missing. The results from this were positive with the newly trained models scoring close to the original models. Finally, to conclude this study, the three high performing models were retrained using different random states to see if there was any major variation when different states were used.

42 ENGINEERING↗

A generalized machine learning workflow to visualize mechanical discontinuity

Accurate detection and mapping of mechanical discontinuity in materials has widespread industrial and research applications. Herein, we developed a generalized machine-learning framework for visualizing single mechanical discontinuity embedded in material of any composition, velocity, density, porosity, and size with limited data. The proposed visualization of discontinuity requires accurate estimations of the length, location, and orientation of the embedded discontinuity by processing multipoint wave-transmission measurements. k-Wave simulator is used to create a large dataset of elastic waveforms recorded during multi-point wave-transmission measurements through materials containing single mechanical discontinuity. k-Wave simulator considers the wave attenuation, dispersion, and mode conversion in wave motion. Discrete wavelet transform (DWT) and statistical feature extraction are essential for data preprocessing prior to the data-driven model development. DWT also minimizes the effect of noise. Using hyper-parameter tuning and cross validation, gradient boosting regression can visualize the mechanical discontinuity with an accuracy of 0.85, in terms of coefficient of determination. A double-layered neural network-based regression has better performance with an accuracy of 0.95. Use of convolutional neural network converts the predictive task from a waveform processing to an image processing problem. Convolutional neural network achieved a generalization performance of 0.91. The proposed generalized workflow requires robust simulation of wave propagation, signal processing, feature engineering, and model evaluation. Sensors closest to the source and those located opposite the source are the most significant for the desired visualization. Notably, the sensors closest to the source capture the non-linear associations, whereas the sensor on the border opposite to the source capture the linear associations between the measured waveforms and the properties of the mechanical discontinuity.

42 ENGINEERING↗

Machine learning-assisted upscaling analysis of reservoir rock core properties based on micro-computed tomography imagery

Optimum solutions for geologic modeling and reservoir simulation in industries such as oil and gas recovery and carbon capture and storage require accurate characterization of reservoir properties, which are often heterogeneous. In this study, high-quality micro-computed tomography (CT) images (1.475-μm/pixel resolution) of a sandstone core acquired from the Bell Creek oil field, USA, were used to provide nondestructive analysis of pore- and core-scale heterogeneity across measurement scales of 94–566 μm. In addition to characterizing the as-received sample, the core sample was flooded with brine to evaluate the capacity of the core sample to receive injected fluids. The micro-CT images were systematically segmented into pore spaces and grains via machine learning (ML) steps including image preprocessing, label creation using a traditional ML method based on limited manual image annotation, and finally U-Net segmentation. The segmented image stacks were reconstructed into digital cubes of various scales of voxel lengths. The 3D porosity values were calculated for all the digital cubes, and the fractal dimensions of the cubes were estimated using a box-counting method. The results showed that smaller cubes had greater heterogeneity and that the porosity values could be accurately estimated by fractal dimension and voxel lengths using ML models. For the core sample with brine flooding, the ratio of pores filled by brine to the total pore space was related to the porosity and could also be accurately estimated by porosity, fractal dimension, and voxel lengths using ML models. In conclusion, the results of this study demonstrate that the concept of fractal dimension can be a useful vector to perform upscaling analysis of sandstone rock heterogeneity from the pore to core scale and that fractal dimensions can be used to estimate porosity values and pore space-filling capacity across those scales.

58 GEOSCIENCES↗

Post-landing major element quantification using SuperCam laser induced breakdown spectroscopy

The SuperCam instrument on the Perseverance Mars 2020 rover uses a pulsed 1064 nm laser to ablate targets at a distance and conduct laser induced breakdown spectroscopy (LIBS) by analyzing the light from the resulting plasma. SuperCam LIBS spectra are preprocessed to remove ambient light, noise, and the continuum signal present in LIBS observations. Prior to quantification, spectra are masked to remove noisier spectrometer regions and spectra are normalized to minimize signal fluctuations and effects of target distance. In some cases, the spectra are also standardized or binned prior to quantification. To determine quantitative elemental compositions of diverse geologic materials at Jezero crater, Mars, we use a suite of 1198 laboratory spectra of 334 well-characterized reference samples. The samples were selected to span a wide range of compositions and include typical silicate rocks, pure minerals (e.g., silicates, sulfates, carbonates, oxides), more unusual compositions (e.g., Mn ore and sodalite), and replicates of the sintered SuperCam calibration targets (SCCTs) onboard the rover. For each major element (SiO 2 , TiO 2 , Al 2 O 3 , FeO T , MgO, CaO, Na 2 O, K 2 O), the database was subdivided into five “folds” with similar distributions of the element of interest. One fold was held out as an independent test set, and the remaining four folds were used to optimize multivariate regression models relating the spectrum to the composition. We considered a variety of models, and selected several for further investigation for each element, based primarily on the root mean squared error of prediction (RMSEP) on the test set, when analyzed at 3 m. In cases with several models of comparable performance at 3 m, we incorporated the SCCT performance at different distances to choose the preferred model. Shortly after landing on Mars and collecting initial spectra of geologic targets, we selected one model per element. Subsequently, with additional data from geologic targets, some models were revised to ensure results that are more consistent with geochemical constraints. The calibration discussed here is a snapshot of an ongoing effort to deliver the most accurate chemical compositions with SuperCam LIBS.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Designing resilient IoT and Edge Computing with federated tinyML

The rapid growth of the Internet of Things (IoT) and Edge Computing (EC) has brought significant conveniences to modern society but has also greatly expanded the cyber attack surfaces, particularly as these technologies are being increasingly integrated into critical systems such as power grids, healthcare, and smart homes. Here, to improve IoT/EC’s cybersecurity posture, we leveraged Artificial Intelligence (AI) and Machine Learning (ML) by employing tinyML to monitor voluminous IoT data for cyber threats while addressing devices’ resource constraints, and utilizing Federated Learning (FL) to share local detection knowledge across the system while preserving privacy. Building on our three-layer architecture combining tinyML and FL to enhance autonomous cyber attack detection, this paper demonstrated that the architecture improves detection accuracy, reduces resource consumption, and enables lightweight, secure IoT device monitoring. These results were validated using the public N-BaIoT dataset as well as real IoT network traffic data collected under multiple attack scenarios from our testbeds. Additionally, we introduced an enhanced FL methodology with a novel preprocessing stage, including federated feature selection and global preprocessor construction, to address IoT/EC data heterogeneity. We developed a physical IoT testbed for attack simulations and data collection, implemented a tinyML-powered detector for realistic model validation, and also built a virtual testbed for scalable evaluations of FL models across diverse network environments.

Cognitive cyber↗

Data-driven method for electric vehicle charging demand analysis: Case study in Virginia

Electric vehicle (EV) adoption in the U.S. will be accelerated by the historic $7.5 billion public investments in EV charging infrastructure. Careful analysis of EV charging demands plays a vital role in understanding the energy requirements, power grid impact, and smart charging management opportunities of EVs. To this end, this paper develops a data-driven trip-chaining-based modeling framework including five steps: Trip data acquisition and preprocessing, EV adoption modeling, travel itinerary synthesis, EV charging demand simulation and EV load profile generation. The developed analysis framework was demonstrated using real-world data for one region in Virginia, U.S. The results show that the proposed modeling framework can work effectively. For the study region in 2040, the predicted number of plug-in EVs is 470,114, resulting in a weekly charging demand of 38,078,127 kWh (55% home, 9% work, and 36% public) in September and 45,920,358 kWh (61% home, 9% work, and 30% public) in February.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Statistical characterization of experimental magnetized liner inertial fusion stagnation images using deep-learning-based fuel–background segmentation

Significant variety is observed in spherical crystal x-ray imager (SCXI) data for the stagnated fuel–liner system created in Magnetized Liner Inertial Fusion (MagLIF) experiments conducted at the Sandia National Laboratories Z-facility. As a result, image analysis tasks involving, e.g., region-of-interest selection (i.e. segmentation), background subtraction and image registration have generally required tedious manual treatment leading to increased risk of irreproducibility, lack of uncertainty quantification and smaller-scale studies using only a fraction of available data. We present a convolutional neural network (CNN)-based pipeline to automate much of the image processing workflow. This tool enabled batch preprocessing of an ensemble of N scans = 139 SCXI images across N exp = 67 different experiments for subsequent study. The pipeline begins by segmenting images into the stagnated fuel and background using a CNN trained on synthetic images generated from a geometric model of a physical three-dimensional plasma. The resulting segmentation allows for a rules-based registration. Our approach flexibly handles rarely occurring artifacts through minimal user input and avoids the need for extensive hand labelling and augmentation of our experimental dataset that would be needed to train an end-to-end pipeline. Here we also fit background pixels using low-degree polynomials, and perform a statistical assessment of the background and noise properties over the entire image database. Our results provide a guide for choices made in statistical inference models using stagnation image data and can be applied in the generation of synthetic datasets with realistic choices of noise statistics and background models used for machine learning tasks in MagLIF data analysis. We anticipate that the method may be readily extended to automate other MagLIF stagnation imaging applications.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Comparing Sensor Fusion and Multimodal Chemometric Models for Monitoring U(VI) in Complex Environments Representative of Irradiated Nuclear Fuel

Optical sensors and chemometric models were leveraged for the quantification of uranium(VI) (0–100 μg mL –1 ), europium (0–150 μg mL –1 ), samarium (0–250 μg mL –1 ), praseodymium (0–350 μg mL –1 ), neodymium (0–1000 μg mL –1 ), and HNO 3 (2–4 M) with varying corrosion product (iron, nickel, and chromium) levels using laser fluorescence, Raman scattering, and ultraviolet–visible–near-infrared absorption spectra. In this paper, an efficient approach to developing and evaluating tens of thousands of partial least-squares regression (PLSR) models, built from fused optical spectra or multimodal acquisitions, is discussed. Each PLSR model was optimized with unique preprocessing combinations, and features were selected using genetic algorithm filters. The 7-factor D-optimal design training set contained just 55 samples to minimize the number of samples. The performance of PLSR models was evaluated by using an automated latent variable selection script. PLS1 regression models tailored to each species outperformed a global PLS2 model. PLS1 models built using fused spectra data and a multimodal (i.e., analyzed separately) approach yielded similar information, resulting in percent root-mean-square error of prediction values of 0.9–5.7% for the seven factors. Further, the optical techniques and data processing strategies established in this study allow for the direct analysis of numerous species without measuring luminescence lifetimes or relying on a standard addition approach, making it optimal for near-real-time, in situ measurements. Nuclear reactor modeling helped bound training set conditions and identified elemental ratios of lanthanide fission products to characterize the burnup of irradiated nuclear fuel. Leveraging fluorescence, spectrophotometry, experimental design, and chemometrics can enable the remote quantification and characterization of complex systems with numerous species, monitor system performance, help identify the source of materials, and enable rapid high-throughput experiments in a variety of industrial processes and fundamental studies.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Absorption Correction for Reliable Pair Distribution Functions from Low Energy X-ray Sources

This paper explores the development and testing of a simple absorption correction model for processing powder X-ray diffraction data from Debye−Scherrer geometry laboratory X-ray experiments. This may be used as a preprocessing step before using PDFGETX3 to obtain reliable pair distribution functions (PDFs). Various experimental and theoretical methods for estimating μR were explored, and the most appropriate μR values for correction were identified for different capillary diameters and X-ray beam sizes. We identify operational ranges of μR where a reasonable signal-to-noise ratio is possible after correction. A user-friendly software package, DIFFPY.LABPDFPROC, is presented that can help estimate μR and perform absorption corrections with a rapid calculation for efficient processing.

Absorption↗

Multiscale Catalytic Fast Pyrolysis of Grindelia Reveals Opportunities for Generating Low Oxygen Content Bio-Oils from Drought Tolerant Biomass

Grindelia squarrosa (curlycup gumweed) biomass possesses unique biochemistry, cell wall composition, and leaf architecture tailored for prolific growth in arid and semiarid climates. Most notably, this plant has developed high levels of extractable resins that have high effective H/C eff ratios ((mol H - 2 x mol O)/mol C), which is hypothesized to lead to low coke formation during catalytic fast pyrolysis (CFP) over the ZSM-5 catalyst. In microscale experiments with high ZSM-5 loadings (biomass-to-catalyst mass ratio (B/C) ~ 0.1), in situ CFP generated high yields of aromatic hydrocarbons (30% carbon yield) while ex situ CFP favored aliphatic hydrocarbons (25% carbon yield). The difference between the two configurations was attributed to the constant catalyst temperature during ex situ CFP. Deactivation leading to partially deoxygenated vapor products occurred rapidly until B/C ≤ 0.5 by the adsorption of organic species blocking access to acid sites inside the micropores of the catalyst. This was followed by more gradual deactivation leading to primary vapor breakthrough, which we attribute to coke formation on acid sites on the external surface of ZSM-5 crystallites. Noncatalytic fast pyrolysis of Grindelia in a bench scale reactor produced oils with oxygen content (18 wt % on dry basis) and carbon yield (33%) comparable to those of CFP of woody biomass. The CFP of Grindelia further reduced the oxygen content to 7 wt % for in situ CFP and 4 wt % for ex situ CFP at B/C of 2-3. The good deoxygenation was attributed to a combination of a high H/C eff ratio and overall better quality of the pyrolysis vapors that were passed over the ZSM-5 catalyst. The high inorganic content of the Grindelia likely catalyzed pyrolysis to remove oxygenated coke precursors. This integrated CFP study demonstrated that Grindelia could be an important feedstock for generating stabilized noncatalytic and CFP oils for downstream processing into fuels and/or extraction of high-value chemicals. The preprocessing of this feedstock will be required to remove inorganics, which cause an irreversible deactivation of ZSM-5.

BIOMASS FUELS↗