Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “pre-processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Hydrocyclone pre-processing of wastewater algae: A strategy for inorganic ash separation

Microalgae cultivation on wastewater can provide remediation and generate valuable feedstocks for biofuel production. Wastewater algae typically have a high percentage of inorganic ash, which can reduce yield and quality of biocrude produced during hydrothermal liquefaction (HTL). Here, in this work, we evaluated the ability of hydrocyclone pre-processing to remove inorganic ash from wastewater algae. The pH of the algae slurry was adjusted to 9.5 to encourage the formation of precipitates and create a density differential between ash particles and algal cells. Hydrocyclone processing successfully concentrated ash particles in the underflow fraction and reduced the total ash percentage in the overflow fraction. Overall, hydrocyclone processing reduced the total ash by 21%, while only 8% of organics were lost. Elemental and mineral analysis showed that Mg and P were concentrated in the underflow in the form of baricite (an isomorph of vivianite). Future research should focus on improving vivianite and/or baricite formation, and therefore ash removal, by providing a reducing environment. The addition of multiple hydrocyclones in series could also improve the removal of ash. We concluded that hydrocyclone treatment of wastewater algae is a feasible method to remove inorganic ash, but further process optimization is required.

09 - BIOMASS FUELS↗

“One Table to Rule Them All”: How a Single Table can Enable Extensive Insights, Analytics and Assessment on Human Mobility Data

While much research has been conducted in Human Mobility Science, most studies on the analytics/insights part generally focus on one of the following: processing and analytics on human stop-trip behavior, design of individual mobility metrics (often in silos), calculation and characterization of only a handful (typically 5-6) of human mobility metrics on geospatial-temporal human mobility data of interest. Although human mobility research offers a vast and diverse array of available metrics, most individual studies typically compute only a small subset of five or six metrics at a time when analyzing trajectory datasets of human mobility across different areas of interest. This paper is motivated by the critical need to repeatedly compute an extensive array of human mobility metrics across several trajectory datasets and perform individual metric-level benchmarking to establish a new, standardized Test and Evaluation (T&E) suite for the field of Human Mobility Science. We first present our findings on the minimal yet sufficient pre-processing required to reliably and efficiently compute a wide range of human mobility metrics. The key findings are specifically related to the proposed Composite Stop Locations table, which serves as a core pre-processing data layer. Subsequently, we present a case study demonstrating how the Composite Stop Locations table facilitates computation of at least 14 distinct human mobility metrics (unlike 5-6 different set of metrics used for studies in the literature) using the popular and open-source OpenPFLOW dataset. Finally, we have also presented an example of our benchmarking methodology to evaluate the quality and performance of the trajectory dataset of interest, assessed across multiple human mobility metrics.

De, Debraj [ORNL] (ORCID:0000000233630020)↗

Automatic building energy model development and debugging using large language models agentic workflow

Building energy modeling (BEM) is a complex process that demands significant time and expertise, limiting its broader application in building design and operations. While Large Language Models (LLMs) agentic workflow have facilitated complex engineering processes, their application in BEM has not been specifically explored. This paper investigates the feasibility of automating BEM using LLM agentic workflow. Here, we developed a generic LLM-planning-based workflow that takes a building description as input and generates an error-free EnergyPlus building energy model. Our robust workflow includes four core agents: 1) Building Description Pre-Processing, 2) IDF Object Information Extraction, 3) Single IDF Object Generator Suite, and 4) IDF Debugging Agent. These agents divide the complex tasks into manageable sub-steps, enabling LLMs to generate accurate and reliable results at each stage. The case study demonstrates the successful translation of a building description into an error-free EnergyPlus model for the iUnit modular building at the National Renewable Energy Laboratory. The effectiveness of our workflow surpasses: 1) naive prompt engineering, 2) other LLM-based workflows, and 3) manual modeling, in terms of accuracy, reliability, and time efficiency. The paper concludes with a discussion on the interplay between foundational models and LLM agent planning design, advocating for the use of fine-tuned, specialized models to advance this field.

97 MATHEMATICS AND COMPUTING↗

Parallel sorting algorithm classification: is manual instrumentation necessary?

Understanding parallel algorithms is crucial for accelerating scientific simulations on complex, distributed memory, high-performance computers. Modern algorithm classification approaches learn semantics directly from source code to differentiate between algorithms, however, accessing source code is not always possible. We can learn about parallel algorithms from observing their performance, as programs running the same algorithms and using the same hardware should exhibit similar performance characteristics. We present an approach to learn algorithm classes from parallel performance data directly in order to classify algorithms without access to the source code. We extend previous work to enable classifying parallel sorting algorithms using automatic instrumentation instead of requiring manual region annotations in the source code. In this work, we design and demonstrate a study for classification of parallel sorting algorithms using parallel performance data collected from automatic instrumentation, and evaluate the performance of our new methodology on classification. We leverage Caliper to collect the performance data, Thicket for our exploratory data analysis (EDA), and PyTorch and Scikit-learn to evaluate the effectiveness of random forests, support vector machines (SVMs), decision trees, neural networks, and logistic regressions on parallel performance data. Additionally, we study noise in parallel performance data, whether the removal of noise and pre-processing of the data is necessary to accurately classify parallel sorting algorithms, and determine the effectiveness of features created from performance data. In conclusion, we demonstrate classification accuracy for these five different models of up to 97.7% across four different parallel algorithm classes.

Algorithm Classification↗

Multi-scale dynamic modeling and validation of radial flow fixed bed contactors for post-combustion CO 2 capture using bench scale and pilot plant data

Here, in this work, a multi-scale model of a radial flow fixed bed contactor packed with a carbon sorbent is developed and validated with laboratory-scale and pilot plant scale dynamic data. For the lab scale system, the model results were compared with low, medium and high gas and sweep flowrates, yielding root mean square error (RMSE) of 0.80, 0.63, 0.96 CO 2 mol%, respectively, for the outlet CO 2 concentration profile considering the entire A-D cycle. For the bed outer temperature profile, maximum RMSE was found to be 5.5 °C considering all flowrates and entire A-D cycles. An experimental campaign was developed and applied to a pilot plant at Technology Center Mongstad (TCM), Norway. Approaches were developed for pre-processing of data including consideration of the effect of gas mixing, measurement delay, and determination of cyclic steady-state conditions. Considering profiles during A-D cycles for all test runs, it was found that the maximum RMSE for pressure drop, temperature for the outer section of the bed, temperature for the middle section of the bed, and outlet CO 2 concentration profile remained less than 1.5 mbar, 3.5 °C, 2.8 °C, and 1.3 CO 2 mol%, respectively. The validated model was used to perform sensitivity studies on several key design operating variables for the adsorption-desorption cycle. It was found that the flow rate and concentration of flue gas have dominant nonlinear effects on the breakthrough time while the desorption time was strongly affected by the sweep gas flowrate for the specific sorbent being evaluated in this study.

20 FOSSIL-FUELED POWER PLANTS↗

Micro-structural features and material properties impact on adhesive metal joints via computational modeling and machine learning

The quality of structural bonding in practical applications depends on various factors arising from materials, pre-processing conditions, and manufacturing. Understanding how these factors influence bonding performance and determining their relative importance are of significant interest. Thus, this study evaluates the effects of microstructural features and material properties on the structural strength of adhesively-bonded metal joints at the submillimeter scale, utilizing a combination of Finite Element Modeling (FEM) and Machine Learning (ML) with Gradient Boosting Regression (GBR). The microstructural features include adhesive thickness, internal voids within the adhesive, adherend-adhesive interfacial voids, void size and volume fraction, and surface roughness. The material properties include the constitutive behavior of the adhesive, as well as the adherend-adhesive interfacial strength and fracture energy. The changes in structural strength and morphologies of the bonded metal structures with respect to different microstructural features and material properties were clarified by FEM. By further leveraging ML-GBR, the sequence of importance of these factors affecting bonding performance across various scenarios was summarized. This work provides valuable insights into the development of improved structural bonding for adhesive joints in industries such as automotive , aerospace, and beyond.

36 MATERIALS SCIENCE↗

Grain2mesh: A Python and cubit mesh generator from unprocessed mesoscale images

Predicting bulk behavior from microscale features constitutes a key objective in multiscale modeling research, often involving numerical models composed of finite elements that capture the diversity of constituent phases, shapes, and orientations within the material. The Grain2mesh toolbox allows the user to input unprocessed mesoscopic images for automatic segmentation, pre-processing, quality control, and numerical mesh generation. The numerical mesh generation incorporates Cubit routines to generate robust multi-phase mesh structure for use in computational mechanics solvers. The python classes developed contain detailed documentation and examples to support standard usage and case-specific alternative options.

58 GEOSCIENCES↗

Rapid characterization of MSW and RDF feedstocks for waste-to-energy process using LIBS and ML techniques

The heterogeneity in the composition of municipal solid wastes (MSW) poses significant challenges in the production of biofuel and bioproducts. This research aims to enhance the accuracy and efficiency of waste analysis and characterization by introducing a fast characterization approach for MSW-derived refuse-derived fuels (RDF) by combining Laser-Induced Breakdown Spectroscopy (LIBS) with advanced machine learning (ML) techniques. The approach combines data pre-processing of LIBS spectra of RDF, and the development of ML models trained on domain and theory-based spectral features for predicting process parameters. These models are adept at predicting key process parameters like High Heating Value (HHV), carbon content, and volatile matter. This approach can achieve an average RRMSE of 2.13% and R 2 of 0.98 or higher for all considered parameters on testing data. This work demonstrates significant potential for improving waste sorting, processing efficiency, and environmental compliance over traditional labor- and time-intensive laboratory waste analysis and characterization.

09 BIOMASS FUELS↗

nmRanalysis: An Open-Source Web Application for Semi-automated NMR Metabolite Profiling

Though data acquisition and initial signal pre-processing of nuclear magnetic resonance (NMR) spectra have achieved high degrees of automation, downstream processing - specifically the profiling of spectra - has bottlenecked the overall NMR analysis workflow. Several efforts have been made to mitigate this bottleneck, but these solutions often trade an increase in automation for limitations elsewhere. Here, in this technical note, we introduce nmRanalysis, a user-friendly web-application that integrates the strengths of existing profiling tools for a more automated profiling workflow. nmRa-nalysis additionally incorporates novel features, including a machine-learning-driven recommender system for me-tabolite identification, further increasing the utility of nmRanalysis over the individual tools that it incorporates.

Flores, Javier E. [Pacific Northwest National Labo↗

Complementing Dynamical Downscaling With Super‐Resolution Convolutional Neural Networks

Despite advancements in Artificial Intelligence (AI) methods for climate downscaling, significant challenges remain for their practicality in climate research. Current AI-methods exhibit notable limitations, such as limited application in downscaling Global Climate Models (GCMs), and accurately representing extremes. To address these challenges, we implement an AI-based methodology using super-resolution convolutional neural networks (SRCNN), trained and evaluated on 40 years of daily precipitation data from a reanalysis and a high-resolution dynamically downscaled counterpart. The dynamical downscaled simulations, constrained using spectral nudging, enable the replication of historical events at a higher resolution. This allows the SRCNN to emulate dynamical downscaling effectively. Modifications, such as incorporating elevation data and data pre-processing enhances overall model performance, while using exponential and quantile loss functions improve the simulation of extremes. Our findings show SRCNN models efficiently and skillfully downscale precipitation from GCMs. Future work will expand this methodology to downscale additional variables for future climate projections.

54 ENVIRONMENTAL SCIENCES↗

Machine learning pipeline for denoising low signal-to-noise ratio and out-of-distribution transmission electron microscopy datasets

High-resolution transmission electron microscopy (HRTEM) is crucial for observing material’s structural and morphological evolution at Angstrom scales, but the electron beam can alter these processes. Devices such as CMOS-based direct-electron detectors operating in electron-counting mode can be utilized to substantially reduce the electron dosage. However, the resulting images often lead to a low signal-to-noise ratio, which requires frame integration that sacrifices temporal resolution. Several machine learning (ML) models have been recently developed to successfully denoise HRTEM images. Yet, these models are often computationally expensive, and their inference speeds on GPUs are outpaced by the imaging speed of advanced detectors, precluding in situ analysis. Furthermore, the performance of these denoising models on datasets with imaging conditions that deviate from the training datasets has not been evaluated. To mitigate these gaps, we propose a new self-supervised ML denoising pipeline specifically designed for time-series HRTEM images. This pipeline integrates a blind-spot convolution neural network with pre-processing and post-processing steps, including drift correction and low-pass filtering. Results demonstrate that our model outperforms various other ML and non-ML denoising methods in noise reduction and contrast enhancement, leading to improved visual clarity of atomic features. Additionally, the model is drastically faster than U-Net-based ML models and demonstrates excellent out-of-distribution generalization. The model’s computational inference speed is in the order of milliseconds per image, rendering it suitable for application in in-situ HRTEM experiments.

36 MATERIALS SCIENCE↗

The gas-phase mass–metallicity relation of dwarf galaxies across large-scale environments using the CAVITY parent sample

Context. The gas-phase mass–metallicity relation (MZR) of galaxies shows a noticeable break in slope and an increased scatter at low stellar masses, suggesting that the physical processes governing chemical enrichment differ between dwarf and high-mass systems. Dwarf galaxies, in particular, are highly susceptible to both internal and environmental mechanisms due to their shallow potential wells. Aims. The primary aim of this work is to assess whether a single, universal MZR can describe dwarf galaxies across diverse large-scale environments, or whether systematic environmental variations emerge. To probe these, we examine the MZR and star formation rate (SFR) of dwarf galaxies with stellar masses in the range of 8.9 < log(M ★ /M ⊙ ) < 9.5. Methods. Using optical spectra from the Sloan Digital Sky Survey, we measured the fluxes of key emission lines via the pyPipe3D full spectral fitting pipeline. Aperture-corrected fluxes, along with multiple metallicity indicators and calibrations, were used to derive the MZR and the SFR for 353, 311, and 22 dwarf galaxies located in voids, filaments, and clusters, respectively. Results. We find a systematic variation in the MZR slope, steeper in voids (0.28 ± 0.03) and progressively flatter in clusters (0.17 ± 0.08), indicating a dependence of the MZR on the large-scale environment in this mass regime. When galaxies are separated by local density, no significant differences are observed between isolated and non-isolated dwarfs in voids. Isolated dwarf galaxies in filaments also exhibit properties similar to those of their counterparts in voids. However, non-isolated filament galaxies exhibit similar MZR slopes comparable to those of cluster dwarfs and flatter slopes than their counterparts in voids. Conclusions. We report both large- and local-scale environmental dependencies in the gas-phase metallicity and in the slope of the MZR for dwarf galaxies. Consistent with the general consensus on the pre-processing of galaxies in filaments, our results indicate that the influence of the local environment becomes increasingly significant within the filamentary regions of the cosmic web, affecting the chemical enrichment and star formation activity of low-mass systems. These findings further suggest that a portion of the scatter commonly observed in the MZR of dwarf galaxies arises from environmental effects.

Bidaran, Bahar [Dpto. de Física Teórica y del Cosm↗

Decoding diffraction and spectroscopy data with machine learning: A tutorial

This Tutorial provides a step-by-step guide on how to apply supervised machine-learning techniques to analyze diffraction and spectroscopy data. This Tutorial details four models—a reconstruction-focused model, a regression-focused model, a hybrid reconstruction/regression model, and a multimodal model—that use x-ray diffraction profiles and vibrational density of states spectra to predict various microstructural descriptors. In this Tutorial, we cover data pre-processing steps, constructions of the models via dimensionality reduction and regression, training, and analysis of these models. Comparisons of the model’s performance are provided, highlighting the strength and weakness of the various approaches utilized.

36 MATERIALS SCIENCE↗

Sharp detection of low-dimensional structure in probability measures via dimensional logarithmic Sobolev inequalities

Identifying low-dimensional structure in high-dimensional probability measures is an essential pre-processing step for efficient sampling. To identify this structure, we approximate the target measure as a perturbation of an arbitrary reference measure along a few directions in $\mathbb{R}^{d}$. These directions are determined by minimizing an upper bound on the Kullback–Leibler (KL) divergence between the target and its approximation. Our contribution improves upon previous works by leveraging dimensional logarithmic Sobolev inequalities to refine the bound on the KL divergence. These inequalities lead to a uniformly tighter bound on the KL divergence, thereby enhancing the identification of the most significant perturbation directions. In particular, when the target and reference are both Gaussian, minimizing the resulting bound is equivalent to minimizing the KL divergence. We further demonstrate the applicability of this analysis to the squared Hellinger distance, where analogous reasoning shows that the dimensional Poincaré inequality offers improved bounds.

Bayesian inference↗

Auriga Streams – I: disrupting satellites surrounding Milky Way-mass haloes at multiple resolutions

In a hierarchically formed Universe, galaxies accrete smaller systems that tidally disrupt as they evolve in the host’s potential. We present a complete catalogue of disrupting galaxies accreted onto Milky Way-mass haloes from the Auriga suite of cosmological magnetohydrodynamic zoom-in simulations. We classify accretion events as intact satellites, stellar streams, or phase-mixed systems based on automated criteria calibrated to a visually classified sample, and match accretions to their counterparts in haloes re-simulated at higher resolution. Most satellites at the present day have lost substantial amounts of stellar mass – 67 per cent have $f_\text{bound} < 0.97$ (our threshold of lost stellar mass to no longer be considered intact), while 53 per cent satisfy a more stringent $f_\text{bound} < 0.8$. Streams typically outnumber intact systems, contribute a smaller fraction of overall accreted stars, and are substantial contributors at intermediate distances from the host centre ($\sim$0.1 to $\sim 0.7R_\text{200m}$, or $\sim$35 to $\sim$250 kpc for the Milky Way). We also identify accretion events that disrupt to form streams around massive intact satellites instead of the main host. Streams are more likely than intact or phase-mixed systems to have experienced pre-processing, suggesting this mechanism is important for setting disruption rates around Milky Way-mass haloes. All of these results are preserved across different simulation resolutions, though we do find some hints that satellites disrupt more readily at lower resolution. The Auriga haloes suggest that disrupting satellites surrounding Milky Way-mass galaxies are the norm and that a wealth of tidal features waits to be uncovered in upcoming surveys.

galaxies: haloes↗

LaueMatching: an approach for rapid and robust indexing of Laue diffraction patterns

Traditional Laue diffraction pattern indexing often struggles with noisy data, weak signals, peak overlap and missing reflections, particularly from complex or deformed microstructures. Here, we introduce LaueMatching, a high-throughput indexing algorithm designed to overcome these limitations. LaueMatching utilizes a fundamentally different approach based on direct pattern correlation: experimentally pre-processed images are compared against a comprehensive pre-computed library of simulated diffraction patterns corresponding to a dense grid of possible orientations. This approach bypasses the need for explicit peak identification and fitting, steps that are often a failure point for traditional methods. The algorithm rapidly and robustly indexes multiple crystallographic orientations and crystal systems simultaneously, even from challenging patterns. LaueMatching's effectiveness and accuracy have been rigorously tested and validated on diverse experimental (Ni, Al, EuAl 2 O 4 ) and simulated diffraction patterns, demonstrating high-fidelity orientation refinement. Code to implement this approach on both CPU and GPU resources can be downloaded from https://github.com/AdvancedPhotonSource/LaueMatching.

36 MATERIALS SCIENCE↗

Descriptor: High Temporal Resolution Meteorological Data at Oak Ridge Reservation (ORR-HiResMet)

Access to continuous, quality assessed meteorological data is critical for understanding the climatology and atmospheric dynamics of a region. Research facilities like Oak Ridge National Laboratory (ORNL) rely on such data to assess site-specific climatology, model potential emissions, establish safety baselines, and prepare for emergency scenarios. To meet these needs, on-site towers at ORNL collect meteorological data at 15-minute and hourly intervals. However, data measurements from meteorological towers are affected by sensor sensitivity, degradation, lightning strikes, power fluctuations, glitching, and sensor failures, all of which can affect data quality. To address these challenges, we conducted a comprehensive quality assessment and processing of five years of meteorological data collected from ORNL at 15-minute intervals, including measurements of temperature, pressure, humidity, wind, and solar radiation. The time series of each variable was pre-processed and gap-filled using established meteorological data collection and cleaning techniques, i.e., the time series were subjected to structural standardization, data integrity testing, automated and manual outlier detection, and gap-filling. The data product and highly generalizable processing workflow developed in Python Jupyter notebooks are publicly accessible online. As a key contribution of this study, the evaluated 5-year data will be used to train atmospheric dispersion models that simulate dispersion dynamics across the complex ridge-and-valley topography of the Oak Ridge Reservation in East Tennessee.

Steckler, Morgan R. [Oak Ridge National Laboratory↗