Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “preprocessed”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Nth-plant scenario for forest resources and short rotation woody crops: Biorefineries and depots in the contiguous US

Estimating the US potential of woody material is of vital importance to ensure cost-effective supply logistics and develop a sustainable bioenergy and bioproducts industry. We analyzed a mature conversion technology for woody resources for the contiguous US that takes advantage of economies of scale: the nth-plant. Here, we developed a database to quantify the total accessible woody biomass within a distributed network of preprocessing depots and biorefineries considering both quality specifications for conversion and a target cost to compete with fossil fuels. We considered two categories of woody biomass: 1) forest residues from trees, tops and limbs produced from conventional thinning and timber harvesting operations as well as non-timber tree removal; and 2) short rotation woody crops such as poplar, willow, pine, and eucalyptus. A mixed integer linear programming model was developed to analyze scenarios with woody feedstock blends at variable biomass ash contents and cost targets at the biorefinery. When considering a target cost of 85.51 dollars/dry ton (2016$) at the biorefinery, the maximum accessible biomass from forest residues in 2040 remained constant at 106 million dry tons regardless of ash targets. Including short rotation woody crops as part of the blend increased the total accessible biomass to 153 and 195 million dry tons at ash targets of 1% and 1.75%, respectively. We concluded from our analysis that woody resources could address about 55% of EPA’s (Environmental Protection Agency) target of 16 billion gallons of cellulosic biofuel.

09 BIOMASS FUELS↗

Biomass supply chain equipment for renewable fuels production: A review

The production of renewable fuels is a critical component of global strategies to reduce greenhouse gas (GHG) emissions. Moreover, the collection of raw materials for its production can provide added benefits such as reduction of wildfire risk, additional income for farmers, and decreased disposal costs. Although there is substantial literature on design and modeling of supply chains, the authors were unable to find a single reference with the information needed for the selection and cost estimation of each type of equipment involved in the supply chain. Therefore, the goal of this research is to gather information necessary for the construction and utilization of models that might drive the identification of a feasible supply chain to produce renewable fuels at a commercial scale. The primary objectives are to 1) understand the supply chain of critical feedstocks for renewable fuels production; 2) identify the equipment commercially available for collection and adequation of feedstock; and 3) consolidate information regarding equipment cost, energy consumption, and efficiency, as well as feedstock storage and transportation systems. This paper provides a compilation for five feedstock types studied for sustainable aviation fuel production: 1) agricultural residues and grasses, 2) forest residues, 3) urban wood waste, 4) oilseeds, 5) fats, oils & greases. All the technologies involved from the field to the gate of the preprocessing or conversion unit were reviewed. The information on fats, oils & greases supply chains and equipment purposely designed for forest thinning and pruning was very limited.

: Feedstock, Collection and Adequation, Renewable ↗

Enhancing biomass flowability for entrained flow Gasification: The role of densification and torrefaction

Gasification presents a key strategy in addressing future energy demands while minimizing environmental impact. This has been recognized as a promising method to convert biomass to higher value products such as biofuels or hydrogen. Among gasification technologies, high-temperature and high-pressure reactors, particularly the R-GAS® system, emerge as an advanced option boasting superior conversion efficiency. However, akin to conventional high-temperature and high-pressure gasifiers, R-GAS® necessitates small particle sizes for optimal carbon conversion, a requirement yet to be fully explored for biomass. Hence, this study investigated the effectiveness of combined mechanical and thermal preprocessing techniques in modifying the physicochemical properties of biomass to suit gasification systems. Mechanical techniques including densification and pulverization, alongside thermal techniques such as torrefaction and steam explosion, were examined. The results demonstrate that torrefaction fosters producing of uniform granular material, enhancing flowability and reducing energy requirements for pulverization compared to steam explosion. Notably, torrefied corn stover exhibited lower internal friction angles and effective cohesion (40.09 ± 0.22° and 0.56 ± 0.01 kPa, respectively) compared to steam exploded corn stover (41.87 ± 0.65° and 0.83 ± 0.06 kPa, respectively), indicative of improved flowability. Additionally, pulverization of torrefied corn stover required approximately 16 % less energy than steam exploded corn stover and 91 % less energy than raw corn stover. Furthermore, the torrefaction-induced alterations in particle size, shape, and packing densities emphasize its potential to optimize flow and handling processes for gasification. These findings underline that densification followed by torrefaction effectively addresses biomass variability, leading to more efficient and sustainable energy conversion.

09 - BIOMASS FUELS↗

A guideline to document occupant behavior models for advanced building controls

The availability of computational power, and a wealth of data from sensors have boosted the development of model-based predictive control for smart and effective control of advanced buildings in the last decade. More recently occupant-behavior models have been developed for including people in the building control loops. However, while important objectives of scientific research are reproducibility and replicability of results, not all information is available from published documents. Therefore, the aim of this paper is to propose a guideline for a thorough and standardized occupant-behavior model documentation. For that purpose, the literature screening for the existing occupant behavior models in building control was conducted, and the occupant behavior modeling processes were studied to extract practices and gaps for each of the following phases: problem statement, data collection, and preprocessing, model development, model evaluation, and model implementation. Here, the literature screening pointed out that the current state-of-the-art on model documentation shows little unification, which poses a particular burden for the model application and replication in field studies. In addition to the standardized model documentation, this work presented a model-evaluation schema that enabled benchmarking of different models in field settings as well as the recommendations on how OB models are integrated with the building system.

Building control↗

Machine learning–assisted prediction of heat fluxes through thermally anisotropic building envelopes

Thermally anisotropic building envelope (TABE) is a novel active building envelope that can save energy use to maintain thermal comfort in buildings by redirecting heat and coolness from building envelopes to thermal loops. Finite element models (FEMs) can be used to compute the heat fluxes through TABEs, but the high computational cost of finite element simulations has prevented parametric studies and design optimizations. This paper proposes a domain knowledge–informed, finite element–based machine learning framework to reduce the computation cost for the energy management of buildings installed with TABE that uses a ground thermal loop. First, the training heat flux data set was generated by FEM simulations with different thermal loop schedules. Then, both shallow learning models (i.e., multivariate linear regression and eXtreme Gradient Boost, or XGBoost) and a deep learning model (i.e., deep neural network, or DNN) were trained to predict the heat fluxes. Domain knowledge was used for data preprocessing and feature selection. Finally, the suitability of the selected machine learning model was tested under different thermal loop schedules. Herein, the case study results showed that: (1) XGBoost can be as accurate as DNN (coefficient of determination equal to 0.81) with much less training time; (2) the annual energy cost savings for different thermal loop schedules obtained by the XGBoost-predicted and FEM-calculated heat fluxes are consistent, having a difference of only 4%; and (3) XGBoost can reduce the computation time for the annual energy analysis of the case study building with a given thermal loop schedule from around 12 h by using FEM to less than 1 min.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

GMT: A deep learning approach to generalized multivariate translation for scientific data analysis and visualization

In scientific visualization, despite the significant advances of deep learning for data generation, researchers have not thoroughly investigated the issue of data translation. We present a new deep learning approach called generalized multivariate translation (GMT) for multivariate time-varying data analysis and visualization. Like V2V, GMT assumes a preprocessing step that selects suitable variables for translation. However, unlike V2V, which only handles one-to-one variable translation during training and inference, GMT enables one-to-many and many-to-many variable translation in the same framework. We leverage the recent StarGAN design from multi-domain image-to-image translation to achieve this generalization capability. We experiment with different loss functions and injection strategies to explore the best choices and leverage pre-training for performance improvement. We compare GMT with other state-of-the-art methods (i.e., Pix2Pix, V2V, StarGAN). Furthermore, the results demonstrate the overall advantage of GMT in translation quality and generalization ability.

97 MATHEMATICS AND COMPUTING↗

Predicting oxidation damage of ultra high-temperature carbide ceramics in extreme environments using machine learning

Determining the oxidation resistance of UHTC carbides in extreme environments is challenging theoretically and experimentally due to the high dimensional complexity of influencing variables and intricate testing setups. Herein we demonstrate the use of machine learning (ML) models trained with experimental literature data to predict the oxide thickness of UHTC carbides exposed to air based on composition, mean grain size, relative densification, holding time, and temperature. A multi-dimensional database with 76 occurrences is created containing experimental results of Hf, Zr, and Ta carbides plus additives. In this study, the preprocessed database is then used to train ML models to predict their oxidation behavior. The trained model predicts the oxidation damage in the form of an average oxide thickness in UHTC carbides with a Mean Absolute Error (MAE) of ±65.45 μm for samples in the testing set that developed thicknesses up to 1000 μm. The model successfully predicted oxidation damage for a recession rate lower than 60 μm/min. It is noticed that the ensemble method MAE is increased to ±134.34 μm while forecasting the oxidation of samples with a recession rate higher than the threshold. The unprecedented approach is a novel way to predict the damage through the oxidation of carbide compounds before processing for a smarter design with room for improvement.

36 MATERIALS SCIENCE↗

Grain boundary structure search by using an evolutionary algorithm with effective mutation methods

Grain boundaries (GBs) accommodate the misorientation between adjacent grains in a polycrystalline material. GBs are geometrically described by the macroscopic and microscopic degrees of freedom. Besides, at the atomistic level, GBs exhibit complicated behaviors under varying thermodynamic conditions. The complexity of atomistic GB structures demands stochastic searching for possible states. The effectiveness of stochastic search methods relies on techniques to recreate and select atomistic structures. In this work, we developed a new mutation operator that can induce direct and collective atomistic structure changes to boost the search efficiency of exploring GB structures with evolutionary algorithms (EA). We implemented the mutation methods along with innovative selection, crossover, boundary condition preprocessing methods to form an EA-based package to explore GB structures in grand canonical ensembles with atomistic simulations. We used this package to study the [001] symmetric tilt grain boundaries (STGBs) in FCC copper (Cu), the [110] STGBs in BCC tungsten (W), and the $[1\bar{2}10]$ STGBs in HCP magnesium (Mg). The results show that our design and implementation based on new mutation procedures, selection, and boundary conditions provide a high-quality search of atomistic GB structures in the grand canonical ensemble for different crystal lattices.

36 MATERIALS SCIENCE↗

Optical image analysis of WSe 2 − thresholding for layer detection

The fast and reliable layer identification of two-dimensional transition metal dichalcogenide (TMD), such as WSe 2 , is essential to investigating their thickness-dependent electronic and optical properties. This article presents efficient optical image thresholding methodology designed to segment the mono, bi, and tri-layer regions of WSe 2 flakes mechanically exfoliated onto a SiO 2 /Si substrate. The optical images were first preprocessed to exclude the background effect and analyzed using the pixel medians and interquartile ranges for fundamental color channels—red, green, and blue (RGB). The analysis of red channel pixel intensities yielded three distinct ranges, serving as thresholds for layer segmentation: monolayer (111.0–118.0), bilayer (103.0–110.0), and tri-layer (93.0–103.0). Similarly, thresholds were established for each color channel, facilitating a comparative study of the segmentation performances. Further, the intersection-over-union ($IoU$) calculations revealed that the red and green channels demonstrated greater than 99 % and 90 % accuracy in differentiating each layer, respectively. This approach yields remarkable results without substantial data calibration that utilizes time-intensive heuristic techniques. Moreover, the proposed methodology offers the flexibility to compare performances across different color channels, expanding the applicability for other 2D material systems.

2D Materials↗

Adaptive continuity-preserving simplification of street networks

Street network data is widely used to study human-based activities and urban structure. Often, these data are geared towards transportation applications, which require highly granular, directed graphs that capture the complex relationships of potential traffic patterns. While this level of network detail is critical for certain fine-grained mobility models, it represents a hindrance for studies concerned with the morphology of the street network. For the latter case, street network simplification — the process of converting a highly granular input network into its most simple morphological form — is a necessary, but highly tedious preprocessing step, especially when conducted manually. In this manuscript, we develop and present a novel adaptive algorithm for simplifying street networks that is both fully automated and able to mimic results obtained through a manual simplification routine. The algorithm — available in the neatnet Python package — outperforms current state-of-the-art procedures when comparing those methods to manually, human-simplified data, while preserving network continuity.

Python↗

Pavement condition and climatic data in southeast Texas: A dataset for evaluating flood impacts on pavement performance

Effective pavement maintenance is essential for economic stability, optimal network performance, and roadway safety. Achieving this requires thorough evaluation of pavement conditions, including structural integrity, surface roughness, and distress characteristics. Pavement performance indicators play a critical role in influencing vehicle safety and ride quality. Recent advances have emphasized the use of data-driven modeling to anticipate pavement behavior, with the goal of optimizing resource allocation and refining Maintenance and Rehabilitation (M&R) strategies through accurate condition assessment. A foundational requirement for these modeling efforts is the availability of standardized, high-quality datasets that can support robust and reproducible infrastructure analysis. This data article presents a comprehensive dataset assembled to facilitate pavement performance prediction, with a geographic focus on Southeast Texas, particularly the flood-vulnerable area of Beaumont. The dataset encompasses pavement and traffic attributes, meteorological records, flood simulation outputs, ground deformation measurements, and topographic indices, enabling detailed examination of both load-associated and non-load-associated degradation mechanisms. Data preprocessing was performed using ArcGIS Pro, Microsoft Excel, and Python to ensure consistency and usability in data-driven modeling applications, including machine learning workflows. Key contributions of this dataset include its utility in analyzing the climatic and environmental factors affecting pavement conditions, identifying critical predictive features, and enabling in-depth correlation analysis across diverse variables. By filling existing gaps in input variable selection resources, this dataset supports the development of predictive tools for estimating future maintenance demand and enhancing the resilience of pavement networks in flood-impacted areas. The resource highlights the importance of standardized datasets for advancing pavement management practices and provides a robust foundation for ongoing infrastructure performance modeling.

42 ENGINEERING↗

A portable application framework for energy management and information systems (EMIS) solutions using Brick semantic schema

This paper introduces a portable framework for developing, scaling and maintaining energy management and information systems (EMIS) applications using an ontology-based approach. Key contributions include an interoperable layer based on Brick schema, the formalization of application constraints pertaining metadata and data requirements, and a field demonstration. The framework allows for querying metadata models, fetching data, preprocessing, and analyzing data, thereby offering a modular and flexible workflow for application development. Its effectiveness is demonstrated through a case study involving the development and implementation of a data-driven anomaly detection tool for the photovoltaic systems installed at the Politecnico di Torino, Italy. During eight months of testing, the framework was used to tackle practical challenges including: (i) developing a machine learning-based anomaly detection pipeline, (ii) replacing data-driven models during operation, (iii) optimizing model deployment and retraining, (iv) handling critical changes in variable naming conventions and sensor availability (v) extending the pipeline from one system to additional ones.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Optimization of harvest and logistics for multiple lignocellulosic biomass feedstocks in the northeastern United States

A mixed-integer linear programming model was developed to optimize the multiple biomass feedstock supply chains, including feedstock establishment, harvest, storage, transportation, and preprocessing. The model was applied for analyses of multiple biomass feedstocks at county level for 13 states in the northeastern United States. In the base case with a demand of 180,000 dry Mg/year of biomass, the delivered costs ranged from $\$67.90$ to $\$86.97$ per dry Mg with an average of $\$79.58$ /dry Mg. The biomass delivered costs by county were from $\$67.90$ to 150.81 per dry Mg across the northeastern U.S. Considered the entire study area, the delivered cost averaged $\$85.30$ /dry Mg for forest residues, $\$84.47$ /dry Mg for hybrid willow, $\$99.68$ for switchgrass and $\$97.87$ per dry Mg for Miscanthus. Seventy seven out of 387 counties could be able to deliver biomass at $84 per dry Mg or less a target set by US DOE by 2022. A sensitivity analysis was also conducted to evaluate the effects of feedstock availability, feedstock price, moisture content, procurement radius, and facility demand on the delivered cost. Lastly, our results showed that procurement radius, facility capacity, and forest residue availability were the most sensitive factors affecting the biomass delivered costs.

09 BIOMASS FUELS↗

Exploratory analysis and performance prediction of big data transfer in High-performance Networks

Big data transfer in large-scale scientific and business applications is increasingly carried out over connections with guaranteed bandwidth provisioned in High-performance Networks (HPNs) via advance bandwidth reservation. Provisioning agents need to carefully schedule data transfer requests, compute network paths, and allocate appropriate bandwidths. Such reserved bandwidths, if not fully utilized, could be simply wasted due to the exclusive access during the approved time window, and cause extra overhead and complexity for resource management. This calls for accurate performance prediction to reserve bandwidths that match actual needs and avoid over-provisioning. We employ machine learning algorithms to predict big data transfer performance based on extensive performance measurements collected in the past several years from data transfer tests using different protocols and toolkits between various end sites on several real-life physical or emulated testbeds. We first analyze the performance patterns in response to a comprehensive list of parameters in end-host systems, network connections, and data transfer applications, which motivate the use of machine learning and also help us identify the effects of latent factors. We then propose threshold- and clustering-based methods to eliminate negative effects of latent factors in data preprocessing and build a robust performance predictor based on customized domain-oriented loss functions. The performance of the proposed methods is verified by extensive experiments using SVR and RFR as well as theoretical analysis of the general performance bound.

97 MATHEMATICS AND COMPUTING↗

Spatial and temporal characterization of municipal solid waste based on resource recovery pathways

This study presents a two-year, quarterly assessment of MSW across four source sectors (residential, schools, restaurants, and grocery stores) from sixteen sites across five U.S. states. MSW was manually sorted into 27 categories and aggregated into pathway fractions: high-moisture (HM) organics, low-moisture (LM) organics, recyclable (RC) materials, and residuals for disposal. Organics represented 89 % of the MSW stream. The largest fraction was HM organics consisting of food waste (31 %) and yard waste (3 %), with large coefficient of variations (CV), 79 and 278 %, respectively, reflecting high seasonal and site variability that varied significantly (p < 0.01) across sampling periods. The HM fraction showed properties favorable for anaerobic digestion, with moisture content ranging from 56 to 95 % and volatile solids ranges of 86-95 %. In contrast, the LM and RC fractions remained more stable (plastics CV = 41 %; paper CV = 53 %) with heating values up to 26.9 MJ/kg across sources, reflecting suitability for gasification. Microstructural analysis revealed less porosity in residential waste sampled at the landfill, which can influence preprocessing efficiency and microbial accessibility. Pathway informed allocations showed that 35 % of MSW is suitable for anaerobic digestion, 36 % for gasification, and 18 % for recycling, leaving 11 % requiring landfill disposal. These results provide quantitative evidence to determine feedstock allocation, waste-to-energy system design, and the development of data-driven sustainability and resource recovery strategies within a circular bioeconomy.

09 BIOMASS FUELS↗

In-depth analysis on parallel processing patterns for high-performance Dataframes

The Data Science domain has expanded monumentally in both research and industry communities during the past decade, predominantly owing to the Big Data revolution. Artificial Intelligence (AI) and Machine Learning (ML) are bringing more complexities to data engineering applications, which are now integrated into data processing pipelines to process terabytes of data. Typically, a significant amount of time is spent on data preprocessing in these pipelines, and hence improving its efficiency directly impacts the overall pipeline performance. The community has recently embraced the concept of Dataframes as the de-facto data structure for data representation and manipulation. However, the most widely used serial Dataframes today (R, pandas) experience performance limitations while working on even moderately large data sets. We believe that there is plenty of room for improvement by taking a look at this problem from a high-performance computing point of view. In a prior publication, we presented a set of parallel processing patterns for distributed dataframe operators and the reference runtime implementation, Cylon. In this paper, we are expanding on the initial concept by introducing a cost model for evaluating the said patterns. Furthermore, we evaluate the performance of Cylon on the ORNL Summit supercomputer.

97 MATHEMATICS AND COMPUTING↗

Expanded analysis of machine learning models for nuclear transient identification using TPOT

Industries around the world are becoming more and more data driven. The nuclear field is no exception with several different applications being proposed. One popular area of research is the use of machine learning in transient detection. This paper seeks to build upon a previous study which made use of the AutoML package TPOT to train traditional machine learning models to classify transient events occurring with a reactor. Synthetic data was once again collected using a GPWR reactor simulator. Data on 12 different events was collected using 15 different initial conditions. Here, a dataset consisting of over 100,000 data points was compiled and used to train 7 different machine learning models using a pre-defined TPOT dictionary with 12 different preprocessing techniques. Three of the trained models were able to produce validation results in the 90s with the expanded dataset. Once the models were trained, it was possible to look into where during the simulation, misclassifications occurred. Using these three models, analysis was done to determine if TPOT could be used to train models that were effective if important features were missing. The results from this were positive with the newly trained models scoring close to the original models. Finally, to conclude this study, the three high performing models were retrained using different random states to see if there was any major variation when different states were used.

42 ENGINEERING↗

A generalized machine learning workflow to visualize mechanical discontinuity

Accurate detection and mapping of mechanical discontinuity in materials has widespread industrial and research applications. Herein, we developed a generalized machine-learning framework for visualizing single mechanical discontinuity embedded in material of any composition, velocity, density, porosity, and size with limited data. The proposed visualization of discontinuity requires accurate estimations of the length, location, and orientation of the embedded discontinuity by processing multipoint wave-transmission measurements. k-Wave simulator is used to create a large dataset of elastic waveforms recorded during multi-point wave-transmission measurements through materials containing single mechanical discontinuity. k-Wave simulator considers the wave attenuation, dispersion, and mode conversion in wave motion. Discrete wavelet transform (DWT) and statistical feature extraction are essential for data preprocessing prior to the data-driven model development. DWT also minimizes the effect of noise. Using hyper-parameter tuning and cross validation, gradient boosting regression can visualize the mechanical discontinuity with an accuracy of 0.85, in terms of coefficient of determination. A double-layered neural network-based regression has better performance with an accuracy of 0.95. Use of convolutional neural network converts the predictive task from a waveform processing to an image processing problem. Convolutional neural network achieved a generalization performance of 0.91. The proposed generalized workflow requires robust simulation of wave propagation, signal processing, feature engineering, and model evaluation. Sensors closest to the source and those located opposite the source are the most significant for the desired visualization. Notably, the sensors closest to the source capture the non-linear associations, whereas the sensor on the border opposite to the source capture the linear associations between the measured waveforms and the properties of the mechanical discontinuity.

42 ENGINEERING↗