Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “preprocessed data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Nonnegative canonical tensor decomposition with linear constraints: nnCANDELINC

Abstract There is an emerging interest for tensor factorization applications in big‐data analytics and machine learning. To speed up the factorization of extra‐large datasets, organized in multidimensional arrays (also known as tensors), easy to compute compression‐based tensor representations, such as, Tucker and tensor train formats, are used to approximate the initial large‐tensor. Further, tensor factorization is used to extract latent features that can facilitate discoveries of new mechanisms and signatures hidden in the data, where the explainability of the latent features is of principal importance. Nonnegative tensor factorization extracts latent features that are naturally sparse and parts of the data, which makes them easily interpretable. However, to take into account available domain knowledge and subject matter expertise, often additional constraints need to be imposed, which lead us to canonical decomposition with linear constraints (CANDELINC), a canonical polyadic decomposition with rank deficient factors. In CANDELINC, Tucker compression is used as a preprocessing step, which lead to a larger residual error but to more explainable latent features. Here, we propose a nonnegative CANDELINC (nnCANDELINC) accomplished via a specific nonnegative Tucker decomposition; we refer to as minimal or canonical nonnegative Tucker. We derive several results required to understand the specificity of nnCANDELINC, focusing on the difficulties of preserving the nonnegative rank of a tensor to its Tucker core and comparing the real valued to nonnegative case. Finally, we demonstrate nnCANDELINC performance on synthetic and real‐world examples.

97 MATHEMATICS AND COMPUTING↗

Woody Feedstock 2022 State of Technology Report

The U.S. Department of Energy promotes production of advanced liquid transportation fuels from lignocellulosic biomass by funding fundamental and applied research that advances the state of technology (SOT). As part of its involvement in this mission, Idaho National Laboratory completes an annual SOT report for n th -plant and 1 st -plant woody biomass feedstock logistics. The purpose of the SOT is to provide the status of feedstock supply system technology development for woody biomass to biofuels relative to technical targets and cost goals from specific design cases, based on data and experimental results. Conventional feedstock supply systems need to be modified to meet the demands of conversion pathways, specifically to have the ability to adjust the quality of the raw biomass materials. Advanced systems incorporate innovative methods of material handling, preprocessing and supply chain configuration. In advanced designs, variability of the raw biomass can be reduced to produce feedstocks of a uniform format, moving toward biomass commoditization. Against this backdrop, the 2022 Woody SOT for low-ash woody feedstocks utilizes feedstock fractionation by incorporating technologies that can separate the biomass into its anatomical fractions (wood, bark, needle, and extrinsic ash) to reduce impurities and attempt to maximize the retention of usable fractions that satisfy downstream quality considerations. By using a series of air classification steps, this strategy can reduce the extrinsic ash in forest residues, separate out a majority of the incoming needles (which can be supplied to alternate markets), and maximize the retention of whitewood in the usable fraction. The fractionated forest residues are then mixed with clean-pine chips in a 50-50 blend to prepare the feedstock for the desired conversion pathway. The n th -plant analysis estimated the delivered cost for the feedstock at $\$$69.23/dry ton (2016$\$$) which represents a $\$$6.64/dry ton decrease compared to the cost estimate of the 2021 Woody SOT supply system for low-ash woody feedstocks. The quality requirements in the 2022 Woody SOT were identical to those of the 2021 Woody SOT at = 1.00 wt % ash and = 50.51 wt% carbon. The cost savings derive primarily from reductions in dry matter losses during air classification. The GHG emissions for the n th -plant analysis were estimated at 178.39 kg CO2e/dry ton compared to 178.71 kg CO 2 e/dry ton in the 2021 Woody SOT, a decrease of 0.32 kg CO2e/dry ton. The small change stems from an increase in emissions attributed to preprocessing and slightly larger savings in emissions from transportation. In the 1 st -plant analysis of the 2022 Woody SOT system, the average throughput was estimated to be approximately 2,128 dry tons/day or 96.51% of the name plate capacity. During the simulation the daily throughput ranged from 1,090 dry tons/day to 2,200 dry tons/day, or 49.43% to 99.75% of the daily nameplate capacity. After the year of operation 722,403 tons of processed feedstock were produced in total without regard to quality considerations (99.64% of the annual nameplate capacity). The variability in throughput was primarily caused by equipment failures in the system. Regular failures, downtime caused by routine maintenance per manufacturer guidelines, contributed to a majority 62.50% of failures and 62.60% of downtime. Failures due to wear were the other cause of disruption within the system, impacting the rotary shear and orbital screen and accounting for 37.50% of the failures and 37.40% of the total downtime. Ultimately the system was on stream for 87.84% during the simulation period, which is only 2.16 percentage points below the nth-plant assumption for on-stream time. The production cost of the system averaged $\$$71.66/dry ton. The costs ranged from a minimum of $\$$71.23/dry ton to a maximum of $\$$2,115.30/dry ton. When dry matter losses (disposed low-quality fractions as well as other losses such as in grinders) were considered the costs increased to an average of $\$$75.11/dry ton with a minimum of $\$$74.69/dry ton and a maximum of $\$$2,136.86/dry ton...

09 BIOMASS FUELS↗

Herbaceous Feedstock 2018 State of Technology Report

The U.S. Department of Energy (DOE) promotes the production of advanced liquid transportation fuels from lignocellulosic biomass by funding fundamental and applied research that advances the state of technology (SOT). As part of its involvement with this mission, Idaho National Laboratory (INL) completes an annual SOT report for biomass feedstock logistics. This report summarizes supply system impacts of Bioenergy Technologies Office (BETO)-funded research and development efforts at INL and elsewhere (such as the High-Tonnage Feedstock Logistics projects (Webb et al. 2013a, Webb et al. 2013b, Webb et al. 2013c, Webb and Sokhansanj 2014, Sokhansanj et al. 2014) that lead to improvements in feedstock supply systems. These include improvements to and observed performance of innovative harvest and collection methods, storage technologies, transportation and handling approaches, and advanced preprocessing technologies. Biomass quality and variability, and the interface between feedstock quality and conversion performance are key drivers in addition to delivered feedstock cost. In this report, we estimate the benefits of R&D improvements to individual supply system unit operations, and present the status of feedstock logistics technology development for converting biomass into biofuels. These analyses are supported by experimental data where possible, and help to align the SOT relative to the cost goals defined in the Multi-Year Program Plan. The 2018 Herbaceous SOT aligned feedstock logistic design with current biorefinery’s design capacity utilized by biochemical conversion platform. Currently biochemical conversion platform utilizes a 725,000 dry ton/year biorefiney design for the techno economic analysis. Hence, feedstock delivered cost in the 2018 Herbaceous SOT is calculated based on biorefinery’s 725,000 dry ton design capacity instead of 800, 000 dry ton capacity utilized in the 2017 Herbaceous SOT. Biomass availabilities in this SOT were updated to year 2018 data from the 2016 Billion-Ton Report (BT16) (DOE 2016a), with the exception of switchgrass, for which the 2018 Herbaceous SOT utilized the 2019 switchgrass availability data from BT16. The BT16 report (DOE 2016a) does not project switchgrass availability in 2018; the soonest switchgrass is available in the BT16 report is 2019. Therefore, availability of switchgrass for this analysis was that projected for 2019. The 2018 Herbaceous SOT incorporates same technologies utilized in the 2017 Herbaceous SOT. However, a sensitivity analysis is performed to understand the impact of variation of process parameters on those technologies on feedstock logistic cost. New R&D data that shows the variations of process parameters affecting process performance is incorporated in the 2018 SOT to measure the variations in delivered feedstock cost. The 2018 Herbaceous SOT has also provided projected delivered feedstock of 2022 design case based on near term technical target under BETO funded R&D project. Finally, updated biorefinery size of 725,000 dry ton/year was incorporated within least-cost formulation model to select optimal siting and depot scales during optimization of the least cost blend. This modification to the optimization algorithm allows the trade-off between the cost of increased supply radius and the savings from selecting biomass from higher producing counties to be assessed. Such optimization has also showed the economic benefit of decentralized depots in comparison to centralized preprocessing co-located with the biorefinery by decoupling the biorefinery and feedstock locations. The 2018 Herbaceous SOT report documents the current modeled cost of a herbaceous feedstock supply system (from harvest to the pretreatment reactor throat, including grower payment) for hydrocarbon fuel production via biochemical conversion, based on equipment and processes now available or potentially available in the near term. The modeled cost also considers both the required quality and the availability of the biomass resources. The 2018 Herbaceous SOT predicts a modeled delivered feedstock cost of $83.67/dry ton (2016$); this is a $0.23/dry ton (2016$) decrease from the 2017 Herbaceous SOT. The modification of biorefinery’s designed capacity and increased projected biomass availability in the same supply shed contributed to this modeled cost reduction. The least-cost formulation model to optimally site and scale local distributed preprocessing depots also contributed to the cost reduction by considering county-level grower payment and distance from the biorefinery as variables in the optimization algorithm. Sensitivity analysis on various process parameters that affect delivered feedstock cost in the 2018 Herbaceous SOT shows that the delivered cost could varies from $80.45-$88.83/dry ton. The top factors that causes such variations are: effective baling rate, bale density, hammer mill throughput, interest rate and storage dry matter loss.

09 BIOMASS FUELS↗

Performance Comparison of Machine Learning Models for Ultrasonic Nondestructive Evaluation of Alkali-Silica Reaction in Concrete

Alkali-silica reaction (ASR) causes concrete degradation, leading to cracking, rebar corrosion, and reduced structural integrity, which raises safety concerns. Ultrasonic nondestructive evaluation (NDE) effectively assesses concrete properties and monitors ASR progression. However, its deployment and analysis require specialized expertise and subjective interpretation. As computational power increases, artificial intelligence (AI) and machine learning (ML) algorithms are increasingly being used to automate NDE data analysis across various industries for AI-assisted automation. Regulatory agencies are adapting to this technological shift, prompting a need to evaluate current ML technologies’ capabilities and limitations in assessing concrete material properties and damage. This report presents a comparative analysis of four ML regression models for predicting concrete material damage induced by ASR expansion using long-term ultrasonic data monitoring. The models investigated include linear regression (LR), support vector regression (SVR), shallow neural networks (NN), and deep neural networks (DNN). LR, SVR, and shallow NN models use features extracted from ultrasonic signals, whereas the DNN model processes time-domain ultrasonic signals and frequency spectra directly. The study systematically compared the models’ performance from various perspectives, including model input, prediction performance, and generalization ability. The findings indicate significant variability in model performance, with some ML algorithms achieving very high or very low prediction accuracy depending on the preprocessing and feature engineering (extraction and selection) applied. Key insights include the observation that shallow ML models (LR, SVR, and shallow NNs) require meticulous preprocessing and feature extraction to achieve high accuracy. In contrast, the DNN model, although it bypasses the need for feature engineering, necessitates extensive preprocessing to mitigate noise and computational demands. The SVR model emerged as the top performer among the shallow models, and the DNN model exhibited superior performance on specific datasets but struggled with generalization across specimens from different batches. Additionally, the SVR model is sensitive to temperature variations, whereas the DNN model is robust in this regard. Using recurrent neural networks is recommended for future ASR expansion prediction studies. Recurrent neural networks’ inherent ability to capture temporal dependencies and long-term patterns makes them well suited for analyzing sequential ultrasonic monitoring data. Overall, the results and conclusions of this study could provide insights into the capabilities and effectiveness of ML when applied to ultrasonic NDE data and help identify best practices for using ML for ultrasonic NDE of concrete material properties.

36 MATERIALS SCIENCE↗

Augmented Human Analysis (AHA)

Radio frequency (RF) signal monitoring generally emphasizes intentionally generated signals, such as WiFi, Bluetooth, or cellular transmissions. However, electronic devices also produce unintended radiated emissions (UREs), which could also be useful in RF spectrum analysis. In either case, deriving intelligence from RF signals is typically a human-intensive process requiring significant domain knowledge. In the Augmented Human Analysis (AHA) project, we investigate the utility of dimensionally aligned signal projection (DASP) and machine learning (ML) algorithms for accelerating RF analysis workflows. We find that while DASP algorithms can indeed highlight signal characteristics relevant for classification tasks, the choice of algorithmic hyperparameters greatly affects performance. To address this challenge, we evaluate the quality of DASP outputs using the silhouette score, which measures how well data points cluster; high silhouette scores indicate good clustering, and thus good hyperparameter values. This approach is critical for machine learning pipelines as the DASP parameters cannot be directly optimized during model training. By identifying good DASP parameters, and thus good DASP outputs, as a preprocessing step, we can decrease the amount of effort required for downstream ML model training. We demonstrate our workflow using a dataset of UREs from common household devices, showing that even without the aid of ML, proper selection of DASP parameters enables clustering by device type.

42 ENGINEERING↗

Jefferson Laboratory C100 Superconducting Radio-Frequency Cavity Fault Data, 2020

The dataset was created to train machine learning models for the task of identifying the (1) cavity and (2) fault type from C100-type cryomodules at the Thomas Jefferson National Accelerator Facility (Jefferson Lab), thereby replacing the time-consuming efforts of a subject matter expert. Superconducting radio-frequency (SRF) cavity trips represent a significant source of accelerator downtime. Real-time – rather than post-mortem – identification of the offending cavity and classification of the fault type would give control room operators valuable feedback for corrective action planning. The anticipated benefit is increased beam-on-target time for users and provides performance metrics that can be used to improve future cavity designs. A series of 17 RF signals are recorded for each of the 8 cavities in a C100 cryomodule every time a cavity trips. These time-series signals are written to file using a specially designed data acquisition system. The dataset represents fault events recorded during Continuous Electron Beam Accelerator Facility (CEBAF) beam operations between January 18, 2019 and March 9, 2020. The following filtering steps were applied to collected data; (1) only 4 of the 17 signals per cavity are retained (GMES, GASK, CRFP, DETA2) (2) only events with data from each of the eight cavities in the cryomodule are kept, (3) only events that were sampled at 5 kHz were kept, (4) events from cryomodule 0L04 were neglected, (5) events occurring between February 4, 3PM and February 5, 12PM were neglected. As a result of preprocessing, the dataset is comprised of 2,375 unique events. The full dataset is comprised of three files: features.csv, cavity_labels.csv, fault_labels.csv. Each instance in faults.csv includes a timestamp (“date_time”), a label for the cryomodule which experienced the trip (“zone_label”), and 192 features (“feature_1”, “feature_2”... “feature_192”). The features correspond to 6 autoregressive features for each of 4 signals per cavity for each of the 8 cavities (6 × 4 signals/cavity × 8 cavities/cryomodule = 192). To deal with the large variation of signal amplitudes, time-series standardization via the z-score (standard score) function was applied prior to computing the features. For each instance, there is an associated label for the (1) cavity which faulted first (cavity_labels.csv) and (2) the type of fault that caused the trip (fault_labels.csv). The cavity identification can take values of [0, 1, 2, 3, 4, 5, 6, 7, 8] and the fault type can take values of [‘Microphonics’, ‘Quench_100ms’, ‘Controls_Fault’, ‘E_Quench’, ‘Quench_3ms’, ‘Single_Cav_Turn_Off’ , ‘Heat_Riser_Choke’, ‘Multi_Cav_Turn_Off’].

43 PARTICLE ACCELERATORS↗

Applications of LIF to Document Natural Variability of Chlorophyll Content and Cu Uptake in Moss

Chlorophyll has long been used as a natural indicator of plant health and photosynthetic efficiency. Laser-induced fluorescence (LIF) is an emerging technique for understanding broad spectrum organic processes and has more recently been used to monitor chlorophyll response in plants. Previous work has focused on developing a LIF technique for imaging moss mats to identify metal contamination with the current focus shifting toward application to moss fronds and aiding sample collection for chemical analysis. Two laser systems (CoCoBi a Nd:YGa pulsed laser system and Chl-SL with two blue continuous semiconductor diodes) were used to collect images of moss fronds exposed to increasing levels of Cu (1, 10, and 100 nmol/cm 2 ) using a CMOS camera. The best methods for the preprocessing of images were conducted before the analysis of fluorescence signatures were compared to a control. The Chl-SL system performed better than the CoCoBi, with dynamic time warping (DTW) proving the most effective for image analysis. Manual thresholding to remove lower decimal code values improved the data distributions and proved whether using one or two fronds in an image was more advantageous. A higher DTW difference from the control correlated to lower chlorophyll a/b ratios and a higher metal content, indicating that LIF, with the aid of image processing, can be an effective technique for identifying Cu contamination shortly after an event.

59 BASIC BIOLOGICAL SCIENCES↗

Evaluation of native Earth system model output with ESMValTool v2.6.0

Earth system models (ESMs) are state-of-the-art climate models that allow numerical simulations of the past, present-day, and future climate. To extend our understanding of the Earth system and improve climate change projections, the complexity of ESMs heavily increased over the last decades. As a consequence, the amount and volume of data provided by ESMs has increased considerably. Innovative tools for a comprehensive model evaluation and analysis are required to assess the performance of these increasingly complex ESMs against observations or reanalyses. One of these tools is the Earth System Model Evaluation Tool (ESMValTool), a community diagnostic and performance metrics tool for the evaluation of ESMs. Input data for ESMValTool needs to be formatted according to the CMOR (Climate Model Output Rewriter) standard, a process that is usually referred to as “CMORization”. While this is a quasi-standard for large model intercomparison projects like the Coupled Model Intercomparison Project (CMIP), this complicates the application of ESMValTool to non-CMOR-compliant climate model output. In this paper, we describe an extension of ESMValTool introduced in v2.6.0 that allows seamless reading and processing of “native” climate model output, i.e., operational output produced by running the climate model through the standard workflow of the corresponding modeling institute. This is achieved by an extension of ESMValTool's preprocessing pipeline that performs a CMOR-like reformatting of the native model output during runtime. Thus, the rich collection of diagnostics provided by ESMValTool is now fully available for these models. For models that use unstructured grids, a further preprocessing step required to apply many common diagnostics is regridding to a regular latitude–longitude grid. Extensions to ESMValTool's regridding functions described here allow for more flexible interpolation schemes that can be used on unstructured grids. Currently, ESMValTool supports nearest-neighbor, bilinear, and first-order conservative regridding from unstructured grids to regular grids. Example applications of this new native model support are the evaluation of new model setups against predecessor versions, assessing of the performance of different simulations against observations, CMORization of native model data for contributions to model intercomparison projects, and monitoring of running climate model simulations. For the latter, new general-purpose diagnostics have been added to ESMValTool that are able to plot a wide range of variable types. Currently, five climate models are supported: CESM2 (experimental; at the moment, only surface variables are available), EC-Earth3, EMAC, ICON, and IPSL-CM6. As the framework for the CMOR-like reformatting of native model output described here is implemented in a general way, support for other climate models can be easily added.

58 GEOSCIENCES↗

Predicting oxidation damage of ultra high-temperature carbide ceramics in extreme environments using machine learning

Determining the oxidation resistance of UHTC carbides in extreme environments is challenging theoretically and experimentally due to the high dimensional complexity of influencing variables and intricate testing setups. Herein we demonstrate the use of machine learning (ML) models trained with experimental literature data to predict the oxide thickness of UHTC carbides exposed to air based on composition, mean grain size, relative densification, holding time, and temperature. A multi-dimensional database with 76 occurrences is created containing experimental results of Hf, Zr, and Ta carbides plus additives. In this study, the preprocessed database is then used to train ML models to predict their oxidation behavior. The trained model predicts the oxidation damage in the form of an average oxide thickness in UHTC carbides with a Mean Absolute Error (MAE) of ±65.45 μm for samples in the testing set that developed thicknesses up to 1000 μm. The model successfully predicted oxidation damage for a recession rate lower than 60 μm/min. It is noticed that the ensemble method MAE is increased to ±134.34 μm while forecasting the oxidation of samples with a recession rate higher than the threshold. The unprecedented approach is a novel way to predict the damage through the oxidation of carbide compounds before processing for a smarter design with room for improvement.

36 MATERIALS SCIENCE↗

Explainable tokamak-agnostic forecasting of fusion plasma instability via megahertz turbulent fluctuations

Scientific applications of artificial intelligence (AI) often remain limited by device-specific training and unexplained “black-box” approaches, creating fundamental barriers to cross-system generalization. This challenge is critical for nuclear fusion, where future reactors will have limited operational data for AI training. Here, we demonstrate that our neural network, trained solely on megahertz-scale turbulence measurements from one machine (DIII-D), forecasts Type-I edge localized mode (ELM) onsets in a different tokamak (KSTAR) through zero-shot weight transfer following physics-consistent preprocessing without device-specific retraining. Through an explainable AI framework combining gradient-weighted class activation mapping with physics validation, we reveal that our network can internalize physics relationships governing the ELM instabilities rather than memorizing device-specific patterns. The network perceives spatiotemporal features that correlate consistently with independently calculated instability growth rates, magnetohydrodynamic stability limits, and pedestal structure dynamics. Statistical analyses of dimensionally-reduced saliency features reveal the identical triangular features between the saliency representations, instability growth rates, and prediction probability across tokamaks, providing evidence that our forecasting system can show tokamak-agnostic generalization. This work contributes to a foundation for explainable scientific AI systems, where cross-system developments are essential for transcending traditional domain-specific constraints.

AI↗

Spatial and temporal characterization of municipal solid waste based on resource recovery pathways

This study presents a two-year, quarterly assessment of MSW across four source sectors (residential, schools, restaurants, and grocery stores) from sixteen sites across five U.S. states. MSW was manually sorted into 27 categories and aggregated into pathway fractions: high-moisture (HM) organics, low-moisture (LM) organics, recyclable (RC) materials, and residuals for disposal. Organics represented 89 % of the MSW stream. The largest fraction was HM organics consisting of food waste (31 %) and yard waste (3 %), with large coefficient of variations (CV), 79 and 278 %, respectively, reflecting high seasonal and site variability that varied significantly (p < 0.01) across sampling periods. The HM fraction showed properties favorable for anaerobic digestion, with moisture content ranging from 56 to 95 % and volatile solids ranges of 86-95 %. In contrast, the LM and RC fractions remained more stable (plastics CV = 41 %; paper CV = 53 %) with heating values up to 26.9 MJ/kg across sources, reflecting suitability for gasification. Microstructural analysis revealed less porosity in residential waste sampled at the landfill, which can influence preprocessing efficiency and microbial accessibility. Pathway informed allocations showed that 35 % of MSW is suitable for anaerobic digestion, 36 % for gasification, and 18 % for recycling, leaving 11 % requiring landfill disposal. These results provide quantitative evidence to determine feedstock allocation, waste-to-energy system design, and the development of data-driven sustainability and resource recovery strategies within a circular bioeconomy.

09 BIOMASS FUELS↗

The Preprocessing of Galaxies in the Early Stages of Cluster Formation in Abell 1882 at z = 0.139

A rare opportunity to distinguish between internal and environmental effects on galaxy evolution is afforded by "SuperGroups," systems that are rich and massive, but include several comparably rich substructures, surrounded by filaments. We present here a multiwavelength photometric and spectroscopic study of the galaxy population in the SuperGroup Abell 1882 (A1882) at z = 0.139, combining new data from the MMT and Hectospec with archival results from the Galaxy And Mass Assembly survey, the Sloan Digital Sky Survey, the Nasa/IPAC Extragalactic Database, the Gemini Multi-Object Spectrograph, and the Galaxy Evolution Explorer. These provide spectroscopic classifications for 526 member galaxies, across wide ranges of local density and velocity dispersion. We identify three prominent filaments along which galaxies seem to be entering the SuperGroup (mostly in E–W directions). A1882 has a well-populated red sequence, containing most galaxies with stellar mass >10 10.5 M Sun , and a pronounced color–density relation even within its substructures. Thus, galaxy evolution responds to the external environment as strongly in these unrelaxed systems as we find in rich and relaxed clusters. From these data, local density remains the primary factor, with a secondary role for distance from the inferred center of the entire structure's potential well. The effects on star formation, as traced by optical and near-UV colors, depend on galaxy mass. We see changes in lower-mass galaxies (M < 10 10.5 M Sun ) at four times the virial radius of major substructures, while the more massive near-UV Green Valley galaxies show low levels of star formation within two virial radii. The suppression of star formation ("quenching") occurs in the infall regions of these structures even before the galaxies enter the denser group environment.

79 ASTRONOMY AND ASTROPHYSICS↗

Spatially Accelerated Winding Numbers for Curved Geometry

The generalized winding number (GWN) is a scalar field that supports robust containment queries on curved geometry, including non-watertight, overlapping, and nested boundary representations. While queries can be easily parallelized over samples, direct evaluation on parametric curves and surfaces remains costly for large and complex models. Fast, state-of-the-art GWN approaches leverage a spatial index to approximate the GWN, typically coupled with a Taylor expansion which approximates the GWN contribution for far clusters of geometric primitives. However, such methods operate only on discrete inputs such as triangle meshes and point clouds, and would introduce containment errors near boundaries if applied to curved input. We extend support for fast GWN evaluation over arbitrary collections of NURBS curves in 2D and trimmed NURBS patches in 3D via a Bounding Volume Hierarchy that stores efficiently precomputed moment data in the hierarchy nodes. When querying the hierarchy, approximations for far clusters are used alongside direct evaluation for nearby NURBS primitives, achieving sub-linear complexity while preserving the geometric features in the vicinity of the query point. Central to our performance improvements is an adaptive subdivision strategy for NURBS primitives during a preprocessing phase, creating better spatial partitions while retaining the same accuracy for containment decisions as a direct evaluation. We demonstrate the performance and accuracy of our approach across a large collection of 2D and 3D datasets.

Computer science↗

Internal calibration of transient kinetic data via machine learning

The temporal analysis of products (TAP) reactor provides a vast amount of transient kinetic information that may be used to describe a variety of chemical features including residence time distributions, kinetic coefficients, number of active sites, reaction mechanism, etc. However, as with any measurement device, the TAP reactor signal is convoluted with noise and drift is common. In order to reduce the uncertainty of the kinetic measurement and any derived parameters or mechanisms, proper preprocessing must be performed prior to any advanced type of analysis. This preprocessing includes baseline correction, i.e., a shift in the voltage response, and calibration, i.e., a scaling of the flux response based on prior experiments. The traditional methodology of preprocessing requires significant user discretion and reliance on separate calibration experiments that may drift over time. Herein we use machine learning techniques combined with physical constraints to understand the noise and drift that is being generated within and between experiments for enhancement of the chemical kinetic signal. As such, the proposed methodology demonstrates clear benefits over the traditional preprocessing approach by eliminating the need for separate calibration experiments or heuristic input from the user.

36 MATERIALS SCIENCE↗

Predicting biomass comminution: Physical experiment, population balance model, and deep learning

An extended population balance model (PBM) and a deep learning-based enhanced deep neural operator (DNO+) model are introduced for predicting particle size distribution (PSD) of comminuted biomass through a large knife mill. Experimental tests using corn stalks with varied moisture contents, mill blade speeds, and discharge screen sizes are conducted to support model development. A novel mechanism in the extended PBM allows for including additional input parameters such as moisture content, which is not possible in the original PBM. The DNO+ model can include influencing factors of different data types such as moisture content and discharge screen size, which significantly extends the engineering applicability of the standard DNO model that only admits feed PSD and outcome PSD. Test results show that both models are remarkably accurate in the calibration or training parameter space and can be used as surrogate models to provide effective guidance for biomass preprocessing design.

09 BIOMASS FUELS↗

Comparing Sensor Fusion and Multimodal Chemometric Models for Monitoring U(VI) in Complex Environments Representative of Irradiated Nuclear Fuel

Optical sensors and chemometric models were leveraged for the quantification of uranium(VI) (0–100 μg mL –1 ), europium (0–150 μg mL –1 ), samarium (0–250 μg mL –1 ), praseodymium (0–350 μg mL –1 ), neodymium (0–1000 μg mL –1 ), and HNO 3 (2–4 M) with varying corrosion product (iron, nickel, and chromium) levels using laser fluorescence, Raman scattering, and ultraviolet–visible–near-infrared absorption spectra. In this paper, an efficient approach to developing and evaluating tens of thousands of partial least-squares regression (PLSR) models, built from fused optical spectra or multimodal acquisitions, is discussed. Each PLSR model was optimized with unique preprocessing combinations, and features were selected using genetic algorithm filters. The 7-factor D-optimal design training set contained just 55 samples to minimize the number of samples. The performance of PLSR models was evaluated by using an automated latent variable selection script. PLS1 regression models tailored to each species outperformed a global PLS2 model. PLS1 models built using fused spectra data and a multimodal (i.e., analyzed separately) approach yielded similar information, resulting in percent root-mean-square error of prediction values of 0.9–5.7% for the seven factors. Further, the optical techniques and data processing strategies established in this study allow for the direct analysis of numerous species without measuring luminescence lifetimes or relying on a standard addition approach, making it optimal for near-real-time, in situ measurements. Nuclear reactor modeling helped bound training set conditions and identified elemental ratios of lanthanide fission products to characterize the burnup of irradiated nuclear fuel. Leveraging fluorescence, spectrophotometry, experimental design, and chemometrics can enable the remote quantification and characterization of complex systems with numerous species, monitor system performance, help identify the source of materials, and enable rapid high-throughput experiments in a variety of industrial processes and fundamental studies.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Intelligent Sampling of Extreme-Scale Turbulence Datasets for Accurate and Efficient Spatiotemporal Model Training

With the end of Moore’s law and Dennard scaling, efficient training increasingly requires rethinking data volume. Can we train better models with significantly less data via intelligent subsampling? To explore this, we develop SICKLE, a sparse intelligent curation framework for efficient learning, featuring a novel maximum entropy (MaxEnt) sampling approach, scalable training, and energy benchmarking. We compare MaxEnt with random and phase-space sampling on large direct numerical simulation (DNS) datasets of turbulence. Evaluating SICKLE at scale on Frontier, we show that subsampling as a preprocessing step can, in many cases, improve model accuracy and substantially lower energy consumption, with observed reductions of up to 38×.

Brewer, Wes [ORNL] (ORCID:0000000236393956)↗

Herbaceous Feedstock 2021 State of Technology Report

The U.S. Department of Energy (DOE) promotes the production of advanced liquid transportation fuels from lignocellulosic biomass by funding fundamental and applied research that advances the State of Technology (SOT). As part of its involvement with this mission, Idaho National Laboratory (INL) completes an annual SOT report for biomass feedstock logistics. This report summarizes supply system impacts of Bioenergy Technologies Office (BETO)-funded research and development efforts at INL and INL collaboration with external partners (e.g. Forest Concepts, Purdue University ) that lead to improvements in feedstock supply systems. These include improvements to and observed performance of innovative harvest and collection methods, storage technologies, transportation and handling approaches, and advanced preprocessing technologies. Biomass quality and variability, and the interface between feedstock quality and conversion performance are key drivers in addition to delivered feedstock cost. In this report, we estimate the benefits of R&D technology improvements to individual supply system unit operations and present the status of feedstock logistics technology development for converting herbaceous biomass into biofuels. These analyses are supported by experimental data where possible and help to align the SOT relative to the cost goals defined in the Multi-Year Plan.

09 BIOMASS FUELS↗