Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “pre-processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Machine learning pipeline for denoising low signal-to-noise ratio and out-of-distribution transmission electron microscopy datasets

High-resolution transmission electron microscopy (HRTEM) is crucial for observing material’s structural and morphological evolution at Angstrom scales, but the electron beam can alter these processes. Devices such as CMOS-based direct-electron detectors operating in electron-counting mode can be utilized to substantially reduce the electron dosage. However, the resulting images often lead to a low signal-to-noise ratio, which requires frame integration that sacrifices temporal resolution. Several machine learning (ML) models have been recently developed to successfully denoise HRTEM images. Yet, these models are often computationally expensive, and their inference speeds on GPUs are outpaced by the imaging speed of advanced detectors, precluding in situ analysis. Furthermore, the performance of these denoising models on datasets with imaging conditions that deviate from the training datasets has not been evaluated. To mitigate these gaps, we propose a new self-supervised ML denoising pipeline specifically designed for time-series HRTEM images. This pipeline integrates a blind-spot convolution neural network with pre-processing and post-processing steps, including drift correction and low-pass filtering. Results demonstrate that our model outperforms various other ML and non-ML denoising methods in noise reduction and contrast enhancement, leading to improved visual clarity of atomic features. Additionally, the model is drastically faster than U-Net-based ML models and demonstrates excellent out-of-distribution generalization. The model’s computational inference speed is in the order of milliseconds per image, rendering it suitable for application in in-situ HRTEM experiments.

36 MATERIALS SCIENCE↗

The gas-phase mass–metallicity relation of dwarf galaxies across large-scale environments using the CAVITY parent sample

Context. The gas-phase mass–metallicity relation (MZR) of galaxies shows a noticeable break in slope and an increased scatter at low stellar masses, suggesting that the physical processes governing chemical enrichment differ between dwarf and high-mass systems. Dwarf galaxies, in particular, are highly susceptible to both internal and environmental mechanisms due to their shallow potential wells. Aims. The primary aim of this work is to assess whether a single, universal MZR can describe dwarf galaxies across diverse large-scale environments, or whether systematic environmental variations emerge. To probe these, we examine the MZR and star formation rate (SFR) of dwarf galaxies with stellar masses in the range of 8.9 < log(M ★ /M ⊙ ) < 9.5. Methods. Using optical spectra from the Sloan Digital Sky Survey, we measured the fluxes of key emission lines via the pyPipe3D full spectral fitting pipeline. Aperture-corrected fluxes, along with multiple metallicity indicators and calibrations, were used to derive the MZR and the SFR for 353, 311, and 22 dwarf galaxies located in voids, filaments, and clusters, respectively. Results. We find a systematic variation in the MZR slope, steeper in voids (0.28 ± 0.03) and progressively flatter in clusters (0.17 ± 0.08), indicating a dependence of the MZR on the large-scale environment in this mass regime. When galaxies are separated by local density, no significant differences are observed between isolated and non-isolated dwarfs in voids. Isolated dwarf galaxies in filaments also exhibit properties similar to those of their counterparts in voids. However, non-isolated filament galaxies exhibit similar MZR slopes comparable to those of cluster dwarfs and flatter slopes than their counterparts in voids. Conclusions. We report both large- and local-scale environmental dependencies in the gas-phase metallicity and in the slope of the MZR for dwarf galaxies. Consistent with the general consensus on the pre-processing of galaxies in filaments, our results indicate that the influence of the local environment becomes increasingly significant within the filamentary regions of the cosmic web, affecting the chemical enrichment and star formation activity of low-mass systems. These findings further suggest that a portion of the scatter commonly observed in the MZR of dwarf galaxies arises from environmental effects.

Bidaran, Bahar [Dpto. de Física Teórica y del Cosm↗

Decoding diffraction and spectroscopy data with machine learning: A tutorial

This Tutorial provides a step-by-step guide on how to apply supervised machine-learning techniques to analyze diffraction and spectroscopy data. This Tutorial details four models—a reconstruction-focused model, a regression-focused model, a hybrid reconstruction/regression model, and a multimodal model—that use x-ray diffraction profiles and vibrational density of states spectra to predict various microstructural descriptors. In this Tutorial, we cover data pre-processing steps, constructions of the models via dimensionality reduction and regression, training, and analysis of these models. Comparisons of the model’s performance are provided, highlighting the strength and weakness of the various approaches utilized.

36 MATERIALS SCIENCE↗

Sharp detection of low-dimensional structure in probability measures via dimensional logarithmic Sobolev inequalities

Identifying low-dimensional structure in high-dimensional probability measures is an essential pre-processing step for efficient sampling. To identify this structure, we approximate the target measure as a perturbation of an arbitrary reference measure along a few directions in $\mathbb{R}^{d}$. These directions are determined by minimizing an upper bound on the Kullback–Leibler (KL) divergence between the target and its approximation. Our contribution improves upon previous works by leveraging dimensional logarithmic Sobolev inequalities to refine the bound on the KL divergence. These inequalities lead to a uniformly tighter bound on the KL divergence, thereby enhancing the identification of the most significant perturbation directions. In particular, when the target and reference are both Gaussian, minimizing the resulting bound is equivalent to minimizing the KL divergence. We further demonstrate the applicability of this analysis to the squared Hellinger distance, where analogous reasoning shows that the dimensional Poincaré inequality offers improved bounds.

Bayesian inference↗

Auriga Streams – I: disrupting satellites surrounding Milky Way-mass haloes at multiple resolutions

In a hierarchically formed Universe, galaxies accrete smaller systems that tidally disrupt as they evolve in the host’s potential. We present a complete catalogue of disrupting galaxies accreted onto Milky Way-mass haloes from the Auriga suite of cosmological magnetohydrodynamic zoom-in simulations. We classify accretion events as intact satellites, stellar streams, or phase-mixed systems based on automated criteria calibrated to a visually classified sample, and match accretions to their counterparts in haloes re-simulated at higher resolution. Most satellites at the present day have lost substantial amounts of stellar mass – 67 per cent have $f_\text{bound} < 0.97$ (our threshold of lost stellar mass to no longer be considered intact), while 53 per cent satisfy a more stringent $f_\text{bound} < 0.8$. Streams typically outnumber intact systems, contribute a smaller fraction of overall accreted stars, and are substantial contributors at intermediate distances from the host centre ($\sim$0.1 to $\sim 0.7R_\text{200m}$, or $\sim$35 to $\sim$250 kpc for the Milky Way). We also identify accretion events that disrupt to form streams around massive intact satellites instead of the main host. Streams are more likely than intact or phase-mixed systems to have experienced pre-processing, suggesting this mechanism is important for setting disruption rates around Milky Way-mass haloes. All of these results are preserved across different simulation resolutions, though we do find some hints that satellites disrupt more readily at lower resolution. The Auriga haloes suggest that disrupting satellites surrounding Milky Way-mass galaxies are the norm and that a wealth of tidal features waits to be uncovered in upcoming surveys.

galaxies: haloes↗

LaueMatching: an approach for rapid and robust indexing of Laue diffraction patterns

Traditional Laue diffraction pattern indexing often struggles with noisy data, weak signals, peak overlap and missing reflections, particularly from complex or deformed microstructures. Here, we introduce LaueMatching, a high-throughput indexing algorithm designed to overcome these limitations. LaueMatching utilizes a fundamentally different approach based on direct pattern correlation: experimentally pre-processed images are compared against a comprehensive pre-computed library of simulated diffraction patterns corresponding to a dense grid of possible orientations. This approach bypasses the need for explicit peak identification and fitting, steps that are often a failure point for traditional methods. The algorithm rapidly and robustly indexes multiple crystallographic orientations and crystal systems simultaneously, even from challenging patterns. LaueMatching's effectiveness and accuracy have been rigorously tested and validated on diverse experimental (Ni, Al, EuAl 2 O 4 ) and simulated diffraction patterns, demonstrating high-fidelity orientation refinement. Code to implement this approach on both CPU and GPU resources can be downloaded from https://github.com/AdvancedPhotonSource/LaueMatching.

36 MATERIALS SCIENCE↗

Descriptor: High Temporal Resolution Meteorological Data at Oak Ridge Reservation (ORR-HiResMet)

Access to continuous, quality assessed meteorological data is critical for understanding the climatology and atmospheric dynamics of a region. Research facilities like Oak Ridge National Laboratory (ORNL) rely on such data to assess site-specific climatology, model potential emissions, establish safety baselines, and prepare for emergency scenarios. To meet these needs, on-site towers at ORNL collect meteorological data at 15-minute and hourly intervals. However, data measurements from meteorological towers are affected by sensor sensitivity, degradation, lightning strikes, power fluctuations, glitching, and sensor failures, all of which can affect data quality. To address these challenges, we conducted a comprehensive quality assessment and processing of five years of meteorological data collected from ORNL at 15-minute intervals, including measurements of temperature, pressure, humidity, wind, and solar radiation. The time series of each variable was pre-processed and gap-filled using established meteorological data collection and cleaning techniques, i.e., the time series were subjected to structural standardization, data integrity testing, automated and manual outlier detection, and gap-filling. The data product and highly generalizable processing workflow developed in Python Jupyter notebooks are publicly accessible online. As a key contribution of this study, the evaluated 5-year data will be used to train atmospheric dispersion models that simulate dispersion dynamics across the complex ridge-and-valley topography of the Oak Ridge Reservation in East Tennessee.

Steckler, Morgan R. [Oak Ridge National Laboratory↗

FiberFlex: Real-time FPGA-based Intelligent and Distributed Fiber Sensor System for Pedestrian Recognition

In recent years, security monitoring of public places and critical infrastructure has heavily relied on the widespread use of cameras, raising concerns about personal privacy violations. To balance the need for effective security monitoring with the protection of personal privacy, we explore the potential of optical fiber sensors for this application. This article proposes FiberFlex, an intelligent and distributed fiber sensor system. Ultizing Field Programmable Gate Arrays (FPGA) high-level synthesis (HLS) acceleration, FiberFlex offers real-time pedestrian detection by co-designing the entire pipeline of optical signal acquisition, processing, and recognition networks based on the principles of optical fiber sensing. As a promising alternative to traditional camera-based monitoring systems, FiberFlex achieves pedestrian detection by analyzing the vibration patterns caused by pedestrian footsteps, enabling security monitoring while preserving individual privacy. FiberFlex comprises three modules: First , fiber-optic sensing system: A fiber-optic distributed acoustic sensing (DAS) system is built and used to measure the ground vibration waves generated by people walking. Second , algorithms: We first collect the training data by measuring the ground vibration waves, label the data, and use the data to train the neural network models to perform pedestrian recognition. Third , hardware accelerators: We use HLS tools to design hardware modules on FPGA for data collection and pre-processing and integrate them with the downstream neural network accelerators to perform in-line real-time pedestrian detection. The final detection results are sent back from FPGA to the host CPU. We implement our system FiberFlex with the in-house built DAS system and AMD/Xilinx Kintex7 FPGA KC705 board and verify the whole system using the real-world collected data. We conduct recognition tests on five test subjects of varying ages, heights, and weights in a fixed sensing area. Each subject experienced 20 real-time recognition tests using their daily walking habits, and the subjects were given adequate rest between tests. After 100 tests on five test subjects, the overall real-time recognition accuracy exceeded \(88.0\%\) . The whole system uses 55 W of power, 33 W in the optical DAS system and 22 W in the FPGA. Relying on its end-to-end interdisciplinary design, FiberFlex seamlessly combines fiber-optic sensors with FPGA accelerators to enable low-power real-time security monitoring without compromising privacy, making it a valuable addition to the existing security monitoring network. According to FiberFlex, more valuable research can be conducted in the future, such as fall monitoring for the elderly, migration of identification networks between different application scenarios, and improvement of anti-interference performance in more complex environments. In future perception networks, where the “eyes” are not feasible, let’s use fiber optic touch instead.

Distributed↗

MAPLE v.1.0

SAND2025-00659O MAPLE is a software tool that uses epigenomic data to predict gene expression. MAPLE uses a set of epigenomic modifications to determine the effect on gene expression in a specific subset of species. The algorithm can be trained on additional species and epigenomic modifications, enhancing its predictive capabilities. EAGLE employs a hybrid neural network architecture, featuring a convolutional front-end and a multi-head attention layer, to process pre-processed signal data as input. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy's National Nuclear Security Administration under contract DE-NA0003525.

Davis IV, Warren↗

Seismic H2: Version 1.0

Seismic-H2 is an integrated software package for geological hydrogen reservoir simulation, optimization, and leakage monitoring. The package includes multiple components: (1) code used for modeling seismic wave propagation in 3D heterogeneous elastic media based on finite-difference method to support detection of geological hydrogen storage reservoir leakage; (2) 3D reservoir simulations of leaks from an underground reservoir and 3D simulations of saline aquifers and depleted gas reservoirs; (3) seismic monitoring costs of passive and active seismic monitoring required for UHS; (4) rock physics calculations and interpolations for converting the reservoir simulations from part (2) into the elastic media models in part (1); (5) pre-processing seismic data; and lastly (6), a GUI interface that combines these different components.

Creasy, Neala↗

HydraGNN_Predictive_GFM_2026 - Ensemble of predictive graph foundation models for atomistic materials modeling

This release contains data and parameters of HydraGNN-based graph foundation models trained as a result of the work published in the pre-print "Exascale Multi-Task Graph Foundation Models for Imbalanced, Multi-Fidelity Atomistic Data" by M. Lupo Pasini et al. (https://arxiv.org/abs/2604.15380). We jointly train on 16 open first-principles datasets (544+ million structures covering 85+ elements) using a multi-task architecture with per-dataset heads and a scalable ADIOS2/DDStore data pipeline. On Frontier, we execute six large-scale DeepHyper hyperparameter optimization campaigns in FP64 and promote the top-performing message-passing models to sustained 2,048-node training, yielding a PaiNN-based lead model. The version of HydraGNN used to generate the outputs provided in this release is HydraGNN v5.0 (https://github.com/ORNL/HydraGNN/releases/tag/v5.0) The list of datasets used for the training of the graph foundation model is the following: 1) Alexandria [1] 2) ANI1x [2] 3) MPTrj [3] 4) Open Catalyst 2020 (OC20) [4] 5) Open Catalyst 2022 (OC22) [5] 6) Open Catalyst 2025 (OC25) [6] 7) Open Direct ir Capture 2023 (ODAC23) [7] 8) Open Materials 2024 (OMat24) [8] 9) Open Molecules 2025 (OMol25) [9] 10) OMol25-neutral (subset of OMol25 that contains only molecules with zero total charge) 11) OMol25-non-neutral (subset of OMol25 that contains only molecules with non-zero total charge) 12) Open Polymers 2026 (OPoly2026) [10] 13) Nabla2DFT [11] 14) QCML [12] 15) QM7X [reference 13] 16) transition1x [14] Dataset references: [1] J. Schmidt et al., “A dataset of 175k stable and metastable materials calculated with the PBEsol and SCAN functionals,” Scientific Data, vol. 9, p. 64, 2022. [2] J. S. Smith et al., “The ANI-1ccx and ANI-1x data sets, coupled-cluster and density functional theory properties for molecules,” Scientific Data, vol. 7, p. 134, 2020. [Online]. Available: https: //www.nature.com/articles/s41597-020-0473-z [3] A. Jain et al., “Commentary: The Materials Project: A materials genome approach to accelerating materials innovation,” APL Materials, vol. 1, no. 1, p. 011002, 07 2013. [Online]. Available: https://doi.org/10.1063/1.4812323 [4] L. Chanussot et al., “Open catalyst 2020 (oc20) dataset and community challenges,” ACS Catalysis, vol. 11, no. 10, pp. 6059–6072, 2021. [Online]. Available: https://doi.org/10.1021/acscatal.0c04525 [5] K. Tran et al., “Open catalyst 2022 (oc22) dataset and challenges for oxidation electrocatalysts,” ACS Catalysis, vol. 13, no. 5, pp. 3066–3084, 2023. [Online]. Available: https://doi.org/10.1021/acscatal.2c05426 [6] S. J. Sahoo et al., “The open catalyst 2025 (oc25) dataset and models for solid-liquid interfaces,” arXiv preprint arXiv:2509.17862, 2025. [Online]. Available: https://arxiv.org/abs/2509.17862 [7] A. Sriram et al., “The open DAC 2023 dataset and challenges for sorbent discovery in direct air capture,” ACS Central Science, vol. 10, no. 5, pp. 923–941, 2024. [8] L. Barroso-Luque et al., “Open materials 2024 (omat24) inorganic materials dataset and models,” 2024. [Online]. Available: https://arxiv.org/abs/2410.12771 [9] D. S. Levine et al., “The open molecules 2025 (OMol25) dataset, evaluations, and models,” 2025. [Online]. Available: https://arxiv.org/abs/2505.08762 [10] D. S. Levine et al., The open polymers 2026 (OPoly26) dataset and evaluations,” arXiv preprint arXiv:2512.23117, 2025. [Online]. Available: https://arxiv.org/abs/2512.23117 [11] K. Khrabrov et al., “Nabla2dft: A universal quantum chemistry dataset of drug-like molecules and a benchmark for neural network potentials,” in NeurIPS 2024 Datasets and Benchmarks Track, 2024. [Online]. Available: https://openreview.net/forum?id=ElUrNM9U8c [12] S. Ganscha et al., “The QCML dataset, quantum chemistry reference data from 33.5M DFT and 14.7B semi-empirical calculations,” Scientific Data, vol. 12, p. 406, 2025. [13] J. Hoja et al., “QM7-X, a comprehensive dataset of quantum-mechanical properties spanning the chemical space of small organic molecules,” Scientific Data, vol. 8, p. 43, 2021. [Online]. Available: https://www.nature.com/articles/s41597-021-00812-2 [14] M. Schreiner et al., “Transition1x - a dataset for building generalizable reactive machine learning potentials,” Scientific Data, vol. 9, p. 779, 2022. The folder "datasets_ADIOS2_format" contains the set of pre-processed datasets in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used for the development and training of GFMs in this work. The "datasets_ADIOS2_format" directory contains 2 sub-directories, one for the version "v1" of the datasets and one for the version "v2" of the datasets. The version "v1" of the datasets provides values of the total energy as they are extracted from the original data as it was released by the respective institutions. The version "v2" of the datasets provides values of the energy that have been realigned. The realignment was performed by training a linear regression model that predicts the total energy as a function of the chemical composition of the atomistic structure, and then subtract such prediction from the original value of the total energy. Both folders "v1" and "v2" contain 16 sub-directories, each corresponding to an ADIOS2-formatted dataset The folder "DeepHyper-results" contains the configurational files and model's parameters for all the 186 HPO trials that were successfully completed by the scalable hyperparameter optimization (HPO) runs on Frontier. The content of the folder "DeepHyper-results" I structured as follows: 1) task-list.txt: list of mpnn name, jobid, and deephyper task id 2) gfm_${MPNN}_${JOBID}_0.${TASKID}: run directory with checkpoint files 3) gfm_${MPNN}: deephyper summary directory (*.csv) for each specific MPNN type 4) deephyper-experiment-${JOBID}: output and error logs for each job The file "deephyper-sorted.csv" contains the details of each HydraGNN model built and tested by HPO, obtained by merging the (*.csv) filed from each HPO run executed. Out of all the HPO trials, we selected 10 to continue the training of the respective HydraGNN models. Due to limited computational budget available in the LRN070 allocation we could not complete the training till convergence for all these 10 selected models. The folder "models" contains multiple sub-folders, one per each HydraGNN model trained. Each model sub-folder contains the parameters of each HydraGNN model, with multiple checkpoint-restarts. The list of sub-folders are as follows: 1) multidataset_hpo-BEST1-fp64 2) multidataset_hpo-BEST2-fp64 3) multidataset_hpo-BEST3-fp64 4) multidataset_hpo-BEST4-fp64 5) multidataset_hpo-BEST5-fp64 6) multidataset_hpo-BEST6-fp64 7) multidataset_hpo-BEST7-fp64 8) multidataset_hpo-BEST8-fp64 9) multidataset_hpo-BEST9-fp64 10) multidataset_hpo-BEST10-fp64 Within each one of these folders, additional auxiliary log files are provided with descriptions about how the training proceeded. The lead PaiNN-model is contained inside "multidataset_hpo-BEST6-fp64". The file "mlp_branch_weights" contains the parameters of the multi-layer perceptron (MLP) used to reconcile the predictions of the 16 output decoding heads of the HydragNN architectures. The MLP takes in input the chemical composition of the atomistic structure and predicts averaging weights to linearly mix the predictions of each output decoding head toward consolidating them into a single one. The folder "1.1billion-structure-inference" contains 1.1 billion atomistic structures randomly generated. Each structures is associated with energy and forces predicted with the lead-PaiNN model combined with the MLP model for reconciliation of the multi-branch predictions generated by the 16 output decoding heads. The folder "1.1billion-structure-inference" contains 9,300 (*.tar.gz) subdirectories, one per Frontier compute node used to execute the inference at exascale. Once uncompressed, each (*.tar.gz) subdirectory contains an ADIOS2 (*.bp) file container, where each atomistic structure is stored as a PyTorch-Geometric Data object. The file "export_dataset_environment_variables.sh" contains the environment variables that need to be set before running the HydraGNN code to reproduce the results provided in this dataset release. The code that can be used to load the ADIOS2 files, load HydraGNN models, and run inference is available at: https://github.com/ORNL/HydraGNN/releases/tag/v5.0

36 MATERIALS SCIENCE↗

X-ray Computed Tomography Data of Dense Metallic Components

The data shared in here are X-ray computed tomography (XCT) scans of a hexagonal fuel nozzle in 3 sections with the Metrotom 800 system at the Manufacturing Demonstration Facility (MDF) at Oak Ridge National Laboratory. The data are used in the paper "Tomographic Sparse View Selection using the View Covariance Loss, by Lin et al. (doi:10.1109/TPAMI.2025.36000720), accepted to the international conference on computational imaging (ICCP 2025). Figures 4-7 in the paper describe the part/XCT scan. File name Descriptions: Bottom section: TCR- Single Channeled SRC L 2019-3-18 12-26-41.hdf5 Medium section: TCR- Single Channeled SRC M 2019-3-18 13-8-9.hdf5 Top section: TCR- Single Channeled SRC T 2019-3-18 13-45-39.hdf5 Each hdf5 file contains projection data, and all the relevant X-ray CT scan setting. The full list of included attributes: distance_unit: Units of all distances specified angle_unit : Units of the angles angles: Array of all angles used voxel_size_xy: Baseline recon (if any) has this voxel size in the in-plane direction voxel_size_z: Baseline recon (if any) has this voxel size in the cross-plane direction det_pixel_size_col: Size of the detector pixels in the column dimension det_pixel_size_row: Size of the detector pixels in the row dimension src_iso_dist: Source to iso-center distance iso_det_dist: Iso-center to detector distance det_angle: If the detector is rotated/tilted, this angle corresponds to that value det_row_offset: Center of rotation offset in the vertical direction det_col_offset: Center of rotation offset in the horizontal direction reconstruction: A baseline reconstruction stored as 3D array BHC params: Beam-hardening parameters - Van De Casteel Model - if it has been used to pre-process the projections We also provided a python script (hdf_io.py) that allows the user to read the relevant data from each hdf5 file.

Ziabari, Amir [Oak Ridge National Laboratory]↗

Site and endmember spectra of terrestrial vegetation and soils for the Colorado Headwaters Ecological Spectroscopy Study, June-July 2025

This dataset provides site and endmember spectra collected during the 2025 Colorado Headwaters Ecological Spectroscopy Study (CHESS) campaign. The site spectra were collected to help validate airborne hyperspectral data acquired by the National Ecological Observatory Network's aerial observation platform (NEON AOP). Endmember spectra were collected to augment existing spectral libraries with additional samples of bare surfaces and non-photosynthetic vegetation. All measurements were acquired with an Analytical Spectral Devices (ASD) FieldSpec4 Hi-Res NG (Next Generation) spectroradiometer, which records radiance at 1nm (nanometer) intervals from the ultraviolet to the short-wave infrared (350-2500 nm). The dataset includes spectra measured at meadow sites where the CHESS team also collected vegetation samples for trait analyses. The site spectra were collected with the ASD FieldSpec4 palm grip attachment using an 8° field-of-view foreoptic. Site spectra are integrated measurements of the entire surface within the foreoptic’s field of view. For site-level spectra, the sun is the illumination source. A Spectralon panel mounted on a tripod was used for instrument optimization and white reference measurements for all site spectra. Site spectra were acquired within two hours of solar noon and within 48 hours of a NEON AOP overflight. Site spectra are labeled by date, sampling area, and site number according to the naming conventions of the CHESS campaign’s data management plan. The dataset also contains endmember spectra in the following categories: photosynthetic vegetation (PV), non-photosynthetic vegetation (NPV), bare (soil/rock), and flowers. Endmember measurements were acquired using either the contact probe or the leaf clip attachments of the ASD FieldSpec4. In these configurations, the bulb inside the spectrometer provides the light source for the measurements. The spectrometer was optimized and white reference measurements were recorded using the circular white pucks attached to the contact probe and leaf clip. Because they do not rely on solar illumination, contact probe and leaf clip measurements were collected during a broader time frame than the palm grip site spectra. Some endmembers were measured at CHESS meadow sites, while others were collected within the larger sampling area or in nearby locations (e.g. Gothic Townsite) with similar characteristics. Radiance, reflectance, and metadata files are split into three subfolders according to measurement type: proximal/palm grip (prx), contact probe (cp), and leaf clip (lc). Radiance spectra are provided in ASD file format (.asd file extension). All ASD files can be opened using the provided scripts. Metadata is provided in two formats: CSV file format (no geolocation) and GEOJSON file format (includes geolocation for each spectra). The dataset includes a set of pre-processed reflectance spectra as CSV files (yyyymmdd_rfl.csv). The python scripts and jupyter notebook used to calculate reflectance spectra from the ASD radiance data is included here and was previously published at: https://doi.org/10.3334/ORNLDAAC/2446. There is also a folder of JPEG photographs corresponding to selected spectra. We include a protocol document with detailed steps for ASD FieldSpec4 assembly and operations. This data additionally contains a file level metadata (flmd.csv) and data dictionary (dd.csv) file. Geospatial information: Geospatial data for mapping measurement site locations are in the files CHESS_polygons_lai_UTM.geojson, CHESS_polygons_shrub_UTM.geojson, and CHESS_polygons_meadow_UTM.geojson in the companion geospatial package for the 2025 CHESS campaign, ‘CHESS 2025: Location data for field observations and sampling’ (Henderson et al., 2026). CHESS Project Description: The Colorado Headwaters Ecological Spectroscopy Study (CHESS) comprised a multi-week airborne remote sensing and field observation campaign in the Upper Gunnison Basin, Colorado, conducted in June and July of 2025. Airborne remote sensing was conducted by the National Ecological Observatory Network Airborne Observation Platform (NEON AOP), concurrent with a field campaign run by the Rocky Mountain Biological Laboratory (RMBL), the Lawrence Berkeley National Laboratory (LBNL) and SLAC National Accelerator Laboratory Watershed Function Science Focus Area (SFA), and NASA-JPL (Jet Propulsion Laboratory) Earth Surface Mineral Dust Source Investigation (EMIT) program. Between June 10 and July 18, 2025, the NEON AOP flight team collected high-resolution aerial imaging spectroscopy and Light Detection and Ranging (LiDAR) data over three domains: the Upper East River (CRBU), Almont Triangle (ALMO), and the Upper Taylor Basin (UPTA). In coordination with the flights, a field campaign acquired ground-truth observations, including observations of vegetation composition, foliar traits, forest demography, and subsurface properties in 18 core sampling areas within the domains. Additional surface water observations were taken at over 380 point locations. All CHESS campaign datasets can be found within the CHESS ESS-DIVE data portal: https://data.ess-dive.lbl.gov/portals/chess. Funding Acknowledgment: This research was carried out at the Jet Propulsion Laboratory, California Institute of Technology, under a contract with the National Aeronautics and Space Administration (80NM0018D0004) and was funded by EMIT Extended Mission Phase E Science.

2018 NEON and 2025 CHESS Campaigns↗

Coal-Waste-Enhanced Filaments for Additive Manufacturing of High-Temperature Plastics and Ceramic Composites

In the United States, coal waste from over a century of mining and burning coal for heat and electricity has accumulated as mountains of coal fly ash and bottom ash and acre-size ponds, coal fines and gob. These materials can be a problem for local communities and water systems. A cost-effective process to utilize high volumes of these coal wastes in a high-value product would be beneficial to those communities by reducing the amount of waste and providing jobs, manufacturing components, and materials from the waste. Many coal-to-products technologies (e.g., carbon fibers, graphene, carbon foam) rely on carefully choosing the starting material and then altering it chemically or thermally to make the products work. Due to the wide variability of composition and coal content in typical coal waste streams, many high-volume coal waste streams are likely to be unsuitable for use in those technologies. Semplastics’ technology has been shown to utilize most types of coal waste successfully without any pre-selection or pre-processing requirements other than a nominal particle-size reduction for wastes like bottom ash. This characteristic of Semplastics’ solution may enable the use of much larger volumes of a wider range of coal wastes than other coal-to-products technologies. In this project, Semplastics leveraged its unique experience with both coal waste (fly ash or coal combustion residuals), resin materials, and 3D printing to develop 3D printer filaments using common coal wastes – bituminous coal fines and fly ash – and researched the feasibility of using other forms of coal waste as fillers. Simple 3D-printed parts were successfully produced from the coal waste enhanced filaments, which were found to have improved strength and stiffness.

01 COAL, LIGNITE, AND PEAT↗

Ring Pull Strain Analysis Version 1.1

This report details an analysis package, Ring Pull Strain Analysis (RPSA), that can be used to present and quantify digital image correlation (DIC) data as it relates to a gaugeless ring pull test. Gaugeless ring pull is a testing technique for mechanical testing of small annular samples, usually cut from a thin-walled tube. DIC data is often necessary for this kind of test because bending moments present on the ring cause a non-uniform strain distribution and localized measurements are necessary. In addition, the annular geometry of a ring lends itself to a polar representation, which is not present with typical DIC analysis methods. RPSA was made to calculate and plot the polar representation of strain from standard pre-processed DIC data of a gaugeless ring pull test. Further analysis can be done on ring pull including a quasi-uniaxial tensile analysis and coating analysis, which are also performed by RPSA. In addition, due to the universality of DIC plotting and ring pull test analysis, RPSA can accommodate a wide variety of tests, though it is tailored for ring pull testing. This report details how RPSA works, including the theory, assumptions, and logic behind the calculations and the structure of the program.

36 MATERIALS SCIENCE↗

Optimizing enzymes for plastic upcycling using machine learning design and high throughput experiments

Plastic use is ubiquitous in the modern world, and polyethylene terephthalate (PET) is one of the most abundantly produced plastics (and the most highly produced polyester), with ~65 million metric tons manufactured annually. To the consumer, PET is likely most recognizable as the plastic used to make beverage bottles. Like many plastics, traditional mechanical or chemical means of PET deconstruction and upcycling are costly and inefficient. Because of these challenges, recycled plastic is generally of lower quality and is more expensive to produce than virgin plastic derived from petroleum. Ultimately, this results in most plastic ending up as waste. We view plastic waste as an underutilized resource which, with the development of more efficient and high-quality recycling processes, could (1) generate significant economic value while (2) decreasing petroleum usage and greenhouse gas emissions, as well as (3) minimizing its negative environmental and health impacts. Biocatalytic recycling, or biomanufacturing the basic building blocks of new plastic from plastic waste, is a promising approach to plastic reuse that complements existing recycling technologies. Recently, biological enzymes capable of breaking down PET have garnered significant attention as an attractive means of dealing with the plastic problem. These enzymes are currently undergoing pilot studies for implementation in industrial-scale enzyme-based recycling. However, there are significant limitations to current enzymes, including the need to perform costly pre-processing of the plastic waste before the enzymes are able to work. Further optimization of these enzymes is necessary to make these technologies competitive, and ultimately incentivise industry-wide adoption of this biology-based green recycling technology. n this work we demonstrate a means to design and generate performant biological enzymes, capable of efficiently deconstructing plastic waste. Specifically, we applied recent advances in artificial intelligence, machine learning, and statistical analysis to design new versions and discover natural enzymes capable of breaking down PET. We focused on optimizing key properties that are important for industrial-scale enzymatic recycling such as pH and thermotolerance. Normal testing of enzymatic plastic-deconstruction is extremely labor intensive and so through this work we also developed a robotic-assisted experimental pipeline capable of characterizing thousands of candidate enzymes. The results of this iterative, AI-guided, multi-discipline approach have led to increases in enzymatic breakdown of over 150X over starting enzymes. This work supports the rapidly developing and transformative field of biocatalytic solutions to environmental problems beyond the discovery and predictive understanding of enzymes for polymer recycling, and has wide implications for tackling numerous energy problems such as carbon capture and fixation (e.g., engineering carbon monoxide dehydrogenase and the rubisco-pathway), biomining (e.g., design of lanthanide-binding proteins) and biomanufacturing (e.g., lignin-deconstruction enzymes).

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Resonance Self-Shielding: Why it is so Important

This paper is one of a series that I am writing to document my 58 years of experience with ENDF and Neutron Transport calculations, beginning when I worked at the National Nuclear Data Center (NNDC), Brookhaven National Laboratory (BNL), from 1967 to 1972. During those years I was the head of the computer unit of NNDC, assigned to develop computer codes to pre-process, view and test ENDF/B data. Since then, I have continued to support the ENDF effort without any official position or monetary compensation, because I realized how important accurate nuclear data is for use in use in our Engineering applications. It is so important to realize that regardless of how accurate or even perfect our application codes may be to transport particles, without accurate nuclear data we are in a “Garbage In = Garbage Out” situation.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

AI-powered topic modeling: comparing LDA and BERTopic in analyzing opioid-related cardiovascular risks in women

Topic modeling is a crucial technique in natural language processing (NLP), enabling the extraction of latent themes from large text corpora. Traditional topic modeling, such as Latent Dirichlet Allocation (LDA), faces limitations in capturing the semantic relationships in the text document although it has been widely applied in text mining. BERTopic, created in 2022, leveraged advances in deep learning and can capture the contextual relationships between words. In this work, we integrated Artificial Intelligence (AI) modules to LDA and BERTopic and provided a comprehensive comparison on the analysis of prescription opioid-related cardiovascular risks in women. Opioid use can increase the risk of cardiovascular problems in women such as arrhythmia, hypotension etc. 1,837 abstracts were retrieved and downloaded from PubMed as of April 2024 using three Medical Subject Headings (MeSH) words: “opioid,” “cardiovascular,” and “women.” Machine Learning of Language Toolkit (MALLET) was employed for the implementation of LDA. BioBERT was used for document embedding in BERTopic. Eighteen was selected as the optimal topic number for MALLET and 23 for BERTopic. ChatGPT-4-Turbo was integrated to interpret and compare the results. The short descriptions created by ChatGPT for each topic from LDA and BERTopic were highly correlated, and the performance accuracies of LDA and BERTopic were similar as determined by expert manual reviews of the abstracts grouped by their predominant topics. The results of the t-SNE (t-distributed Stochastic Neighbor Embedding) plots showed that the clusters created from BERTopic were more compact and well-separated, representing improved coherence and distinctiveness between the topics. Our findings indicated that AI algorithms could augment both traditional and contemporary topic modeling techniques. In addition, BERTopic has the connection port for ChatGPT-4-Turbo or other large language models in its algorithm for automatic interpretation, while with LDA interpretation must be manually, and needs special procedures for data pre-processing and stop words exclusion. Therefore, while LDA remains valuable for large-scale text analysis with resource constraints, AI-assisted BERTopic offers significant advantages in providing the enhanced interpretability and the improved semantic coherence for extracting valuable insights from textual data.

Research & Experimental Medicine↗