Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “pre-processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Deeplynx Dag Repository

The DeepLynx DAG repository will contain several Airflow DAGs (Directed Acyclic Graphs) which will be used in the context of DeepLynx's deployed Apache Airflow instance. These DAGs will be used for multiple data management tasks for DeepLynx data, including but not limited to: - bringing data from various sources and tools into DeepLynx - managing sequential data workflows, such as running Python scripts on data to perform analysis and returning the results to DeepLynx - performing any necessary transformation or pre-processing on data coming into DeepLynx from external sources or out of DeepLynx to go to external applications

Brownlee, JarenM.↗

MAPLE v.1.0

SAND2025-00659O MAPLE is a software tool that uses epigenomic data to predict gene expression. MAPLE uses a set of epigenomic modifications to determine the effect on gene expression in a specific subset of species. The algorithm can be trained on additional species and epigenomic modifications, enhancing its predictive capabilities. EAGLE employs a hybrid neural network architecture, featuring a convolutional front-end and a multi-head attention layer, to process pre-processed signal data as input. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy's National Nuclear Security Administration under contract DE-NA0003525.

Davis IV, Warren↗

Seismic H2: Version 1.0

Seismic-H2 is an integrated software package for geological hydrogen reservoir simulation, optimization, and leakage monitoring. The package includes multiple components: (1) code used for modeling seismic wave propagation in 3D heterogeneous elastic media based on finite-difference method to support detection of geological hydrogen storage reservoir leakage; (2) 3D reservoir simulations of leaks from an underground reservoir and 3D simulations of saline aquifers and depleted gas reservoirs; (3) seismic monitoring costs of passive and active seismic monitoring required for UHS; (4) rock physics calculations and interpolations for converting the reservoir simulations from part (2) into the elastic media models in part (1); (5) pre-processing seismic data; and lastly (6), a GUI interface that combines these different components.

Creasy, Neala↗

HydraGNN_Predictive_GFM_2026 - Ensemble of predictive graph foundation models for atomistic materials modeling

This release contains data and parameters of HydraGNN-based graph foundation models trained as a result of the work published in the pre-print "Exascale Multi-Task Graph Foundation Models for Imbalanced, Multi-Fidelity Atomistic Data" by M. Lupo Pasini et al. (https://arxiv.org/abs/2604.15380). We jointly train on 16 open first-principles datasets (544+ million structures covering 85+ elements) using a multi-task architecture with per-dataset heads and a scalable ADIOS2/DDStore data pipeline. On Frontier, we execute six large-scale DeepHyper hyperparameter optimization campaigns in FP64 and promote the top-performing message-passing models to sustained 2,048-node training, yielding a PaiNN-based lead model. The version of HydraGNN used to generate the outputs provided in this release is HydraGNN v5.0 (https://github.com/ORNL/HydraGNN/releases/tag/v5.0) The list of datasets used for the training of the graph foundation model is the following: 1) Alexandria [1] 2) ANI1x [2] 3) MPTrj [3] 4) Open Catalyst 2020 (OC20) [4] 5) Open Catalyst 2022 (OC22) [5] 6) Open Catalyst 2025 (OC25) [6] 7) Open Direct ir Capture 2023 (ODAC23) [7] 8) Open Materials 2024 (OMat24) [8] 9) Open Molecules 2025 (OMol25) [9] 10) OMol25-neutral (subset of OMol25 that contains only molecules with zero total charge) 11) OMol25-non-neutral (subset of OMol25 that contains only molecules with non-zero total charge) 12) Open Polymers 2026 (OPoly2026) [10] 13) Nabla2DFT [11] 14) QCML [12] 15) QM7X [reference 13] 16) transition1x [14] Dataset references: [1] J. Schmidt et al., “A dataset of 175k stable and metastable materials calculated with the PBEsol and SCAN functionals,” Scientific Data, vol. 9, p. 64, 2022. [2] J. S. Smith et al., “The ANI-1ccx and ANI-1x data sets, coupled-cluster and density functional theory properties for molecules,” Scientific Data, vol. 7, p. 134, 2020. [Online]. Available: https: //www.nature.com/articles/s41597-020-0473-z [3] A. Jain et al., “Commentary: The Materials Project: A materials genome approach to accelerating materials innovation,” APL Materials, vol. 1, no. 1, p. 011002, 07 2013. [Online]. Available: https://doi.org/10.1063/1.4812323 [4] L. Chanussot et al., “Open catalyst 2020 (oc20) dataset and community challenges,” ACS Catalysis, vol. 11, no. 10, pp. 6059–6072, 2021. [Online]. Available: https://doi.org/10.1021/acscatal.0c04525 [5] K. Tran et al., “Open catalyst 2022 (oc22) dataset and challenges for oxidation electrocatalysts,” ACS Catalysis, vol. 13, no. 5, pp. 3066–3084, 2023. [Online]. Available: https://doi.org/10.1021/acscatal.2c05426 [6] S. J. Sahoo et al., “The open catalyst 2025 (oc25) dataset and models for solid-liquid interfaces,” arXiv preprint arXiv:2509.17862, 2025. [Online]. Available: https://arxiv.org/abs/2509.17862 [7] A. Sriram et al., “The open DAC 2023 dataset and challenges for sorbent discovery in direct air capture,” ACS Central Science, vol. 10, no. 5, pp. 923–941, 2024. [8] L. Barroso-Luque et al., “Open materials 2024 (omat24) inorganic materials dataset and models,” 2024. [Online]. Available: https://arxiv.org/abs/2410.12771 [9] D. S. Levine et al., “The open molecules 2025 (OMol25) dataset, evaluations, and models,” 2025. [Online]. Available: https://arxiv.org/abs/2505.08762 [10] D. S. Levine et al., The open polymers 2026 (OPoly26) dataset and evaluations,” arXiv preprint arXiv:2512.23117, 2025. [Online]. Available: https://arxiv.org/abs/2512.23117 [11] K. Khrabrov et al., “Nabla2dft: A universal quantum chemistry dataset of drug-like molecules and a benchmark for neural network potentials,” in NeurIPS 2024 Datasets and Benchmarks Track, 2024. [Online]. Available: https://openreview.net/forum?id=ElUrNM9U8c [12] S. Ganscha et al., “The QCML dataset, quantum chemistry reference data from 33.5M DFT and 14.7B semi-empirical calculations,” Scientific Data, vol. 12, p. 406, 2025. [13] J. Hoja et al., “QM7-X, a comprehensive dataset of quantum-mechanical properties spanning the chemical space of small organic molecules,” Scientific Data, vol. 8, p. 43, 2021. [Online]. Available: https://www.nature.com/articles/s41597-021-00812-2 [14] M. Schreiner et al., “Transition1x - a dataset for building generalizable reactive machine learning potentials,” Scientific Data, vol. 9, p. 779, 2022. The folder "datasets_ADIOS2_format" contains the set of pre-processed datasets in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used for the development and training of GFMs in this work. The "datasets_ADIOS2_format" directory contains 2 sub-directories, one for the version "v1" of the datasets and one for the version "v2" of the datasets. The version "v1" of the datasets provides values of the total energy as they are extracted from the original data as it was released by the respective institutions. The version "v2" of the datasets provides values of the energy that have been realigned. The realignment was performed by training a linear regression model that predicts the total energy as a function of the chemical composition of the atomistic structure, and then subtract such prediction from the original value of the total energy. Both folders "v1" and "v2" contain 16 sub-directories, each corresponding to an ADIOS2-formatted dataset The folder "DeepHyper-results" contains the configurational files and model's parameters for all the 186 HPO trials that were successfully completed by the scalable hyperparameter optimization (HPO) runs on Frontier. The content of the folder "DeepHyper-results" I structured as follows: 1) task-list.txt: list of mpnn name, jobid, and deephyper task id 2) gfm_${MPNN}_${JOBID}_0.${TASKID}: run directory with checkpoint files 3) gfm_${MPNN}: deephyper summary directory (*.csv) for each specific MPNN type 4) deephyper-experiment-${JOBID}: output and error logs for each job The file "deephyper-sorted.csv" contains the details of each HydraGNN model built and tested by HPO, obtained by merging the (*.csv) filed from each HPO run executed. Out of all the HPO trials, we selected 10 to continue the training of the respective HydraGNN models. Due to limited computational budget available in the LRN070 allocation we could not complete the training till convergence for all these 10 selected models. The folder "models" contains multiple sub-folders, one per each HydraGNN model trained. Each model sub-folder contains the parameters of each HydraGNN model, with multiple checkpoint-restarts. The list of sub-folders are as follows: 1) multidataset_hpo-BEST1-fp64 2) multidataset_hpo-BEST2-fp64 3) multidataset_hpo-BEST3-fp64 4) multidataset_hpo-BEST4-fp64 5) multidataset_hpo-BEST5-fp64 6) multidataset_hpo-BEST6-fp64 7) multidataset_hpo-BEST7-fp64 8) multidataset_hpo-BEST8-fp64 9) multidataset_hpo-BEST9-fp64 10) multidataset_hpo-BEST10-fp64 Within each one of these folders, additional auxiliary log files are provided with descriptions about how the training proceeded. The lead PaiNN-model is contained inside "multidataset_hpo-BEST6-fp64". The file "mlp_branch_weights" contains the parameters of the multi-layer perceptron (MLP) used to reconcile the predictions of the 16 output decoding heads of the HydragNN architectures. The MLP takes in input the chemical composition of the atomistic structure and predicts averaging weights to linearly mix the predictions of each output decoding head toward consolidating them into a single one. The folder "1.1billion-structure-inference" contains 1.1 billion atomistic structures randomly generated. Each structures is associated with energy and forces predicted with the lead-PaiNN model combined with the MLP model for reconciliation of the multi-branch predictions generated by the 16 output decoding heads. The folder "1.1billion-structure-inference" contains 9,300 (*.tar.gz) subdirectories, one per Frontier compute node used to execute the inference at exascale. Once uncompressed, each (*.tar.gz) subdirectory contains an ADIOS2 (*.bp) file container, where each atomistic structure is stored as a PyTorch-Geometric Data object. The file "export_dataset_environment_variables.sh" contains the environment variables that need to be set before running the HydraGNN code to reproduce the results provided in this dataset release. The code that can be used to load the ADIOS2 files, load HydraGNN models, and run inference is available at: https://github.com/ORNL/HydraGNN/releases/tag/v5.0

36 MATERIALS SCIENCE↗

X-ray Computed Tomography Data of Dense Metallic Components

The data shared in here are X-ray computed tomography (XCT) scans of a hexagonal fuel nozzle in 3 sections with the Metrotom 800 system at the Manufacturing Demonstration Facility (MDF) at Oak Ridge National Laboratory. The data are used in the paper "Tomographic Sparse View Selection using the View Covariance Loss, by Lin et al. (doi:10.1109/TPAMI.2025.36000720), accepted to the international conference on computational imaging (ICCP 2025). Figures 4-7 in the paper describe the part/XCT scan. File name Descriptions: Bottom section: TCR- Single Channeled SRC L 2019-3-18 12-26-41.hdf5 Medium section: TCR- Single Channeled SRC M 2019-3-18 13-8-9.hdf5 Top section: TCR- Single Channeled SRC T 2019-3-18 13-45-39.hdf5 Each hdf5 file contains projection data, and all the relevant X-ray CT scan setting. The full list of included attributes: distance_unit: Units of all distances specified angle_unit : Units of the angles angles: Array of all angles used voxel_size_xy: Baseline recon (if any) has this voxel size in the in-plane direction voxel_size_z: Baseline recon (if any) has this voxel size in the cross-plane direction det_pixel_size_col: Size of the detector pixels in the column dimension det_pixel_size_row: Size of the detector pixels in the row dimension src_iso_dist: Source to iso-center distance iso_det_dist: Iso-center to detector distance det_angle: If the detector is rotated/tilted, this angle corresponds to that value det_row_offset: Center of rotation offset in the vertical direction det_col_offset: Center of rotation offset in the horizontal direction reconstruction: A baseline reconstruction stored as 3D array BHC params: Beam-hardening parameters - Van De Casteel Model - if it has been used to pre-process the projections We also provided a python script (hdf_io.py) that allows the user to read the relevant data from each hdf5 file.

Ziabari, Amir [Oak Ridge National Laboratory]↗

X-ray nano-holotomography reconstruction with simultaneous probe retrieval

In conventional tomographic reconstruction, the pre-processing step includes flat-field correction, where each sample projection on the detector is divided by a reference image taken without the sample. When using coherent X-rays as a probe, this approach overlooks the phase component of the illumination field (probe), leading to artifacts in phase-retrieved projection images, which are then propagated to the reconstructed 3D sample representation. The problem intensifies in nano-holotomography with focusing optics, which, due to various imperfections creates high-frequency components in the probe function. Here, we present a new iterative reconstruction scheme for holotomography, simultaneously retrieving the complex-valued probe function. Implemented on GPUs, this algorithm results in 3D reconstruction resolving twice thinner layers in a 3D ALD standard sample measured using nano-holotomography.

Nikitin, Viktor↗

Data and scripts associated with the manuscript evaluating the hydrologic responses of the Pacific Northwest watersheds to wildfires (v2)

This data package is associated with the publication “Evaluating Post-fire Watershed Response to Varying Burn Severity and Precipitation Regimes Using Fully-distributed and Integrated Hydrologic Models” submitted to Journal of Hydrology (Li et al. 2025). In this study, we employed the Advanced Terrestrial Simulator (ATS), an integrated watershed model that couples surface flow, subsurface flow, and canopy biophysical processes, to investigate post-fire hydrologic responses in a few selected watersheds with varying burn severity.The data package contains the required input data (meteorological forcing, Leaf Area Index, wildfire burn severities, etc.) to run the model, configuration files, the Jupyter notebooks in Python to pre-process and post-process data, the figures in the manuscript, and the modeling output files. The variables include watershed-averaged evapotranspiration, watershed-averaged surface/subsurface/canopy water content, and river discharge at watershed outlet.The data package contains a file-level metadata that lists and describes all the files contained in the data package (ATS_flmd.csv), a data dictionary file that defines columns headers across all csv files contained in the data package (ATS_dd.csv), a data package level readme file (the current file), and four zipped folders.The ‘data’ folder provides data needed to run the model in .h5, .i2s, .xyz, .shp, and .exo formats. The sub-folders are for each data types. The ‘model’ folder provides input files (.xml format) and essential model outputs. Each sub-folder provides the files from each simulated watershed. The ‘notebooks’ folder provides the Jupyter notebooks (.ipynb format) for pre- and post- processing model files, and for producing the figures in the manuscript. The ‘figures’ folder provides the figures associated with manuscript in .pdf and .png formats.The ‘model’ folder and the ‘data’ folder have been split into 5GB-large pieces using the Linux command ‘split -b 5120m model.zip model.zip.’ and ‘split -b 5120m data.zip data.zip.’, respectively. They can be merged back using the Linux command ‘cat model.zip.* > model.zip’ and ‘cat data.zip.* > data.zip’, respectively.

54 ENVIRONMENTAL SCIENCES↗

Building heights and urban canopy parameters for urban modeling

GLObal Building heights for Urban Studies (UT-GLOBUS) is a random forest model based framework that provides a level-of-detail-1 (LoD-1) building height dataset. The primary objective of UT-GLOBUS is not to precisely predict the height and footprint of individual buildings, but rather to offer a functional framework for generating building level information using open-source datasets for modeling applications. Specifically, UT-GLOBUS is tailored to meet the requirements of deriving urban canopy parameters (UCPs) for the multi-layer model within the Weather Research and Forecasting (WRF) model and building heights for the SOLWEIG and SUEWS model. Building-level data is accessible in vector file format (GeoPackage: .gpkg), which can be converted into raster file format (geoTIFF). The vector files employ the Universal Transverse Mercator (UTM) projection. The vector files are compatible with GIS platforms like QGIS and ArcGIS, and can be imported for analysis using programming languages such as Python. We are also providing UCPs required by the multi-layer urban model in the urban WRF in binary file format. Additionally, we provide the urban fractions calculated using ESA world cover dataset (https://esa-worldcover.org/en) for WRF model in binary file format. These binary files can be directly incorporated into the WRF pre-processing system (WPS).

54 ENVIRONMENTAL SCIENCES↗

Site and endmember spectra of terrestrial vegetation and soils for the Colorado Headwaters Ecological Spectroscopy Study, June-July 2025

This dataset provides site and endmember spectra collected during the 2025 Colorado Headwaters Ecological Spectroscopy Study (CHESS) campaign. The site spectra were collected to help validate airborne hyperspectral data acquired by the National Ecological Observatory Network's aerial observation platform (NEON AOP). Endmember spectra were collected to augment existing spectral libraries with additional samples of bare surfaces and non-photosynthetic vegetation. All measurements were acquired with an Analytical Spectral Devices (ASD) FieldSpec4 Hi-Res NG (Next Generation) spectroradiometer, which records radiance at 1nm (nanometer) intervals from the ultraviolet to the short-wave infrared (350-2500 nm). The dataset includes spectra measured at meadow sites where the CHESS team also collected vegetation samples for trait analyses. The site spectra were collected with the ASD FieldSpec4 palm grip attachment using an 8° field-of-view foreoptic. Site spectra are integrated measurements of the entire surface within the foreoptic’s field of view. For site-level spectra, the sun is the illumination source. A Spectralon panel mounted on a tripod was used for instrument optimization and white reference measurements for all site spectra. Site spectra were acquired within two hours of solar noon and within 48 hours of a NEON AOP overflight. Site spectra are labeled by date, sampling area, and site number according to the naming conventions of the CHESS campaign’s data management plan. The dataset also contains endmember spectra in the following categories: photosynthetic vegetation (PV), non-photosynthetic vegetation (NPV), bare (soil/rock), and flowers. Endmember measurements were acquired using either the contact probe or the leaf clip attachments of the ASD FieldSpec4. In these configurations, the bulb inside the spectrometer provides the light source for the measurements. The spectrometer was optimized and white reference measurements were recorded using the circular white pucks attached to the contact probe and leaf clip. Because they do not rely on solar illumination, contact probe and leaf clip measurements were collected during a broader time frame than the palm grip site spectra. Some endmembers were measured at CHESS meadow sites, while others were collected within the larger sampling area or in nearby locations (e.g. Gothic Townsite) with similar characteristics. Radiance, reflectance, and metadata files are split into three subfolders according to measurement type: proximal/palm grip (prx), contact probe (cp), and leaf clip (lc). Radiance spectra are provided in ASD file format (.asd file extension). All ASD files can be opened using the provided scripts. Metadata is provided in two formats: CSV file format (no geolocation) and GEOJSON file format (includes geolocation for each spectra). The dataset includes a set of pre-processed reflectance spectra as CSV files (yyyymmdd_rfl.csv). The python scripts and jupyter notebook used to calculate reflectance spectra from the ASD radiance data is included here and was previously published at: https://doi.org/10.3334/ORNLDAAC/2446. There is also a folder of JPEG photographs corresponding to selected spectra. We include a protocol document with detailed steps for ASD FieldSpec4 assembly and operations. This data additionally contains a file level metadata (flmd.csv) and data dictionary (dd.csv) file. Geospatial information: Geospatial data for mapping measurement site locations are in the files CHESS_polygons_lai_UTM.geojson, CHESS_polygons_shrub_UTM.geojson, and CHESS_polygons_meadow_UTM.geojson in the companion geospatial package for the 2025 CHESS campaign, ‘CHESS 2025: Location data for field observations and sampling’ (Henderson et al., 2026). CHESS Project Description: The Colorado Headwaters Ecological Spectroscopy Study (CHESS) comprised a multi-week airborne remote sensing and field observation campaign in the Upper Gunnison Basin, Colorado, conducted in June and July of 2025. Airborne remote sensing was conducted by the National Ecological Observatory Network Airborne Observation Platform (NEON AOP), concurrent with a field campaign run by the Rocky Mountain Biological Laboratory (RMBL), the Lawrence Berkeley National Laboratory (LBNL) and SLAC National Accelerator Laboratory Watershed Function Science Focus Area (SFA), and NASA-JPL (Jet Propulsion Laboratory) Earth Surface Mineral Dust Source Investigation (EMIT) program. Between June 10 and July 18, 2025, the NEON AOP flight team collected high-resolution aerial imaging spectroscopy and Light Detection and Ranging (LiDAR) data over three domains: the Upper East River (CRBU), Almont Triangle (ALMO), and the Upper Taylor Basin (UPTA). In coordination with the flights, a field campaign acquired ground-truth observations, including observations of vegetation composition, foliar traits, forest demography, and subsurface properties in 18 core sampling areas within the domains. Additional surface water observations were taken at over 380 point locations. All CHESS campaign datasets can be found within the CHESS ESS-DIVE data portal: https://data.ess-dive.lbl.gov/portals/chess. Funding Acknowledgment: This research was carried out at the Jet Propulsion Laboratory, California Institute of Technology, under a contract with the National Aeronautics and Space Administration (80NM0018D0004) and was funded by EMIT Extended Mission Phase E Science.

2018 NEON and 2025 CHESS Campaigns↗

Analytical methods for online data quality assessment

This chapter provides a comprehensive overview of the main steps for algorithmic sensor signal quality assessment, which can enhance the decision-making process for water resource recovery facility (WRRF) operation and optimization. It introduces the concept of redundancy as the basis for data quality assessment. It also explains the typical data processing pipeline, which consists of preliminary analysis, data pre-processing, and specific algorithmic approaches. Each of these processes is presented and discussed in three separate sections. Importantly, this chapter introduces the main approaches for data quality assessment, provides guidelines for selecting the most suitable one and the key performance indicators to evaluate them and explains how to collect metadata through such an algorithmic approach.

Aguado, Daniel↗

Interparticle Characterization of Mechanical Biomass Particle-Particle and Particle-Wall Interactions

The biomass materials industry faces significant challenges in managing material variability and its impact on storage and handling systems. Physical properties such as moisture content, particle size, and density fluctuate considerably, leading to operational issues like bridging and ratholing that disrupt material flow. These variations create a complex cascade effect throughout the process chain, affecting transportation, storage, and conversion processes. The economic consequences of this variability manifest in increased operational costs, maintenance requirements, and system downtime. Environmental factors further complicate the situation, as weather conditions and seasonal availability influence material properties and system performance. Engineers employ specialized equipment design, material characterization protocols, and pre-processing steps like size reduction and homogenization to address these challenges. A critical knowledge gap exists between continuous-level constitutive models and particle-scale behavior. This project developed a novel device to quantify interparticle mechanics between biomass particles, measuring friction and adhesion forces between particles and wall materials. The research focused on corn stover and southern pine forest residue, creating a comprehensive database of particle interactions. This breakthrough enables direct application in particle-based computational modeling, advancing the field's understanding of biomass handling characteristics and supporting the development of more reliable and efficient storage and handling systems. The project's outcomes contribute significantly to understanding biomass's mechanical and flow characteristics, particularly how variability at the particle level affects larger-scale handling operations. This knowledge is crucial for engineering feedstock supply systems that consistently meet quality and cost specifications for various conversion processes. The innovative experimental setup developed through this research represents a significant advancement in biomass characterization methodology. Providing precise measurements of particle-level interactions establishes a foundation for more accurate predictive modeling of bulk material behavior. This enhanced understanding of fundamental particle mechanics enables engineers to anticipate better and address handling challenges before they manifest in full-scale operations. This research opens new avenues for optimizing biomass handling systems through data-driven design approaches. The comprehensive database of particle interactions serves as a valuable resource for future research and development efforts, potentially leading to more efficient and cost-effective biomass processing solutions. This advancement in particle-level mechanics could revolutionize how biomass handling systems are designed and operated, contributing to more sustainable and reliable renewable energy production.

09 BIOMASS FUELS↗

Coal-Waste-Enhanced Filaments for Additive Manufacturing of High-Temperature Plastics and Ceramic Composites

In the United States, coal waste from over a century of mining and burning coal for heat and electricity has accumulated as mountains of coal fly ash and bottom ash and acre-size ponds, coal fines and gob. These materials can be a problem for local communities and water systems. A cost-effective process to utilize high volumes of these coal wastes in a high-value product would be beneficial to those communities by reducing the amount of waste and providing jobs, manufacturing components, and materials from the waste. Many coal-to-products technologies (e.g., carbon fibers, graphene, carbon foam) rely on carefully choosing the starting material and then altering it chemically or thermally to make the products work. Due to the wide variability of composition and coal content in typical coal waste streams, many high-volume coal waste streams are likely to be unsuitable for use in those technologies. Semplastics’ technology has been shown to utilize most types of coal waste successfully without any pre-selection or pre-processing requirements other than a nominal particle-size reduction for wastes like bottom ash. This characteristic of Semplastics’ solution may enable the use of much larger volumes of a wider range of coal wastes than other coal-to-products technologies. In this project, Semplastics leveraged its unique experience with both coal waste (fly ash or coal combustion residuals), resin materials, and 3D printing to develop 3D printer filaments using common coal wastes – bituminous coal fines and fly ash – and researched the feasibility of using other forms of coal waste as fillers. Simple 3D-printed parts were successfully produced from the coal waste enhanced filaments, which were found to have improved strength and stiffness.

01 COAL, LIGNITE, AND PEAT↗

Ring Pull Strain Analysis Version 1.1

This report details an analysis package, Ring Pull Strain Analysis (RPSA), that can be used to present and quantify digital image correlation (DIC) data as it relates to a gaugeless ring pull test. Gaugeless ring pull is a testing technique for mechanical testing of small annular samples, usually cut from a thin-walled tube. DIC data is often necessary for this kind of test because bending moments present on the ring cause a non-uniform strain distribution and localized measurements are necessary. In addition, the annular geometry of a ring lends itself to a polar representation, which is not present with typical DIC analysis methods. RPSA was made to calculate and plot the polar representation of strain from standard pre-processed DIC data of a gaugeless ring pull test. Further analysis can be done on ring pull including a quasi-uniaxial tensile analysis and coating analysis, which are also performed by RPSA. In addition, due to the universality of DIC plotting and ring pull test analysis, RPSA can accommodate a wide variety of tests, though it is tailored for ring pull testing. This report details how RPSA works, including the theory, assumptions, and logic behind the calculations and the structure of the program.

36 MATERIALS SCIENCE↗

Optimizing enzymes for plastic upcycling using machine learning design and high throughput experiments

Plastic use is ubiquitous in the modern world, and polyethylene terephthalate (PET) is one of the most abundantly produced plastics (and the most highly produced polyester), with ~65 million metric tons manufactured annually. To the consumer, PET is likely most recognizable as the plastic used to make beverage bottles. Like many plastics, traditional mechanical or chemical means of PET deconstruction and upcycling are costly and inefficient. Because of these challenges, recycled plastic is generally of lower quality and is more expensive to produce than virgin plastic derived from petroleum. Ultimately, this results in most plastic ending up as waste. We view plastic waste as an underutilized resource which, with the development of more efficient and high-quality recycling processes, could (1) generate significant economic value while (2) decreasing petroleum usage and greenhouse gas emissions, as well as (3) minimizing its negative environmental and health impacts. Biocatalytic recycling, or biomanufacturing the basic building blocks of new plastic from plastic waste, is a promising approach to plastic reuse that complements existing recycling technologies. Recently, biological enzymes capable of breaking down PET have garnered significant attention as an attractive means of dealing with the plastic problem. These enzymes are currently undergoing pilot studies for implementation in industrial-scale enzyme-based recycling. However, there are significant limitations to current enzymes, including the need to perform costly pre-processing of the plastic waste before the enzymes are able to work. Further optimization of these enzymes is necessary to make these technologies competitive, and ultimately incentivise industry-wide adoption of this biology-based green recycling technology. n this work we demonstrate a means to design and generate performant biological enzymes, capable of efficiently deconstructing plastic waste. Specifically, we applied recent advances in artificial intelligence, machine learning, and statistical analysis to design new versions and discover natural enzymes capable of breaking down PET. We focused on optimizing key properties that are important for industrial-scale enzymatic recycling such as pH and thermotolerance. Normal testing of enzymatic plastic-deconstruction is extremely labor intensive and so through this work we also developed a robotic-assisted experimental pipeline capable of characterizing thousands of candidate enzymes. The results of this iterative, AI-guided, multi-discipline approach have led to increases in enzymatic breakdown of over 150X over starting enzymes. This work supports the rapidly developing and transformative field of biocatalytic solutions to environmental problems beyond the discovery and predictive understanding of enzymes for polymer recycling, and has wide implications for tackling numerous energy problems such as carbon capture and fixation (e.g., engineering carbon monoxide dehydrogenase and the rubisco-pathway), biomining (e.g., design of lanthanide-binding proteins) and biomanufacturing (e.g., lignin-deconstruction enzymes).

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Resonance Self-Shielding: Why it is so Important

This paper is one of a series that I am writing to document my 58 years of experience with ENDF and Neutron Transport calculations, beginning when I worked at the National Nuclear Data Center (NNDC), Brookhaven National Laboratory (BNL), from 1967 to 1972. During those years I was the head of the computer unit of NNDC, assigned to develop computer codes to pre-process, view and test ENDF/B data. Since then, I have continued to support the ENDF effort without any official position or monetary compensation, because I realized how important accurate nuclear data is for use in use in our Engineering applications. It is so important to realize that regardless of how accurate or even perfect our application codes may be to transport particles, without accurate nuclear data we are in a “Garbage In = Garbage Out” situation.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

An intelligent Data Delivery Service for and beyond the ATLAS experiment

The intelligent Data Delivery Service (iDDS) has been developed to cope with the huge increase of computing and storage resource usage in the coming LHC data taking. It has been designed to intelligently orchestrate workflows and data management systems, decoupling data pre-processing, delivery, and primary processing in large scale workflows. It is an experiment-agnostic service that has been deployed to serve data carousel (orchestrating efficient processing of tape-resident data), machine learning hyperparameter optimization, active learning, and other complex multi-stage workflows defined via DAG (Directed Acyclic Graph), CWL (Common Workflow Language) and other descriptions, including a growing number of analysis workflows. We will at first introduce some deployed use cases in a summary. Then we will focus on new improvements and use cases under developments in ATLAS, Rubin Observatory and sPHENIX, together with future efforts.

97 MATHEMATICS AND COMPUTING↗

AI-powered topic modeling: comparing LDA and BERTopic in analyzing opioid-related cardiovascular risks in women

Topic modeling is a crucial technique in natural language processing (NLP), enabling the extraction of latent themes from large text corpora. Traditional topic modeling, such as Latent Dirichlet Allocation (LDA), faces limitations in capturing the semantic relationships in the text document although it has been widely applied in text mining. BERTopic, created in 2022, leveraged advances in deep learning and can capture the contextual relationships between words. In this work, we integrated Artificial Intelligence (AI) modules to LDA and BERTopic and provided a comprehensive comparison on the analysis of prescription opioid-related cardiovascular risks in women. Opioid use can increase the risk of cardiovascular problems in women such as arrhythmia, hypotension etc. 1,837 abstracts were retrieved and downloaded from PubMed as of April 2024 using three Medical Subject Headings (MeSH) words: “opioid,” “cardiovascular,” and “women.” Machine Learning of Language Toolkit (MALLET) was employed for the implementation of LDA. BioBERT was used for document embedding in BERTopic. Eighteen was selected as the optimal topic number for MALLET and 23 for BERTopic. ChatGPT-4-Turbo was integrated to interpret and compare the results. The short descriptions created by ChatGPT for each topic from LDA and BERTopic were highly correlated, and the performance accuracies of LDA and BERTopic were similar as determined by expert manual reviews of the abstracts grouped by their predominant topics. The results of the t-SNE (t-distributed Stochastic Neighbor Embedding) plots showed that the clusters created from BERTopic were more compact and well-separated, representing improved coherence and distinctiveness between the topics. Our findings indicated that AI algorithms could augment both traditional and contemporary topic modeling techniques. In addition, BERTopic has the connection port for ChatGPT-4-Turbo or other large language models in its algorithm for automatic interpretation, while with LDA interpretation must be manually, and needs special procedures for data pre-processing and stop words exclusion. Therefore, while LDA remains valuable for large-scale text analysis with resource constraints, AI-assisted BERTopic offers significant advantages in providing the enhanced interpretability and the improved semantic coherence for extracting valuable insights from textual data.

Research & Experimental Medicine↗

Tensor Extraction of Latent Features (TELF)

Tensor ELF is a user-friendly parallel tensor decomposition Python toolbox that includes a suite of machine learning algorithms for CPU and GPU architectures for the analysis of sparse and dense data including utility tools for pre-processing and post-processing.

Eren, Maksim↗