Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data Distribution”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

EJFAT Scientific Perspective

Presented new computing model to the test by deploying the EJFAT system alongside a data-stream processing framework running the production-level CLAS12 event reconstruction application. In this experiment, a continuous stream of CLAS12 Level-1 identified events was processed in real-time using the EJFAT load balancer, distributing the workload across 90 computing nodes located across the U.S. This marks the first-ever large-scale, real-time distributed data stream processing experiment, demonstrating that scientific data-streaming pipelines can efficiently scale across four dimensions, thanks to EJFAT’s advanced hardware and software capabilities.

Gyurjyan, Vardan [Thomas Jefferson National Accele↗

Artificial intelligence models, photos, and data associated with the manuscript “Quantifying Streambed Grain Size, Uncertainty, and Hydrobiogeochemical Parameters Using Machine Learning Model YOLO” (v2)

This data package is associated with the manuscript “Quantifying Streambed Grain Size, Uncertainty, and Hydrobiogeochemical Parameters Using Machine Learning Model YOLO” published in Water Resources Research (Chen et al., 2024). This data package includes the training, validation, testing, and prediction data used by the artificial intelligence (AI) model for automated grain size and hydro-biogeochemistry quantification using streambed photos. The grain size data are extracted for each photo using You Look Only Once (YOLO), a pre-trained object detection model. This data package was originally published in October 2023. It was updated August 2025 (v2; new and modified files). File and folder names were not revised to indicate changes. See the change history section in the readme for more details. Please see flmd.csv for a list of all files contained in this data package and descriptions for each. Please see dd.csv for a data dictionary that defines the column headers of .csv files in the data package. This dataset is comprised of one data folder containing (1) file-level metadata; (2) data dictionary; (3) readme; and (4) six subfolders. Subfolders 1 to 4 include the training, validation, testing, and prediction data. Subfolder 5_Summary includes the summary results of different combinations of training, validation, testing, and prediction data. Subfolder 6_SupplementalData includes additional data downloaded from public sources (Kaufman et al., 2023a; Kaufman et al., 2023b; Garefalakis et al., 2023; Mair et al., 2024; https://github.com/river-corridors-sfa/Geospatial_variables). In total, the data package includes 110 folders and 44,283 files. These files include 9,047 .jpg photos, 1 .png photo, 3 .tif photos; 26,639 photo labels and individual grain sizes and probability from AI (.txt); 8,447 grain size distribution data (.dat); and 126 CSV files for results summary, and 14 required metadata files (.xlsx). The summary CSV files contain 68 columns and approximately 2,200 rows that represent photo names, site locations, recording time, GPS coordinates, grains sizes (D10, D50, D60, and D84), number of grains, and additional hydro-biogeochemical data such as water depth, flow velocity, Manning’s coefficient, friction factor, hydraulic conductivity, permeability, streambed interstitial velocity magnitude, mass transfer rate, and nitrate uptake velocity. The photos were obtained from 75 sites in the Yakima River Basin and the Columbia River shorelines, and other associated data from samples and sensors obtained when the photos were taken are publicly available (Fulton et al. 2022; Grieger et al. 2023). All files are .csv, .txt, .dat, .jpg, or .pdf. We acknowledge the Yakama Nation as owners and caretakers of the lands where we collected some of these data. We thank the Confederated Tribes and Bands of the Yakama Nation Tribal Council and Yakama Nation Fisheries for working with us to facilitate sample collection and optimization of data usage according to their values and worldview.

54 ENVIRONMENTAL SCIENCES↗

Dynamics modeling of molten salt reactor with reduced and expanded representations of delayed neutron precursors

Molten salt reactors (MSRs) present unique challenges in dynamic behavior due to the mobility of their fuel. In these reactors, delayed neutron precursors (DNPs) drift with the fuel circulation through the primary loop. As a result, a fraction of DNPs decays outside the core, effectively reducing the available delayed neutron population for reactivity control. Consequently, precise modeling of the distribution and behavior of DNPs is critical for accurate reactor dynamics simulations. In this study, the System Dynamics Analysis Tool (SDAT) was used to simulate a thermal-spectrum MSR under steady-state conditions and following transients. The effects of using reduced and expanded representations of DNPs with fewer or more groups than the conventional 6-group model were investigated. Their impact on the simulated distribution of precursors in the primary loop, reactivity loss value, and reactor response to transients was analyzed. Simulation results showed that reduced models lead to the loss of the actual DNPs distribution data, resulting in less accurate estimates of reactivity loss. Reactor power predictions using these reduced models showed significant deviations compared to those using the conventional 6-group model in transient simulations. Expanded models offered a more accurate representation of the distribution of DNPs and reactivity loss estimates. Reactor power predictions using expanded models showed minimal deviation from the conventional 6-group model during the simulated transients.

analysis↗

SITCOMTN-154: Initial studies of photometric redshifts with LSSTComCam from DP1

This technote holds reports based on the first analyses of the Data Preview 1 (DP1) data by the Science Unit for photometric redshifts. Although photometric redshifts are not an official DP1 data product, the "Photo-z Science Unit" generated photo-z estimates for every galaxy in DP1 using the available multi-band imaging on a best-effort basis. This work included developing training and test datasets by matching DP1 data to high-quality reference redshifts obtained with spectroscopy, Grism data, and multi-band photometry. The Science Unit used the RAIL software package to make photometric redshift estimates using eight different algorithms, developed simple scientific performance metrics, used those metrics to explore how the performance of the algorithms varied with configuration changes, derived more optimized configurations of the algorithms and tested the performance of those configurations. This work, the resulting data products and expected data distribution mechanism are all described there.

79 ASTRONOMY AND ASTROPHYSICS↗

Updates to the n+ 63,65 Cu Angular Distributions [Slides]

The performance of the benchmark suite is very sensitive to changes in the angular distribution data. The quasi-differential measurements performed by Blain et al at RPI provide a valuable constraint. Furthermore, ENDF/B-VIII.0 disagrees with their measurements consistently at 300 keV (the transition between RRR and high energy). The next step is to conservatively adjust the Legendre coefficients near 300 keV, validating performance against both critical benchmarks and the RPI quasi-differential measurements.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

The persistence and conversion of coastal foredune and swale vegetation community distributions 63 years later

Vegetation shifts can directly alter habitat dynamics and indirectly impact habitat stability relative to disturbance response. Barrier island dune habitats exhibit spatiotemporally dynamic topography that is affected by vegetation. However, vegetation distribution data can be rare and vegetation persistence is largely unknown such that species turnover over in communities can occur unnoticed. This is true despite concerns and documented cases of woody encroachment related to climate change in these ecogeomorphic habitats where physical stability to resist storm erosion varies with vegetation distribution and density. In 1956 and 1957, the vegetation of Island Beach State Park, NJ, was mapped as a permanent record for subsequent ecological study. In June 2020, we remapped the vegetation of 5.4 of 17 km north to south, seaward of the thicket community boundary. We maintained the same classification system as the historic record, physically mapping vegetation patches via GPS. We quantified changes in thicket, heather, and grass community distribution between the two time periods. Habitat persistence and conversion varied in the 63 years. Heath communities saw 80%–97% habitat loss in conversion to woody thicket. Conversely, thicket community distribution drastically increased, replacing heath where it was previously prevalent. This represents the first known documented instances of woody encroachment for Morella pensylvanica. When they did not expand, thicket communities receded landward or underwent turnover to an invasive species. Woody species of interest for dune stabilization occupied areas of similar habitat characteristics. Dune vegetation distribution persistence from 1957 to 2020 was relatively consistent and stable with the exception of heath habitat conversion. Barrier island stability is directly related to vegetation stability such that understanding where community shifts might occur over time can aid in managing and modeling efforts surrounding these dynamic ecogeomorphic habitats.

54 ENVIRONMENTAL SCIENCES↗

FEDERATED LEARNING ON STOCHASTIC NEURAL NETWORKS

Federated learning is a machine learning paradigm that leverages edge computing on client devices to optimize models while maintaining user privacy by ensuring that local data remain on the device. However, since all data are collected by clients, federated learning is susceptible to latent noise in local datasets. Factors such as limited measurement capabilities or human errors may introduce inaccuracies in client data. To address this challenge, we propose the use of a stochastic neural network as the local model within the federated learning framework. Stochastic neural networks not only facilitate the estimation of the true underlying states of the data but also enable the quantification of latent noise. We refer to our federated learning approach, which incorporates stochastic neural networks as local models, as federated stochastic neural networks. In this work we will present numerical experiments demonstrating the performance and effectiveness of our method, particularly in handling nonindependent and identically distributed data.

97 MATHEMATICS AND COMPUTING↗

Preventing Failures By Dataset Shift Detection in Safety-Critical Graph Applications

Dataset shift refers to the problem where the input data distribution may change over time (e.g., between training and test stages). Since this can be a critical bottleneck in several safety-critical applications such as healthcare, drug-discovery, etc., dataset shift detection has become an important research issue in machine learning. Though several existing efforts have focused on image/video data, applications with graph-structured data have not received sufficient attention. Therefore, in this paper, we investigate the problem of detecting shifts in graph structured data through the lens of statistical hypothesis testing. Specifically, we propose a practical two-sample test based approach for shift detection in large-scale graph structured data. Our approach is very flexible in that it is suitable for both undirected and directed graphs, and eliminates the need for equal sample sizes. Using empirical studies, we demonstrate the effectiveness of the proposed test in detecting dataset shifts. We also corroborate these findings using real-world datasets, characterized by directed graphs and a large number of nodes.

97 MATHEMATICS AND COMPUTING↗

Establishing Pb-203 production from electrodeposited Tl targets at Brookhaven National Laboratory

Background: Promising developments in Pb-212 radiopharmaceutical therapies have increased demand for Pb-203 diagnostic agents. Building on previous work from various isotope production facilities, this study optimized Pb-203 production from electrodeposited Tl targets at Brookhaven National Laboratory (BNL). The additional supply of Pb-203 may help meet growing preclinical and clinical demands. Results: Two Tl targets were irradiated at the Brookhaven Linac Isotope Producer facility with 30 ± 1 MeV protons, measured using previously published cross section data. Distribution coefficients for Pb Resin in acetate media were investigated for both Na + and K + cations, where potassium acetate was ~ 4 times more effective at stripping Pb from the Pb Resin. The Tl electrodeposition was optimized to deposit 350 mg of Tl (~ 60 mg/cm 2 ) on Au backing in under 6 h. The proposed separation process was completed in < 1.5 h and achieved > 98% and 92 ± 3% recovery of Tl and Pb, respectively, with an overall Tl-Pb separation factor of 6 × 10 5 . The experimentally measured half-life of Pb-203 was 52.4 ± 0.7 h, agreeing with 51.93 ± 0.02 h reported by the National Nuclear Data Center. The radioisotopic purity of the Pb fraction at 24 h post end of bombardment (EOB) from a 24 h irradiation was 66% Pb-203, 28% Pb-201, and 6% Pb-200. Following chemical separation, the Pb-203 produced in this work (21 MBq Pb-203 EOB) achieved apparent molar activities of 10 ± 5 and 0.9 ± 0.5 GBq/µmol for [ 203 Pb]Pb-DOTAM and [ 203 Pb]Pb-DO3A, respectively, decay corrected to EOB. Data derived from this work suggests BNL can produce > 10’s GBq Pb-203 with > 99% radiochemical and radioisotopic purity from Tl-205 for worldwide distribution. Conclusions: The production and separation of Pb-203 from natural Tl target material was successfully demonstrated at BNL. Existing methods were adapted and optimized for the facilities at BNL. Results from this work will guide future large-scale Pb-203 production opportunities at BNL for clinical applications.

38 RADIATION CHEMISTRY, RADIOCHEMISTRY, AND NUCLEA↗

Learning Coagulation Processes With Combinatorial Neural Networks

Abstract Simulating the evolution of a coagulating aerosol or cloud of droplets in a key problem in atmospheric science. We present a proof of concept for modeling coagulation processes using a novel combinatorial neural network (CombNN) architecture. Using two types of data from a high‐detail particle‐resolved aerosol simulation, we show that CombNN models outperform standard neural networks and are competitive in accuracy with traditional state‐of‐the‐art sectional models. These CombNN models could have application in learning coarse‐grained coagulation models for multi‐species aerosols and for learning coagulation models from observed size‐distribution data.

54 ENVIRONMENTAL SCIENCES↗

Scale-up Unlearnable Examples Learning with High-performance Computing

Recent advancements in AI models, like ChatGPT, are structured to retain user interactions, which could inadvertently include sensitive healthcare data. In the healthcare field, particularly when radiologists use AI-driven diagnostic tools hosted on online platforms, there is a risk that medical imaging data may be repurposed for future AI training without explicit consent, spotlighting critical privacy and intellectual property concerns around healthcare data usage. Addressing these privacy challenges, a novel approach known as Unlearnable Examples (UEs) has been introduced, aiming to make data unlearnable to deep learning models. A prominent method within this area, called Unlearnable Clustering (UC), has shown improved UE performance with larger batch sizes but was previously limited by computational resources (e.g., a single workstation). To push the boundaries of UE performance with theoretically unlimited resources, we scaled up UC learning across various datasets using Distributed Data Parallel (DDP) training on the Summit supercomputer. Our goal was to examine UE efficacy at high-performance computing (HPC) levels to prevent unauthorized learning and enhance data security, particularly exploring the impact of batch size on UE’s unlearnability. Utilizing the robust computational capabilities of the Summit, extensive experiments were conducted on diverse datasets such as Pets, MedMNist, Flowers, and Flowers102. Our findings reveal that both overly large and overly small batch sizes can lead to performance instability and affect accuracy. However, the relationship between batch size and unlearnability varied across datasets, highlighting the necessity for tailored batch size strategies to achieve optimal data protection. The use of Summit’s high-performance GPUs, along with the efficiency of the DDP framework, facilitated rapid updates of model parameters and consistent training across nodes. Our results underscore the critical role of selecting appropriate batch sizes based on the specific characteristics of each dataset to prevent learning and ensure data security in deep learning applications. The source code is publicly available at https: // github. com/ hrlblab/ UE_ HPC .

Zhu, Yanfan [Vanderbilt University, Nashville, TN,↗

C 12 ( n , n 1 ′ γ ) partial γ -ray cross section measured using the GENESIS array

Improved neutron inelastic scattering cross sections have repeatedly been identified as a top priority nuclear data need, important for basic science and a range of applications in nuclear energy, stockpile stewardship, and proliferation detection. For the C 12 ( n , n ′ γ ) reaction in particular, recent measurements have unveiled some structural discrepancies, demonstrating incongruities among themselves and in relation to the ENDF/B-VIII.0 nuclear data evaluation. To help resolve these disagreements, a measurement was performed at the 88-Inch Cyclotron at Lawrence Berkeley National Laboratory using a broad-spectrum neutron beam and a 99.8% pure natural carbon target. The Gamma Energy Neutron Energy Spectrometer for Inelastic Scattering (GENESIS) was employed to measure energy-differential γ -ray emission spectra as a function of incident neutron energy in the energy range of 5.5 to 16.7 MeV. The C 12 partial γ -ray cross sections were extracted at 63 ∘ , 122 . 5 ∘ , and 150 ∘ with respect to the incoming neutron beam and integrated using angular distribution data available in the literature. The data show agreement with a recent literature measurement and evaluation from 11 to 15 MeV, but indicate a larger cross section for incident neutron energies between 5.5 and 8.5 MeV. The measured relative angular distributions are also reported and were found to agree with evaluation. Published by the American Physical Society 2025

Gordon, J. M. (ORCID:0009000789886897)↗

Towards a Robust Adaptive Digital Twin for Fusion Applications

The development of a digital twin system for fusion applications is essential for enhancing the prediction, analysis, and optimization of complex plasma processes. Machine learning (ML), particularly deep learning has demonstrated strong capabilities in modeling such highly nonlinear and intricate systems. However, two critical challenges limit the deployment of deep learning-based digital twins: Uncertainty Quantification (UQ) and data drift. UQ is vital for ensuring trustworthy predictions, especially in decision-support scenarios. Additionally, data-driven models are often sensitive to changes in the underlying data distribution, such as shot-to-shot variations in fusion experiments, which can lead to performance degradation over time. To address these challenges, we are developing an uncertainty-aware, adaptive digital twin framework. Our approach incorporates deep learning models enhanced with Gaussian Process approximations for predictive uncertainty estimation, coupled with an online learning mechanism that enables continuous model adaptation to new experimental data. This adaptive capability allows the data driven models to respond effectively to evolving plasma behaviors and equipment conditions. Specifically, to mitigate the effects of shot-to-shot drift, our system updates itself incrementally as new data becomes available, improving both robustness and fidelity. Our vision is to evolve this data driven model into a self-sustaining digital twin system that leverages UQ based feedback to continuously refine itself and potentially support real-time decision making. This presentation will cover a brief background on uncertainty quantification for ML, our ongoing effort on development of UQ capabilities for ML, our data science pipeline from data collection to model development and analysis and online learning framework for modeling coil deflection at DIII-D. I will also briefly touch upon opportunities and challenges in development of digital twin framework.

Sammuli, Brian [General Atomics]↗

Effects of Void Morphology on Detonation Initiation of PBX 9502

Accurate predictive modeling of high explosives in abnormal conditions, such as a fire, is critical to personnel safety. Modeling depends heavily on precise physical characterization of the high explosive in question. The aim of this thesis was to characterize the detonation response to shock and the microstructure at high-temperature (250°C) of PBX 9502 and use these data to attempt to model these results using a new reactive burn model called SURF (Scaled Uniform Reactive Front). In order to characterize the shock response of PBX 9502 at high-temperature a 1D gas gun experiment was designed and executed that provided shock to detonation transition data. These data were then used to do a preliminary calibration of SURF for PBX 9502 at 250°C. Small-angle neutron scattering (SANS) was employed to characterize the void morphology of PBX 9502 as a function of temperature. This was done in order to feed a modified version of SURF called physically informed - SURF (π-SURF) which uses the void size and quantity distribution data as measured by SANS along with the deflagration and detonation characteristics of a high explosive to inform the SURF model. The π-SURF model was then used to model both ambient and high-temperature PBX 9502 SDT response and the results were compared to the physically measured results. It was found that π-SURF works in ambient cases. At high-temperature, however, it was determined that not enough information is known about hot spot formation to use the π-SURF model for high-temperature scenarios.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

Multidimensional scaling informed by F -statistic: Visualizing grouped microbiome data with inference

Multidimensional scaling (MDS) is a widely used dimensionality reduction technique in microbial ecology data analysis that captures the multivariate structure of the data while preserving pairwise distances between samples. While improvements in MDS have enhanced the ability to reveal group-specific data patterns, these MDS-based methods require prior assumptions for inference, limiting their application in general microbiome analysis. Here, in this study, we introduce a new MDS-based ordination method, “F-informed MDS,” which configures the data distribution based on the F-statistic, the ratio of dispersion between groups sharing common and different characteristics. Using semisynthetic datasets, we demonstrate that the proposed method is robust to hyperparameter selection while maintaining statistical significance throughout the ordination process. Various quality metrics for evaluating dimensionality reduction confirm that F-informed MDS is comparable to state-of-the-art methods in preserving both local and global data structures. Its application to a diatom-associated bacterial community suggests the role of this new method in interpreting the community’s response to the host. Our approach offers a well-founded refinement of MDS that aligns with statistical test results, which can be beneficial for broader multidimensional data analyses in microbiology and ecology. This new visualization tool can be incorporated into standard microbiome data analyses.

Biological and medical sciences↗

Measurement of particulated matter (PM1, PM2.5, PM10) using a PM sensor during the SAIL campaign in Gothic, CO and Mt. Crested Butte, CO

We deployed another PM sensor (Modulair-PM, QuantAQ) to measure the mass concentration of particulate matter (PM) for three size cuts at both the main site (M1) and supplementary site (S2) of the SAIL from 14 June 2022 to 14 June 2023. The instruments provide the mass concentration of PM1, PM2.5 and PM10. Mass concentrations were calculated based on measurements made with a nephelometer and an OPC. We used Quant-AQ's algorithm [please see their documentation here and the references within] to determine the mass concentrations reported. Here, we present the time series of sample relative humidity, temperature, pressure, and the mass concentration of PM1, PM2.5 and PM10. We also present the size distribution of the aerosol particles within a diameter range of 0.35 - 40 micron, measured by the OPC of the pm-modulair. However, we encourage caution while using the size distribution data, since it is operated only with the factory calibration. Abstract and description of the campaign can be found here : https://www.arm.gov/research/campaigns/amf2022ssb.

54 ENVIRONMENTAL SCIENCES↗

Orchestration of materials science workflows for heterogeneous resources at large scale

In the era of big data, materials science workflows need to handle large-scale data distribution, storage, and computation. Any of these areas can become a performance bottleneck. We present a framework for analyzing internal material structures (e.g., cracks) to mitigate these bottlenecks. We demonstrate the effectiveness of our framework for a workflow performing synchrotron X-ray computed tomography reconstruction and segmentation of a silica-based structure. Our framework provides a cloud-based, cutting-edge solution to challenges such as growing intermediate and output data and heavy resource demands during image reconstruction and segmentation. Specifically, our framework efficiently manages data storage, scaling up compute resources on the cloud. The multi-layer software structure of our framework includes three layers. A top layer uses Jupyter notebooks and serves as the user interface. A middle layer uses Ansible for resource deployment and managing the execution environment. A low layer is dedicated to resource management and provides resource management and job scheduling on heterogeneous nodes (i.e., GPU and CPU). At the core of this layer, Kubernetes supports resource management, and Dask enables large-scale job scheduling for heterogeneous resources. The broader impact of our work is four-fold: through our framework, we hide the complexity of the cloud’s software stack to the user who otherwise is required to have expertise in cloud technologies; we manage job scheduling efficiently and in a scalable manner; we enable resource elasticity and workflow orchestration at a large scale; and we facilitate moving the study of nonporous structures, which has wide applications in engineering and scientific fields, to the cloud. While we demonstrate the capability of our framework for a specific materials science application, it can be adapted for other applications and domains because of its modular, multi-layer architecture.

97 MATHEMATICS AND COMPUTING↗

Utah FORGE: Well 16B(78)-32 Distributed Temperature Sensing Data from April and May 2024

This dataset includes Neubrex Energy Services fiber optic distributed temperature sensing (DTS) data from well 16B(78)-32 during stimulation and circulation, including interaction with well 16A(78)-32, during April and May 2024. The DTS data are stored in HDF5 file format and are accompanied by a PowerPoint report on the study. All times in this dataset are in UTC. Depths are in MD relative to Kelly Bushing Height, and temperatures are in degrees Fahrenheit. All DTS measurements were made using a Yokogawa 3000DTSX Distributed Temperature Sensing Interrogator Unit, with a spatial sampling interval of 3.28 feet and a temporal sampling rate of 129 seconds. The third-party Pressure-Temperature Gauge data should be used with caution after April 20, 2024, as its performance is not considered reliable beyond this date.

15 GEOTHERMAL ENERGY↗