Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data generation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Developing ML/AI Methods for High-Throughput Characterization of Multiple-Sensor Streams of Tokamak Dynamics for High-Speed Control (Final Report)

This project evaluated and developed new mathematical and algorithmic techniques capable of handling (in real-time) the growing amounts of data generated by modern fusion research. While existing numerical linear algebra (NLA) methods provide the backbone to classical data analysis and algorithms, these methods fundamentally do not port to distributed architectures nor do they allow low-latency data reduction for control. Motivated by the needs for modern fusion reactors, this project explored and implemented new numerical methods to characterize plasma dynamics, respond in real-time to discharge evolution, and to process massive-scale data accurately and rapidly more fully. This project links expertise in multiple-sensor diagnostics of tokamak plasma dynamics from Columbia University’s Plasma Physics Laboratory with expertise in massive-scale data reduction and extreme data control algorithms at Columbia University’s Data Science Institute. This interdisciplinary project (i) applied machine learning methods, (ii) implemented a properly-trained neural-network for very fast processing of high-speed plasma videography, and (ii) developed the applied mathematical methods, based on randomized-NLA (rNLA) routines, for data analysis, reduction, and real-time control. The Columbia University High Beta Tokamak-Extended Pulse (HBT-EP) facility provided data to test new algorithms and partnership with Columbia University's Data Sciences Institute evaluated the broader use of new algorithms for many challenging control applications.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Allometric Scaling of Hyporheic Respiration Across Basins in the Pacific Northwest United States

Abstract Hyporheic zones regulate biogeochemical processes in streams and rivers, but high spatiotemporal heterogeneity makes it difficult to predict how these processes scale from individual reaches to river basins. Recent work applying allometric scaling (i.e., power‐law relationships between size and function) to river networks provides a new paradigm for understanding cumulative hyporheic biogeochemical processes. We used previously published model predictions of reach‐scale hyporheic aerobic respiration to explore patterns in allometric scaling across two climatically divergent basins with differing characteristics in the Pacific Northwest, United States. In the model, hydrologic exchange fluxes (HEFs) regulate hyporheic respiration, so we examined how HEFs might influence allometric scaling of respiration. We found consistent scaling behaviors where HEFs were either very low or very high, but differences between basins when HEFs were moderate. Our findings provide initial model‐generated hypotheses for factors influencing allometric scaling of hyporheic respiration. These hypotheses can be used to optimize new data generation efforts aimed at developing predictive understanding of allometries that can, in turn, be used to scale biogeochemical dynamics across watersheds.

59 BASIC BIOLOGICAL SCIENCES↗

A Versatile Simulated Data Transport Layer for in Situ Workflows Performance Evaluation

In situ processing does not only allow scientific applications to face the explosion in data volume and velocity but also to address the time constraints of many simulation-analysis workflows by providing scientists with early insights about their applications at runtime. Multiple frameworks implement the concept of a data transport layer (DTL) to enable such in situ workflows. These tools are very versatile, directly or indirectly access the data generated on the same node, another node of the same compute cluster, or a completely distinct node, and allow data publishers and subscribers to run on the same computing resources or not. This versatility puts on researchers the onus of taking key decisions related to resource allocation and how to transport data to ensure the most efficient execution of their in situ workflows. However, domain scientists and workflow practitioners lack the appropriate tools to assess the respective performance of particular design and deployment options. In this paper we introduce a versatile simulated DTL designed to provide researchers with insights on the respective performance of different execution scenarios of in situ workflows. This open-source, standalone library builds on the SimGrid toolkit and can be linked to any SimGrid-based simulator. It facilitates the evaluation of the performance behavior, at scale, of different data transport configurations and the study of the effects of resource allocation strategies. We demonstrate the scalability, versatility, and accuracy of this simulated DTL by reproducing the execution of two synthetic benchmarks and of a real-world in situ workflow composed of an MPI application and a parallel data analysis. Results of simulations run on a single core show that the proposed library can simulate the interactions of tens of thousands of simulated processes deployed on two interconnected commodity clusters in a few seconds, and the execution by a thousand simulated processes of an in situ workflow in less than three minutes.

Suter, Fred [ORNL] (ORCID:0000000319021955)↗

CTGAN-TVAE

SAND2026-18914O CTGAN-TVAE (Conditional Tabular Generative Adversarial Networks-Tabular Variational Autoencoders) generates extensive sets of variable generation data through a hybrid framework. It enhances latent space representation by combining TVAE's robust feature-embedding with CTGAN's ability to condition categorical variables such as time. CTGAN-TVAE employs a fully connected neural network within a conditional generative adversarial network framework to manage continuous and categorical data effectively, capturing complex feature interactions without needing sequential modeling. This was developed as part of NNSA-MSIPP: Minority Serving Institution Partnership Program, Grant Number DE-NA0004016. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy's National Nuclear Security Administration under contract DE-NA0003525.

Newlun, Cody [Sandia National Lab. (SNL-CA), Liver↗

Using PyBioNetFit to leverage qualitative and quantitative data in biological model parameterization and uncertainty quantification

Data generated in studies of cellular regulatory systems are often qualitative. For example, measurements of signaling readouts in the presence and absence of mutations may reveal a rank ordering of responses across conditions but not the precise extents of mutation-induced differences. Qualitative data are often ignored by mathematical modelers or are considered in an ad hoc manner, as in the study of Kocieniewski and Lipniacki (2013) [Phys Biol 10: 035006], which was focused on the roles of MEK isoforms in ERK activation. In this earlier study, model parameter values were tuned manually to obtain consistency with a combination of qualitative and quantitative data. This approach is not reproducible, nor does it provide insights into parametric or prediction uncertainties. Here, starting from the same data and the same ordinary differential equation (ODE) model structure, we generate formalized statements of qualitative observations, making these observations more reusable, and we improve the model parameterization procedure by applying a systematic and automated approach enabled by the software package PyBioNetFit. We also demonstrate uncertainty quantification (UQ), which was absent in the original study. Our results show that PyBioNetFit enables qualitative data to be leveraged, together with quantitative data, in parameterization of systems biology models and facilitates UQ. These capabilities are important for reliable estimation of model parameters and model analyses in studies of cellular regulatory systems and reproducibility.

59 BASIC BIOLOGICAL SCIENCES↗

Laboratory time series moisture manipulative experiment from sediment across the contiguous US: time series aerobic respiration and geochemistry (v2)

This dataset supports a broader study examining the effects of wetting and drying on hyporheic zone respiration across the contiguous United States (CONUS). The dataset provides data generated from a laboratory moisture manipulation experiment. The contents include time series aerobic respiration and moisture; dissolved oxygen; sediment geochemistry data; and field metadata (including qualitative information on instream and river corridor characteristics). Samples were collected as part of the WHONDRS CONUS-Scale Model-Sample Study (CM). This study was designed following ICON (integrated, coordinated, open, and networked) principles to facilitate a model-experiment (ModEx) iteration approach, leveraging crowdsourced sampling across the CONUS. The data package associated with the CM study is available at https://data.ess-dive.lbl.gov/view/doi:10.15485/1923689. CM sampling began in April 2022 and ended in October 2023. This study uses subsamples from a subset of CM samples collected between June 2022 and June 2023. The original field samples were labeled as CM_###. Subsequent subsamples for this study were labeled as EC_###. The labels from the field samples and the EC subsamples can be mapped directly based on the digits following the prefix and underscore (i.e., EC_001 is a subsample from CM_001). See the critical details section below for more details on sample naming. This data package was originally published in August 2024. It was updated in February 2026 (v2; new and modified files). See the change history section in the readme for more details. For details on how to navigate this data package, see this infographic from the River Corridor SFA https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. This dataset is comprised of one folder of raw Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) data and one main data folder containing (1) file-level metadata; (2) data dictionary; (3) field metadata; (4) readme; (5) field protocol; and a (6) a subfolder with sediment sample data from the incubation experiment. The sample data subfolder contains (1) dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC); (2) total nitrogen (TN); (3) adenosine triphosphate (ATP); (4) percent carbon and nitrogen; (5) effect size; (6) iron (II); (7) gravimetric moisture; (8) respiration rates and raw dissolved oxygen values; (9) specific conductance; (10) pH; (11) temperature; (12) a summary containing median values of each data type for each treatment (wet and dry); (13) methods codes; (14) FTICR-MS methods; and (15) a subfolder of 9.4 Tesla FTICR-MS data. This folder contains three subfolders, one containing the sediment .xml data files, one containing the sediment CoreMS output files, the other containing instructions and scripts for processing the files in CoreMS (https://github.com/EMSL-Computing/CoreMS). All files are .csv, .pdf, .R, .ref, or .xml.

54 ENVIRONMENTAL SCIENCES↗

WHONDRS laboratory time series moisture manipulative experiment from soil core layers across eastern contiguous US: time series aerobic respiration, geochemistry, and aggregates

This dataset supports a broader study examining the effects of wetting and drying on soil layers across the eastern contiguous United States (CONUS). The dataset provides data generated from a laboratory moisture manipulation experiment. The contents include time series aerobic respiration and moisture; dissolved oxygen; sediment geochemistry data; and field metadata. Samples were collected as part of a collaboration between WHONDRS (Worldwide Hydrobiogeochemistry Observation Network for Dynamic River Systems; https://whondrs.pnnl.gov) and MONet (Molecular Observation Network; https://www.emsl.pnnl.gov/monet). The field samples (soil cores) were labeled as MEL_##_COR and subsequent subsamples begin with MEL_##. Additional subsamples were taken for the laboratory experiment and were labeled as EL_##. The labels from the MEL field samples and the EL subsamples can be mapped directly based on the digits following the prefix and underscore (i.e., EL_01 is a subsample from MEL_01). See the critical details section below for more details on sample naming and experimental design.For details on how to navigate this data package, see this infographic from the River Corridor SFA https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About.In addition to this readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes all column/row headers and variable definitions.This dataset is comprised of (1) a folder containing environmental context photos; (2) file-level metadata; (3) data dictionary; (4) field metadata; (5) readme; (6) international generic sample number (IGSN) mapping file; and (7) a subfolder with soil sample data from field samples and the incubation experiment. The sample data subfolder contains (1) effect size; (2) gravimetric moisture from field samples and incubation experiment; (3) respiration rates, raw dissolved oxygen values, and plots; (4) specific conductance, pH, and temperature from the incubation; (5) soil aggregates; (6) a summary containing median values of each data type for each treatment (wet and dry) in the incubation; (7) a summary containing averages for each data type of each soil layer; and (8) methods codes. All files are .csv, .pdf, .jpeg, or .jpg.

54 ENVIRONMENTAL SCIENCES↗

Recent Metallic Fuel Data Recovery in FIPD

The Metallic Fuels Irradiation and Physics Database (FIPD) [1] is an organized collection of metallic fuel test pin data (U-xPu-yZr, = 0 ~ 28; y = 2 ~ 10) and documentation available to industry. FIPD mainly contains three types of data: (1) Fuel pin fabrication data, including fuel slug diameter, fuel slug length, cladding diameter, smear density, etc. (2) Fuel pin operation conditions, including axial distributions for power, temperatures, fluences, burnup, and isotopic densities, etc. and (3) Fuel pin post-irradiation examination (PIE) data, including fission gas release and gas chemistry, profilometry, and neutron radiography, etc. The operating conditions for pins with PIE data available in FIPD span significant ranges across key parameters. The fuel peak burnup extends from less than 5% up to 20 at%. The cladding peak temperature varies from about 490°C to 660°C. Finally, the cladding peak DPA shows a wide range from less than 5 to 120. These broad ranges reflect the diverse testing conditions and operational parameters captured in the available PIE data. More detail about FIPD can be found in ref. [2]. The database development is an ongoing effort covering metallic fuel experiments from the Experimental Breeder Reactor II (EBR-II) and the Fast Flux Test Facility (FFTF). As reported in the ref. [3, 4], most of the PIE data generated during the IFR program [5] has been collected, reviewed, processed, and integrated into FIPD. The most recently added PIE data can be found in ref. [4], which shows the collection of over 95% of the PIE data by the time of this paper. The recent improvements to the database are summarized in this paper.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Stochastic Ensemble Generation for Improved Characterization of Representing Geologic Variability in a Reservoir: IBDP Case Study for SMART Initiative

This document is a poster covering the findings from activities on training data generation, specifically geologic ensemble generation. The generated geologic realizations captured the range of possible permeability distributions of the subsurface at the Illinois Basin - Decatur Project (IBDP) site, based on available well log variabilities. The percentages of reservoirs and baffles in the injection zone and a truncation of baffle permeability led to more variance in the simulations. This will be used to build forward modeling, history matching, and optimization workflows. The geologic realizations were also ranked according to dynamic measures of hydraulic diffusivity, and simulations confirm a greater contrast between the reservoir and the baffles during injection.

stochastic ensemble generation↗

Hydropower Capacity Factor Trends & Analytics for the United States

This data repository contains all code, input data, and data generated for Turner et al. (2024)—“Hydropower capacity factors trending down in the United States”. File descriptions: – hydro-cf-trends-inputs.zip: Full set of input data used in this study, organized for direct entry into “/data” directory of hydro-cf-trends data processing pipeline. – hydro-cf-trends.zip: Full data processing pipeline, coded using the R {targets} framework. This is a snapshot release (v1.0) of the code repository stored at https://code.ornl.gov/turnersw/hydro-cf-trends/. – hydro-cf-trends-results.zip: Provides all dam level results required to reproduce results and graphics in Turner et al. (2024). Dams are identified by the “complxID” (root of the hydropower plant ID in the Existing Hydropower Assets Database, inherited from HILARRI). Results include: • dam_CF_trends.csv: Table of long-term trends in annualized capacity factors for 610 dams and modeled annualized capacity factors for 362 modeled dams (naturalized and assimilated flows). • dam_annualized_CF_gen.csv: Annualized time series of the following variables for each of 610 hydropower dams with nameplate > 5MW – Reported nameplate capacity (MW) – Implied maximum annual generation (MWh) – Reported net generation (MWh) – Computed annual capacity factor – Modeled annual capacity factor (362 modeled plants only)

13 HYDRO ENERGY↗

An atomic cluster expansion potential for twisted multilayer graphene

Twisted multilayer graphene, characterized by its moiré patterns arising from inter-layer rotational misalignment, serves as a rich platform for exploring quantum phenomena. Machine learning interatomic potentials (MLIPs) are a promising approach to model such systems. Our work develops a method to generate training and test datasets for fitting MLIPs that capture all possible misalignments but remain small-scale to facilitate efficient data generation and parameter estimation. To achieve this, we generate configurations with periodic boundary conditions suitable for density functional theory calculations, and then introduce an internal twist and shift within those supercell structures. Using this technique, supplemented with an active learning workflow, we fit an Atomic Cluster Expansion potential for simulating twisted multilayer graphene and test it for accuracy and robustness on a range of simulation tasks.

2D materials↗

Streaming Compression of Scientific Data via Weak-SINDy

Here, in this paper, a streaming weak-SINDy algorithm is developed specifically for compressing streaming scientific data. The production of scientific data, either via simulation or experiments, is undergoing a stage of exponential growth, which makes data compression important and often necessary for storing and utilizing large scientific data sets. As opposed to classical “offline” compression algorithms that perform compression on a readily available data set, streaming compression algorithms compress data “online” while the data generated from simulation or experiments is still flowing through the system. This feature makes streaming compression algorithms well suited for scientific data compression, where storing the full data set offline is often infeasible. This work proposes a new streaming compression algorithm, streaming weak-SINDy, which takes advantage of the underlying data characteristics during compression. The streaming weak-SINDy algorithm constructs feature matrices and target vectors in the online stage via a streaming integration method in a memory efficient manner. The feature matrices and target vectors are then used in the offline stage to build a model through a regression process that aims to recover equations that govern the evolution of the data. For compressing high-dimensional streaming data, we adopt a streaming proper orthogonal decomposition (POD) process to reduce the data dimension and then use the streaming weak-SINDy algorithm to compress the temporal data of the POD expansion. We propose modifications to the streaming weak-SINDy algorithm to accommodate the dynamically updated POD basis. By combining the built model from the streaming weak-SINDy algorithm and a small amount of data samples, the full data flow could be reconstructed accurately at a low memory cost, as shown in the numerical tests.

97 MATHEMATICS AND COMPUTING↗

Multi-Scale 3D Imaging for Machine Learning Property Upscaling: Mt. Simon Sandstone Case Study

Petrographic properties of principal target reservoirs for carbon sequestration, such as the Mt. Simon Sandstone, are relevant to broad interest groups. The Mt. Simon Sandstone is a deep, saline, regionally extensive Cambrian sandstone, overlain by low permeability sealing formations, making it one of the viable geologic carbon storage reservoirs in the Midwestern US. Its thickness (exceeding 2400 ft in some localities), depth, and lateral extent, combined with high porosity and permeability make it a high-priority target of multiple ongoing geologic carbon sequestration efforts in the United States of America. The National Energy Technology Laboratory in Morgantown, West Virginia, has been engaged in characterization efforts of the Mt. Simon for over a decade, with a strong focus on Computed Tomographic data acquisition. Data generated during this period has been hitherto not accessible to the public. This archival effort focused on preservation of historical CT data and associated metadata, and facilitating their accessibility, culminating with the publication of the entire dataset on NETL’s Energy Data eXchange (EDX) and the associated Gill et. al (2024) paper.

Gill, Magdalena K.↗

Unbinned extraction of $γ$ from $B\to DK$ with normalizing flows

We introduce an unbinned method for extracting the CKM angle $γ$ from the decay chain $B^\pm \to (D \to K_S π^+ π^-) K^\pm$ using normalizing flows (NFs). The NFs, trained on $D$ decay data, learn a faithful continuous representation of the amplitude and strong phase variation over the $D\to K_Sπ^+π^-$ Dalitz plot whose fidelity improves with increased data sample sizes. With this input, the $B$ decay data can be used to extract the parameters $r_B$, $δ_B$, and $γ$. We test the method on Monte Carlo generated data, where it successfully recovers the injected value of $γ$ within uncertainties. The present implementation propagates statistical uncertainties from finite training data via an ensemble of independently trained flows, and does not attempt to capture the effects of systematic experimental errors. We explore two versions of the method that differ in how the trigonometric constraint on phase variation is encoded, and comment on the possible extension to Bayesian NFs, which would provide direct uncertainty estimates on the learned densities without requiring ensemble training.

Grossman, Yuval [Cornell U., LEPP]↗

Intelligent experiments through real-time AI: Fast Data Processing and Autonomous Detector Control for sPHENIX and future EIC detectors (Phase-I)

With an ever increasing demand for high precision data from modern detectors for discovery science and precision measurements, all major high energy nuclear and particle experiments, current and future, are facing the challenge on how to deal with the large volume of raw data generated from sophisticated state-of-the-art detectors in high rate collisions. These goals need to be balanced with available hardware and cost limits on DAQ (Data AcQuisition system) bandwidth and offline computing resources to capture, store and process the signal events. Two prototypical examples are the upcoming sPHENIX experiment, the DOE next generation heavy ion physics experiment at the Relativistic Heavy Ion Collider at BNL, and the future EIC experiments that are planned to be online circa 2030.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

A Comparison of Three Neodymium Atomic Data Sets for Kilonova Modeling

We examine the impact of input neodymium (Nd) atomic data on the light curves and spectra of kilonovae (KNe), probing the sensitivity of kilonova observables to the atomic physics of this important lanthanide element. We use the SuperNu Monte Carlo radiative transfer code, simulating a simple semianalytic 1D kilonova (KN) with a pure Nd atmosphere, fixing the radiative transfer method while using input atomic data generated by three different codes: the LANL suite of atomic physics codes, HULLAC, and Autostructure. We see that the choice of atomic data significantly shapes the resulting light curves and spectra. Peak bolometric luminosities differ by a ratio of nearly 1.5 between HULLAC/Autostructure and LANL data sets. Moreover, we observe significant near- to mid-IR differences in the structure of the spectra. We specifically attribute these differences to the choice of atomic data for neutral Nd I. Many of the results here have been adapted from a presentation at “Radiative Transfer and Atomic Physics of Kilonovae” in Stockholm, 2023. We additionally present a LANL data set with energies calibrated to available values in the NIST Atomic Spectra Database, and demonstrate that this calibration also significantly affects IR spectral structure at late time. The substantial differences in KN observables that arise from tuning the atomic data of just one lanthanide element highlight the special attention that must be paid to atomic physics uncertainties when modeling KNe, from AT2017gfo to beyond.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Mic-hackathon 2024: hackathon on machine learning for electron and scanning probe microscopy

Microscopy is one of the primary sources of information on materials structure and functionality at the nanometer and atomic scales. The data generated through microscopy is often contained in well-structured datasets, enriched with extensive metadata and sample histories, although not always with the same level of detail or storage format. The broad incorporation of data management plans by major funding agencies ensures the preservation and accessibility of this data. However, deriving insights from these rich datasets remains challenging due to the lack of established code ecosystems, standardized benchmarks, and integration strategies. Correspondingly, the efficiency of data usage is very low, and time expenditures at the analysis stage are enormous. In addition to post-acquisition data analysis, the emergence of application programming interfaces by major microscope manufacturers now creates opportunities for real-time ML-based data analytics to enable automated decision making, and particularly ML-agent controlled real-time microscope operation. Despite these opportunities, there is a significant gap in integrating the ML community with the broader microscopy community, limiting the value that these methods bring to physics and materials discovery and materials optimization. Hackathons address these challenges by fostering collaboration between ML experts and microscopy professionals, encouraging the development of innovative solutions that leverage ML for microscopy and preparing the workforce of the future both for microscopy-intensive domains areas, instrument manufacturers, and ML scientists interested in real world applications for fundamental research, materials optimization, and manufacturing. The hackathon generated benchmark datasets and digital twins of microscopes that further contribute to the development of the field and establish data analysis ecosystems. All the codes can be found at GitHub(https://github.com/KalininGroup/Mic-hackathon-2024-codes-publication/tree/1.0.0.1) and Zenodo (https://zenodo.org/records/15579940).

97 MATHEMATICS AND COMPUTING↗

Binding energy of the 𝑇 𝑏⁢𝑏 tetraquark from lattice QCD with relativistic and nonrelativistic heavy-quark actions

We present a new determination of the $b\bar{b}$𝑢⁢𝑑 (𝐽 𝑃 = 1 + , 𝐼 = 0) tetraquark binding energy using lattice quantum chromodynamics (QCD) with domain-wall light quarks and a nonperturbatively tuned three-parameter anisotropic-clover “relativistic” action for the 𝑏 quarks. We also perform a direct comparison with a reanalysis of data generated in prior work using a lattice-nonrelativistic QCD (NRQCD) action for the 𝑏 quarks and otherwise identical parameters. Using the new data with relativistic 𝑏 quarks from seven different ensembles with multiple lattice spacings and pion masses, we perform combined chiral and continuum extrapolations and obtain (𝑚 𝑇 𝑏⁢𝑏 −𝑚 𝐵 −𝑚 𝐵* ) RHQ =(−76 ±23) MeV. For the NRQCD data from five ensembles, we perform chiral-only extrapolations and obtain (𝑚 𝑇 𝑏⁢𝑏 −𝑚 𝐵 −𝑚 𝐵* ) NRQCD = (−74 ±17 ±10) MeV. The lower magnitude of the results obtained here, compared to the original analysis in [Phys. Rev. D 100, 014503 (2019)], is due to the use of the symmetric parts of the correlation matrices with local four-quark operators only.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗