Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data dependencies”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

The Timing of Potential Last Nucleosynthetic Injections into the Protosolar Molecular Cloud Inferred from 41 Ca– 26 Al Systematics of Bulk CAIs

Short-lived radionuclides (SLRs) provide important information about the chronology of the early solar system. Among them, 41 Ca, due to its decay to 41 K with a half-life of only 0.1 Ma, is particularly valuable in constraining the timescales and origins of both SLRs and the formation of the oldest solar system materials, the Ca–Al-rich inclusions (CAIs). The initial abundance of 41Ca in the solar system, expressed as the ( 41 Ca/ 40 Ca)I ratio, is the key to unveiling the origin of this nuclide. Here, we report a new solar system ( 41 Ca/ 40 Ca)I ratio of 2.0 × 10 −8 derived from the K isotope compositions of two CAIs. This new ratio is about four times higher than the previous value inferred from a mineral isochron. Such a high ( 41 Ca/ 40 Ca)I ratio in the CAIs exceeds that expected for the protosolar molecular cloud by ∼1000×, implying very late injection of the 41 Ca (and possibly other SLRs) into the protosolar molecular cloud. The correlated enrichments of 41 Ca and 26 Al in the bulk CAI samples hint at a common stellar origin of both SLRs. The injection time estimated from our new data depends on the stellar source—it ranges from 0.6 Ma for a Wolf–Rayet wind to 1.0 Ma for a TP-AGB star ejecta.

79 ASTRONOMY AND ASTROPHYSICS↗

MLOps for Beam Controls

Machine learning operations (MLOps) is the standardization and streamlining of the ML development lifecycle to address the challenges associated with large-scale machine learning applications. The full MLOps pipeline consists of open-source tools: DataHub, MinIO and MLflow. It is being used for dataset management and model development to handle changing data dependencies, varying business needs, reproducibility, and diverse teams working with differing tools and skills. To demonstrate the completion of an MLOps pipeline for particle accelerator operations, we are deploying a simple script that computes settings for the Booster’s gradient magnet power supply. Once the demonstration is complete, we will develop and deploy ML-based optimization algorithms to improve Booster’s overall efficiency. This MLOps pipeline opens the gate to systematically develop and deploy ML applications for accelerator controls and diagnostics.

43 PARTICLE ACCELERATORS↗

MLOps for Beam Controls

Machine learning operations (MLOps) is the standardization and streamlining of the ML development lifecycle to address the challenges associated with large-scale machine learning applications. The full MLOps pipeline consists of open-source tools: DataHub, MinIO and MLflow. It is being used for dataset management and model development to handle changing data dependencies, varying business needs, reproducibility, and diverse teams working with differing tools and skills. To demonstrate the completion of an MLOps pipeline for particle accelerator operations, we are deploying a simple script that computes settings for the Booster’s gradient magnet power supply. Once the demonstration is complete, we will develop and deploy ML-based optimization algorithms to improve Booster’s overall efficiency. This MLOps pipeline opens the gate to systematically develop and deploy ML applications for accelerator controls and diagnostics.

43 PARTICLE ACCELERATORS↗

LSAFE: a Lightweight Static Analysis Framework for binary Executables

Static analysis is a widely used technique for analyzing various aspects of programs. However, as programs become more complex, static analysis tools require larger resources, such as CPU time and memory, to perform the same tasks. Moreover, the source code of programs may not always be accessible, requiring static analysis to be performed on the binary executable code directly. To overcome these challenges, we propose a lightweight static analysis framework called LSAFE, which constructs control flow graphs (CFGs) and data dependency graphs (DDGs) of target programs with optimized performance in terms of CPU and memory usage. We evaluated the proposed framework using both Spec benchmark programs and real-world industrial applications, and found that it outperformed Angr, an existing state-of-the-art static analysis tool. Additionally, we demonstrate a case study that utilizes the CFG generated by LSAFE to detect memory leaks.

Qu, Guangzhi↗

Powering Data Centers with Clean Energy: A Techno-Economic Case Study of Nuclear and Renewable Energy Dependability

Rising data demands from artificial intelligence (AI) and large language models (LLMs) generating images, videos, and text have prompted increased need for larger and more robust data centers in the United States. Major companies interested in these larger data centers face the choice of linking them to existing regional grids, building stand-alone power supplies onsite, or a combination of both. The request, review, and approval process for new transmission lines to grids in the United States, however, has grown in recent years to times spans rivaling those of new construction for nuclear power plants. Building an islanded power supply for each data center is therefore becoming a prominent option. In this case study, several technologies are modeled in techno-economic simulations for long-term system costs subject to fixed electricity demand from a singular data center. A 250 MWe data center is assumed with additional 50 MWe for resiliency. Techno-economic simulations are conducted using the Holistic Energy Resource Optimization Network (HERON) software, which is a part of the Framework for Optimization of Resources and Economics (FORCE) tool suite. Technologies considered include solar, wind, lithium-ion batteries, and several types of nuclear reactors: large-scale reactors, small modular reactors, and microreactors. A low- and high-cost estimate for each technology is assumed to develop a range of expected economic performance. Low-cost estimates included several clean energy production tax credits. Different combinations of renewable energy generators with nuclear reactors are considered, ranging from a fully renewable-powered data center to a fully nuclear-powered data center. Historic time series of wind and solar availability from the Texas grid are used to train a reduced order model; this model then generates unique time series with similar characteristics of the training dataset. Multiple scenarios of weather and subsequent operations are simulated for each renewable-nuclear combination to determine total costs throughout the project lifetime. Fully renewable-powered configurations required large amounts of installed capacity (GW scale) in the simulations to meet the fixed demand of the data center. This is due to some scenarios in the historical dataset which captured low-wind and low-solar days, requiring over-building of these technologies as well as batteries to compensate for the low amounts of electricity generation. Fully nuclear-powered configurations outperformed the fully renewable and mixed renewable-nuclear configurations in terms of cost, with ranges between $1B and $10B in 2023 USDs compared to $40B+ for fully renewable configurations. Of the nuclear technologies, small modular reactors performed better economically than large-scale nuclear models due to lower projected capital costs, and both performed better than the microreactor models. These results demonstrate the applicability of firm, dispatchable electricity resources from baseload generators like nuclear power plants for operating facilities that run at constant power without daily variability.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Temperature-Dependent Constrained Diffusion of Micro-Confined Alkylimidazolium Chloride Ionic Liquids

Alkylimidazolium chloride ionic liquids (ILs) have many uses in a variety of separation systems, including micro-confined separation systems. To understand the separation mechanism in these systems, the diffusion properties of analytes in ILs under relevant operating conditions, including micro-confinement dimension and temperature, should be known. For example, separation efficiencies for various IL-based microextraction techniques are dependent on the sample volume and temperature. Temperature-dependent (20–100 °C) fluorescence recovery after photobleaching (FRAP) was utilized to determine the diffusion properties of a zwitterionic, hydrophilic dye, ATTO 647, in alkylimidazolium chloride ILs in micro-confined geometries. These micro-confined geometries were generated by sandwiching the IL between glass substrates that were separated by ~1 to 100 μm. From the measured temperature-dependent FRAP data, we note alkyl chain length-, thickness-, and temperature-dependent diffusion coefficients, with values ranging from 0.021 to 46 μm 2 /s. Deviations from Brownian diffusion are observed at lower temperatures and increasingly less so at elevated temperatures; the differences are attributed to alterations in intermolecular interactions that reduce temperature-dependent nanoscale structural heterogeneities. Furthermore, the temperature- and thickness-dependent data provide a useful foundation for efficient design of micro-confined IL separation systems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Covariate Dependent Sparse Functional Data Analysis

This study proposes a method to incorporate covariate information into sparse functional data analysis. The method aims at cases where each subject has a limited number of longitudinal measurements and is associated with static covariates. This research is motivated by several use cases in practice. One representative example is void swelling, a nuclear-specific material degradation mechanism. Void swelling is affected by many covariates, including alloy composition and irradiation type. How to accurately model the complicated joint effects of such covariates on the swelling process is the key to mitigating the effect of swelling and ensuring safe operation. Unlike most of the existing methods, the proposed method can handle high-dimensional covariates with the informative covariate identification procedure and sparse and irregularly spaced measurements, that is, does not require complete or dense observations. The main innovation of the proposed method is that we model the variation coming from covariates and the variation left conditioned on covariates, such that the functional principal component analysis and Gaussian process can be conducted in a unified manner. Further, we also propose a systematic approach to identify important covariates in the hypothesis testing context. The methodology is demonstrated on applications in nuclear engineering and healthcare and simulation studies.

42 ENGINEERING↗

Generation of Enrichment-Dependent Thermal Neutron Scattering Data

This work details the generation of enrichment-dependent thermal neutron scattering cross sections for several crucial uranium fuel compounds. The evaluations of the thermal scattering law (TSL) and associated cross sections for uranium dioxide (UO 2 ), uranium carbide (UC), and uranium nitride (UN) were performed using standard ab initio lattice dynamics (AILD) methods. The data for uranium metal was produced using a novel hybrid approach of molecular dynamics combined with lattice dynamics methods. 235 U enrichments of 5%, 10% (LEU+), 19.75% (HALEU), 93% (HEU), and 100% were considered, in addition to natural uranium. The enrichment-dependent masses and free atom cross sections were used in the generation of elastic and inelastic thermal neutron scattering cross sections, while the calculation of the phonon density of states (DOS) and resulting TSL considered only the natural isotopic composition of uranium. The use of an identical DOS for all enrichments is expected to have minimal impact on the final data, as the small change in uranium mass should not significantly affect lattice vibrations. The cross sections are shown to exhibit significant dependence on 235 U enrichment. The submission of this data to the National Nuclear Data Center (NNDC) for release in the ENDF/B-VIII.1 database should support the design of advanced reactor concepts.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Regression of exchangeable relational arrays

Relational arrays represent measures of association between pairs of actors, often in varied contexts or over time. Trade flows between countries, financial transactions between individuals, contact frequencies between school children in classrooms and dynamic protein-protein interactions are all examples of relational arrays. Elements of a relational array are often modelled as a linear function of observable covariates. Uncertainty estimates for regression coefficient estimators, and ideally the coefficient estimators themselves, must account for dependence between elements of the array, e.g., relations involving the same actor. Existing estimators of standard errors that recognize such relational dependence rely on estimating extremely complex, heterogeneous structure across actors. This paper develops a new class of parsimonious coefficient and standard error estimators for regressions of relational arrays. Here we leverage an exchangeability assumption to derive standard error estimators that pool information across actors, and are substantially more accurate than existing estimators in a variety of settings. This exchangeability assumption is pervasive in network and array models in the statistics literature, but not previously considered when adjusting for dependence in a regression setting with relational data. We demonstrate improvements in inference theoretically, via a simulation study, and by analysis of a dataset involving international trade.

97 MATHEMATICS AND COMPUTING↗

Niche-DE: niche-differential gene expression analysis in spatial transcriptomics data identifies context-dependent cell-cell interactions

Existing methods for analysis of spatial transcriptomic data focus on delineating the global gene expression variations of cell types across the tissue, rather than local gene expression changes driven by cell-cell interactions. We propose a new statistical procedure called niche-differential expression (niche-DE) analysis that identifies cell-type-specific niche-associated genes, which are differentially expressed within a specific cell type in the context of specific spatial niches. We further develop niche-LR, a method to reveal ligand-receptor signaling mechanisms that underlie niche-differential gene expression patterns. Niche-DE and niche-LR are applicable to low-resolution spot-based spatial transcriptomics data and data that is single-cell or subcellular in resolution.

59 BASIC BIOLOGICAL SCIENCES↗

Verification of Upcoming MCNP Features For Estimating Nuclear Data Sensitivities in Fixed Source Simulations [Abstract]

Predictive simulation codes, like the Monte Carlo N-Particle (MCNP) transport code, are used throughout the nuclear community. These simulations are based on nuclear data. Maximizing the accuracy and precision of nuclear data maximizes the accuracy and precision of the overall simulation. This is imperative to applications that rely on simulations. For example, improving nuclear data for special nuclear material improves simulation accuracy in stockpile stewardship applications, which results in larger safety margins and decreased operational costs. The improvement and validation of nuclear data is completed through integral benchmark experiments. Past benchmarks have primarily been limited to focus on the effective multiplication factor ($\kappa$ eff ); broadening the purview of benchmarks beyond $\kappa$ eff -dependent nuclear data addresses nuclear data deficiencies. Different response types depend on different areas of nuclear data. This dependence is quantified as nuclear data sensitivity: the change in response due to perturbation of a contributing parameter. The larger the nuclear data sensitivity of a response, the more the experiment is influenced by the uncertainties of the nuclear data. The optimization of nuclear data sensitivities in future benchmarks would result in more detailed validation of lesser studied areas of nuclear data. Currently, direct sensitivity capabilities are not easily found for all experiment types and parameters. An MCNP tool to directly estimate the cross section sensitivities of tallied values is under development. Additionally, updates have been made to the perturbation feature of MCNP, which can be used in a less direct approach to estimating sensitivities. This work verifies these features to estimate nuclear data sensitivities in fixed source simulations of a 4.5-kg sphere of alpha- phase weapons-grade plutonium surrounded by differing amounts of copper and polyethylene. Integrated estimates made using MCNP’s tools were found to statistically agree with integrated estimates made from manual perturbation of nuclear data proving the validity of the MCNP tools.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

What's New in the NSRDB

The National Solar Radiation Database (NSRDB) provides solar resource data across the globe at a high temporal and spatial resolution. This data is primarily used in solar energy modeling. The NSRDB is updated annually for the United States and North, Central and South America and the data is currently available from 1998-2021. In 2022 the NSRDB was updated using the latest version of the underlying Physical Solar Model (PSM). This update includes improved surface albedo and gap-filling of cloud properties. The inclusion of these updates reduced the uncertainty in the data compared to previous versions of the NSRDB. The Himawari and Meteosat Indian Ocean Data Coverage (IODC) satellites were added to the Geostationary Operational Environmental Satellite (GOES) and made our coverage global. While standard data from the GOES continues to be served at an hourly 4km x 4km resolution, full resolution data has also been made available to the user. The NSRDB now contains over 200Tb of data with nearly 40Tb being added annually. We provide significant flexibility for data download depending on the amount of data required by the users. In this paper we provide an update on the current status on the NSRDB.

photovoltaic systems↗

ESS-DIVE guidelines for archiving terrestrial model data

This dataset contains supporting documents and images for ESS-DIVE terrestrial model data archiving guidelines.Terrestrial models are broadly defined as numerical models that couple both land dynamics and energy, water, carbon, or nutrient fluxes. We created these guidelines based on input from the U.S. Department of Energy’s Biological and Environmental Research land modeling community. The guidelines are intended to help modelers determine which components of their terrestrial model data associated with publication should be archived. Based on input from the land modeling community, the guidelines recommend archiving both model input and testing data, as well as code, script, and metadata. The guidelines also recommend archiving model data output, depending on the limitations set by data repositories. Lastly, we provide recommendations for bundling data files for publication as well as a discussion about tools that can facilitate model data archiving and reuse.This dataset is an archive of the associated GitHub repository for our model archiving guidelines (https://github.com/ess-dive-community/essdive-model-data-archiving-guidelines). The ‘README.pdf’ file gives a general introduction to the guidelines, and the ‘instructions.pdf’ file provides more detailed steps for following the guidelines. We also provide 2 figures in this data package: 1) a decision tree (model_data_guidelines_decision_tree.png) that can help users determine which components of their model data to archive. and 2) the ‘model_data_guidelines_flmd.png’ file depicts the different files that can be archived in addition to the model data itself. Lastly, we include 3 digitized tables from our associated manuscript and 3 CSV files with anonymized input from DOE scientists about the importance of different aspects of model data archiving from which we developed the guidelines.Dataset updates for v1.1.0: We updated this data package on 2021-11-22 in response to review comments on our related manuscript. In this update we removed one figure so that the model archiving guidelines are conveyed in text rather than an image. We updated the file-level metadata (FLMD) figure to be in accord with the most recent FLMD recommendations. We made minor edits to the README file to update the recommended citation and added two co-authors. We also added 6 new data files (3 are anonymized input from DOE scientists that helped to inform guidelines, and 3 are digitized tables from our manuscript.

54 ENVIRONMENTAL SCIENCES↗

Monte Carlo Global QCD Analyses of the Pion Parton Distribution Functions

As the lightest hadron, the pion presents itself as a dichotomy. While being the pseudo Goldstone boson associated with chiral symmetry breaking, it is simultaneously regarded as the lightest pseudoscalar meson typically composed of constituent quark and antiquark. As the chiral symmetry shares the same global symmetry as QCD, studying the pion and its internal structure is key in understanding one of the four fundamental forces of nature, namely the quantum chromodynamic strong interaction. We study the pion's parton distribution functions (PDFs), which are universal quantities that describe the structure of the pion in terms of its constituent quarks, antiquarks, and gluons. Our goal is to discover what the available data reveal for these universal quantities, namely the PDFs. Through the use of factorization theorems and perturbative QCD, we use Monte Carlo (MC) methods to extract the PDFs from the available Drell-Yan (DY) and leading neutron (LN) data. While the DY process involves two hadrons colliding, the LN electroproduction involves an electron beam incident on a target nucleon. We can constrain well the pion PDFs at large momentum fraction using DY data, and at low momentum fraction using LN data. We also use threshold resummation in the DY process to gather predictable higher order terms associated with the soft gluon radiation because their contributions to the cross section are nontrivial. In this dissertation, we parametrize the pion PDFs and fit those parameters to the available data. We make use of Bayesian inference and use MC techniques to sample the parameter space. Three sets of results are presented. We first extract the pion PDFs by fitting the DY and LN data. Then, we include transverse momentum dependent DY data in conjunction with the DY and LN data to extract the pion PDFs. Finally, we apply various methods of threshold resummation to the DY cross section and extract pion PDFs. The pion PDFs presented here are the first global QCD analyses performed as well as the first MC extracted pion PDFs, which are at the forefront of the pion PDF community.

Barry, Patrick↗

Energy Dependence in the Neutrino Scattering Data of MINERvA

As we prepare for lower neutrino energy beams where physics is highly energy-dependent, it is important to isolate factors that can contribute largely to these low energy neutrino experiments. The world of neutrino oscillation experiments uses a wide range of energies: from 0.7 GeV beam of T2K and MicroBooNE, 2 GeV beam of NOvA, to the 0.6 to 6.0 GeV beam produced by DUNE in the future. MINERvA’s data are currently the best place to test the upper end of the range for DUNE and can be extrapolated into both DUNE’s and NOvA’s oscillation maximum. The goal of the research is to constrain energy dependence in the neutrino scattering experiment using data from MINERvA and then determine what factors contribute to the observed energy dependence. The research is divided into two parts. First, I will analyze different theoretical models in different channels of interaction, such as the quasi-elastic and delta resonance channels, in terms of neutrino energy dependence. The goal is to study the models at the level of the structure functions, which emphasizes the W2 structure function (the C function for QE) dominates the MINERvA data, but the W3 is increasingly important for NOvA and DUNE. The second half of the research is data-driven and focuses on the different detectors and systematic effects. An experimental effect, the angle acceptance, is larger than the structure function effects. However, it is a detector geometry effect and is well measured and well modeled. Other effects such as the muon energy scale are small and localized. Additionally, the research will determine if the MINERvA GENIE model correctly predicts all of the energy dependence seen in the data, and identify the remaining unmodeled energy dependence between the MINERvA data and its best simulation.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

From multivariate to functional data analysis: Fundamentals, recent developments, and emerging areas

Functional data analysis (FDA), which is a branch of statistics on modeling infinite dimensional random vectors resided in functional spaces, has become a major research area for Journal of Multivariate Analysis. We review some fundamental concepts of FDA, their origins and connections from multivariate analysis, and some of its recent developments, including multi-level functional data analysis, high-dimensional functional regression, and dependent functional data analysis. Here, we also discuss the impact of these new methodology developments on genetics, plant science, wearable device data analysis, image data analysis, and business analytics. Two real data examples are provided to motivate our discussions.

97 MATHEMATICS AND COMPUTING↗

Exploring Ion Mobility Mass Spectrometry Data File Conversions to Leverage Existing Tools and Enable New Workflows

Ion mobility (IM) is often combined with LC-MS experiments to provide an additional dimension of separation for complex sample analysis. While highly complex samples are better characterized by the full dimensionality of LC-IM-MS experiments to uncover new information, downstream data analysis workflows are often not equipped to properly mine the additional IM dimension. For many samples the data acquisition benefits of including IM separations are all that is necessary to uncover sample information and the full dimensionality of the data is not required for data analysis. Post-acquisition reduction and adaptation of the dimensions of LC-IM-MS and IM-MS experiments into an LC-MS format opens the possibility to use a plethora of existing software tools. In this work, we developed data file conversion tools to reduce the complexity of IM data analysis. Three data file transformations are introduced in the PNNL PreProcessor software: 1) mapping the IM axis to the LC axis for IM-MS data, 2) converting the drift time vs. m/z space to CCS/z vs m/z space, and 3) transforming All Ions IM/MS mobility aligned fragmentation data to a standard LC-MS DDA data file format. Finally, these new data file conversions are demonstrated with corresponding lipidomics and proteomics workflows that leverage existing LC-MS data analysis software to highlight the benefits of the data transformations.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

How different power plant types contribute to electric grid reliability, resilience, and vulnerability: a comparative analytical framework

Abstract This work explores the dependability tradeoffs provided by the most common types of central power plants in the United States. Historically, the electricity sector has lacked consensus on how reliability , resilience , and vulnerability differ and how those metrics change depending on the power plant fleet composition. We propose distinct definitions for these metrics and an analytical framework to evaluate power plant fleet dependability. Using data analysis and literature review, we identify fifteen dependability attributes across which we rank eleven power plant types relative to natural gas combined-cycle (NGCC) plants. We use NGCC as the benchmark because it is common to many locations and is of relatively recent vintage. The framework shows that each power plant type has unique dependability benefits and drawbacks. We provide examples of how researchers may use the framework to evaluate grid dependability qualitatively under different scenarios. We find that assuming all attributes that contribute to grid dependability are equally important and additive, electric grid dependability is best supported when power plant fleets include a mixture of power generation technologies. Then, we discuss scenario characteristics that could alter the prioritization and relationships of attributes. We also find that if current capacity installation trends continue to favor low- and zero-carbon power plants, US power grids may benefit from increased resilience and reduced vulnerability at the cost of decreased reliability. We conclude by recommending methods for adapting the framework and quantifying relationships between attributes in individual scenarios.

Ramirez-Meyers, K. (ORCID:0000000291216952)↗