Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “large data sets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Stable Methane Isotopologues From Northern Lakes Suggest That Ebullition Is Dominated by Sub‐Lake Scale Processes

Abstract Stable isotopes have emerged as popular study targets when investigating emission of methane (CH 4 ) from lakes. Yet little is known on how isotopic patterns conform to variations in emission magnitudes—a highly relevant question. Here, we present a large multiyear data set on stable isotopes of CH 4 ebullition (bubbling) from three small adjacent subarctic lakes. The δ 13 C‐CH 4 and δD‐CH 4 range from −78.4‰ to −53.1‰ and from −369.8‰ to −218.8‰, respectively, and vary greatly among the lakes. The signatures suggest dominant hydrogenotrophic methanogenesis, particularly in the deep zones, but there are also signals of seemingly acetoclastic production in some high fluxing shallow areas, possibly fueled by in situ vegetation, but in‐sediment anaerobic CH 4 oxidation cannot be ruled out as an alternative cause. The observed patterns, however, are not consistent across the lakes. Neither do they correspond to the spatiotemporal variations in the measured bubble CH 4 fluxes. Patterns of acetoclastic and hydrogenotrophic production plus oxidation demonstrate that gains and losses of sediment CH 4 are dominated by sub‐lake scale processes. The δD‐CH 4 in the bubbles was significantly different depending on measurement month, likely due to evaporation effects. On a larger scale, our isotopic data, combined with those from other lakes, show a significant difference in bubble δD‐CH 4 between postglacial and thermokarst lakes, an important result for emission inventories. Although this characteristic theoretically assists in source partitioning studies, most hypothetical future shifts in δD‐CH 4 due to high‐latitude lake area or production pathway are too small to lead to atmospheric changes detectable with current technology.

54 ENVIRONMENTAL SCIENCES↗

A Statistical Framework for Evaluating Rain Microphysics in Model Simulations and Disdrometer Observations

Abstract Statistical analyses of a large disdrometer data set and a diverse set of model simulations for convection using the Regional Atmospheric Modeling System were conducted, with the mutual goal of providing insights into precipitation formation and microphysical processes. We demonstrate that a two‐moment bulk microphysical model successfully captures the dominant observed modes of variability in rainfall related to rainfall intensity and raindrop size distributions. The model reproduced the general distribution of observed precipitation groups (PGs) derived from Principal Component Analysis. The multi‐variable analysis also uncovered some shortcomings in the model as well as limitations of the disdrometer data. The model solutions were constrained in their predicted drop size distributions (DSDs) due to the fixed DSD parameters assumed in a two‐moment microphysics scheme. A case study from the Mid‐latitude Continental Clouds and Convection Experiment field project demonstrated how model results can be used to contextualize the disdrometer observations which are limited in sample size, spatial coherence, and detection of small drops and low drop concentrations. The case study also showed that the spatial patterns of the statistically derived PGs revealed by the model are consistent with the hypothesized microphysical processes that determine surface rain DSDs. This work demonstrates how leveraging the strengths of observations and models together can improve our understanding and representation of rain microphysical processes.

54 ENVIRONMENTAL SCIENCES↗

Inferring Plant Acclimation and Improving Model Generalizability With Differentiable Physics‐Informed Machine Learning of Photosynthesis

Net photosynthesis (A N ) is a key component of the global carbon cycle influencing climate feedback over decadal scales. Although plant acclimation to environmental changes can modify A N , traditional vegetation models in Earth system models (ESMs) often rely on plant functional type (PFT)-specific parameterizations or simplified acclimation assumptions limiting generalizability across time, space, and PFTs. In this study, we developed a differentiable photosynthesis model to learn the environmental dependencies of V c,max25 (maximum carboxylation rate at 25°C, representing photosynthetic capacity), as this genre of hybrid physics-informed machine learning can seamlessly train neural networks and process-based equations together. Compared to PFT-specific parameterization of V c,max25 , learning the environment dependencies of key photosynthetic parameters improved model spatiotemporal generalizability. Applying environmental acclimation to V c,max25 led to substantial variations in global mean A N indicating the need to address acclimation in ESMs. The model effectively captured multivariate observations (V c,max25 , A N , and stomatal conductance (g s )) simultaneously with multivariate constraints, improving generalization across space and PFTs. It also learned sensible acclimation relationships of V c,max25 to different environmental conditions. The model explained more than 54%, 57%, and 62% of the variance of A N , g s , and V c,max25 , respectively, presenting a first global-scale spatial test benchmark of A N and g s . These results highlight the potential for differentiable modeling to enhance process-based modules in ESMs and effectively leverage information from large, multivariate data sets.

54 ENVIRONMENTAL SCIENCES↗

Discovering the Multisectoral Impacts of Global Energy Sector Outcomes Through Multiple Ensemble Aggregation Measures

Understanding complex human-Earth system interactions often involves analyzing large scenario ensembles that encompass a wide range of plausible futures. These ensembles often require aggregation to summarize information based on specific criteria or conditions. However, previous research using global change scenario ensembles has largely overlooked how the choice of aggregation method influences the interpretation of results. To address this gap, we leverage a large ensemble data set designed to capture broad energy system dynamics generated using the Global Change Analysis Model. We first explore how energy-related uncertainties are propagated to both global and regional water-energy-food sectors. We then conduct a rank correlation analysis across seven ensemble aggregation measures and demonstrate the need to consider multiple measures in global change scenarios. Our results suggest that global water and food sector outcomes in the 21st century vary widely depending on different scenario assumptions. The global energy productivity is projected to improve by the end of the century across all scenarios. Moreover, regions facing water scarcity challenges in 2100 do not always overlap with those facing extreme energy and food sector outcomes. Although rank correlations across seven aggregation measures are relatively stable across sectors, we identify cases where relying on a single measure leads to losing critical information in the full ensemble. Reliance on a single aggregation measure can distort the interpretation of global change scenario outcomes. Instead, adopting multiple ensemble aggregation measures provides a more holistic understanding of global change scenario ensembles.

Kim, Gijoo↗

A Greening Future Elevates Flash Drought Risk in Northern Mid‐to‐High Latitudes

Flash droughts have become a growing concern, as they can emerge rapidly and increase the risk of crop failure. Although past studies have investigated the meteorological drivers and future changes of flash drought, why flash drought is more frequent over humid and vegetated regions remains underexplored. This study delves further into the mechanism by which vegetation regulates flash drought and its future change using observations from multiple data sets and large ensemble simulations from three Earth system models. On an interannual timescale, both observations and simulations show robust increases in flash drought frequency and a higher flash-to-sub-seasonal drought ratio during spring or antecedent conditions with dense vegetation, supporting the important role of vegetation in flash drought occurrence, especially in the northern mid-to-high latitudes. In the latter regions, the large ensemble simulations show robust increases in flash drought (e.g., 67% and 46% increases in Eastern U.S. and North Asia in 2050–2100 relative to 1950–2000 under the high emission scenario), where the growing season is lengthening. Although greening might suggest reduced drought stress, it drives precipitation-soil moisture-evapotranspiration decoupling by increasing evapotranspiration partitioning to transpiration. As transpiration can access deep soil water through the plant root system, its increased portion can weaken the constraints of concurrent precipitation on evapotranspiration, thus accelerating soil moisture depletion under high evaporative demand, driving a slow-to-rapid drought transition. How vegetation regulates flash drought by regulating surface moisture budget is supported by observations and simulations. Although warming supports early planting, agriculture may increasingly be threatened by surging flash drought risk.

Drought↗

A Mountain Glacier Perspective on the Bipolar Seesaw

A global record of mountain glacier terminations during the last deglaciation (∼19–11 ka) dated by a large, uncurated data set of cosmogenic-nuclide exposure ages highlights a statistically significant asynchrony in termination ages between the Northern and Southern Hemispheres. This interhemispheric offset in the timing of glacier terminations is consistent with previously correlated ice core records that show a systematic interhemispheric lag in the timing of abrupt climate events, with the Southern Hemisphere leading the Northern Hemisphere by ∼300–3,500 years. Our analysis (a) aggregates cosmogenic-nuclide exposure ages from a global data set of moraines to discern climatically driven peaks in moraine emplacement events, and (b) utilizes a Monte Carlo simulation based on a null hypothesis that moraine emplacement is interhemispherically synchronous to estimate the statistical significance of the observed offset. The observed lag of Northern Hemisphere emplacement events compared to the Southern Hemisphere is statistically significant and is consistent with the “bipolar seesaw” pattern observed in ice core records.

58 GEOSCIENCES↗

Predicting fault slip via transfer learning

Abstract Data-driven machine-learning for predicting instantaneous and future fault-slip in laboratory experiments has recently progressed markedly, primarily due to large training data sets. In Earth however, earthquake interevent times range from 10’s-100’s of years and geophysical data typically exist for only a portion of an earthquake cycle. Sparse data presents a serious challenge to training machine learning models for predicting fault slip in Earth. Here we describe a transfer learning approach using numerical simulations to train a convolutional encoder-decoder that predicts fault-slip behavior in laboratory experiments. The model learns a mapping between acoustic emission and fault friction histories from numerical simulations, and generalizes to produce accurate predictions of laboratory fault friction. Notably, the predictions improve by further training the model latent space using only a portion of data from a single laboratory earthquake-cycle. The transfer learning results elucidate the potential of using models trained on numerical simulations and fine-tuned with small geophysical data sets for potential applications to faults in Earth.

58 GEOSCIENCES↗

A measurement of stellar surface gravity hidden in radial velocity differences of comoving stars

The gravitational redshift induced by stellar surface gravity is notoriously difficult to measure for non-degenerate stars, since its amplitude is small in comparison with the typical Doppler shift induced by stellar radial velocity. In this study, we make use of the large observational data set of the Gaia mission to achieve a significant reduction of noise caused by these random stellar motions. By measuring the differences in velocities between the components of the pairs of comoving stars and wide binaries, in this work we are able to statistically measure the combined effects of gravitational redshift and convective blueshifting of spectral lines, and nullify the effect of the peculiar motions of the stars. For the subset of stars considered in this study, we find a positive correlation between the observed differences in Gaia radial velocities and the differences in surface gravity and convective blueshift inferred from effective temperature and luminosity measurements. The results rule out a null signal at the 5σ level for our full data set. Additionally, we study the subdominant effects of binary motion, and possible systematic errors in radial velocity measurements within Gaia. Results from the technique presented in this study are expected to improve significantly with data from the next Gaia data release. Such improvements could be used to constrain the mass–luminosity relation and stellar models that predict the magnitude of convective blueshift.

(stars:) binaries: general↗

The intrinsic alignment of red galaxies in DES Y1 redMaPPer galaxy clusters

ABSTRACT Clusters of galaxies trace the most non-linear peaks in the cosmic density field. The weak gravitational lensing of background galaxies by clusters can allow us to infer their masses. However, galaxies associated with the local environment of the cluster can also be intrinsically aligned due to the local tidal gradient, contaminating any cosmology derived from the lensing signal. We measure this intrinsic alignment in Dark Energy Survey (DES) Year 1 redMaPPer clusters. We find evidence of a non-zero mean radial alignment of galaxies within clusters between redshifts 0.1–0.7. We find a significant systematic in the measured ellipticities of cluster satellite galaxies that we attribute to the central galaxy flux and other intracluster light. We attempt to correct this signal, and fit a simple model for intrinsic alignment amplitude (AIA) to the measurement, finding AIA = 0.15 ± 0.04, when excluding data near the edge of the cluster. We find a significantly stronger alignment of the central galaxy with the cluster dark matter halo at low redshift and with higher richness and central galaxy absolute magnitude (proxies for cluster mass). This is an important demonstration of the ability of large photometric data sets like DES to provide direct constraints on the intrinsic alignment of galaxies within clusters. These measurements can inform improvements to small-scale modelling and simulation of the intrinsic alignment of galaxies to help improve the separation of the intrinsic alignment signal in weak lensing studies.

79 ASTRONOMY AND ASTROPHYSICS↗

Marginal unbiased score expansion and application to CMB lensing

Here, we present the marginal unbiased score expansion (MUSE) method, an algorithm for generic high-dimensional hierarchical Bayesian inference. MUSE performs approximate marginalization over arbitrary non-Gaussian latent parameter spaces, yielding Gaussianized asymptotically unbiased and near-optimal constraints on global parameters of interest. It is computationally much cheaper than exact alternatives like Hamiltonian Monte Carlo (HMC), excelling on funnel problems which challenge HMC, and does not require any problem-specific user supervision like other approximate methods such as variational inference or many simulation-based inference methods. MUSE makes possible the first joint Bayesian estimation of the delensed Cosmic Microwave Background (CMB) power spectrum and gravitational lensing potential power spectrum, demonstrated here on a simulated data set as large as the upcoming South Pole Telescope 3G 1500 deg 2 survey, corresponding to a latent dimensionality of ~6 million and of order 100 global bandpower parameters. On a subset of the problem where an exact but more expensive HMC solution is feasible, we verify that MUSE yields nearly optimal results. We also demonstrate that existing spectrum-based forecasting tools which ignore pixel-masking underestimate predicted error bars by only ~10%. This method is a promising path forward for fast lensing and delensing analyses which will be necessary for future CMB experiments such as SPT-3G, Simons Observatory, or CMB-S4, and can complement or supersede existing HMC approaches. The success of MUSE on this challenging problem strengthens its case as a generic procedure for a broad class of high-dimensional inference problems.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Characterizing skyrmion flow phases with principal component analysis

Principal component analysis (PCA) is a powerful method that can identify patterns in large, complex data sets by constructing low-dimensional order parameters from higher-dimensional feature vectors. There are increasing efforts to use space-and-time-dependent PCA to detect transitions in nonequilibrium systems that are difficult to characterize with equilibrium methods. Here, we demonstrate that feature vectors incorporating the position and velocity information of driven skyrmions moving through random disorder permit PCA to resolve different types of disordered skyrmion motion as a function of driving force and the ratio of the Magnus force to the dissipation. Since the Magnus force creates gyroscopic motion and a finite Hall angle, skyrmions can exhibit a greater range of flow phases than what is observed in overdamped driven systems with quenched disorder. We show that in addition to identifying previously known skyrmion flow phases, PCA detects several additional phases, including different types of channel flow, moving fluids, and partially ordered states. Guided by the PCA analysis, we further characterize the disordered flow phases to elucidate the different microscopic dynamics and show that the changes in the PCA-derived order parameters can be connected to features in bulk transport measures, including the transverse and longitudinal velocity-force curves, differential conductivity, topological defect density, and changes in the skyrmion Hall angle as a function of drive. We discuss how asymmetric feature vectors can be used to improve the resolution of the PCA analysis, and how this technique can be extended to find disordered phases in other nonequilibrium systems with time-dependent dynamics.

36 MATERIALS SCIENCE↗

GPU-based Image Compression for Efficient Compositing in Distributed Rendering Applications

Visualizations of large-scale data sets are often created on graphics clusters that distribute the rendering task amongst many processes. When using real-time GPU-based graphics algorithms, the most time-consuming aspect of distributed rendering is typically the com-positing phase - combining all partial images from each rendering process into the final visualization. Compo siting requires image data to be copied off the GPU and sent over a network to other processes. While compression has been utilized in existing distributed rendering compositors to reduce the data being sent over the network, this compression tends to occur after the raw images are transferred from the GPU to main memory. In this paper, we present work that leverages OpenGL / CUDA interoperability to compress raw images on the GPU prior to transferring the data to main memory. This approach can significantly reduce the device-to-host data transfer time, thus enabling more efficient compositing of images generated by distributed rendering applications.

Lipinksi, Riley↗

Training data selection for event classification in a highly variable environment

A problem of interest for nuclear nonproliferation is monitoring activities at nuclear facilities, where proliferation events may only take place a few times and often under variable conditions. Machine learning has revolutionized data analytics by enabling the use of measurable signatures to generate predictive models of facility operations. However, traditional methods for training these models require large, reliable data sets with labeled observations, a challenge for nonproliferation. Highly variable conditions further complicate this as events from training data may have occurred in conditions quite different from the event of interest. Our hypothesis is that when events occur in a highly variable environment, careful training data selection for each test event could outperform the standard approach of using all available training data. We developed a method to optimize training data selection for the given test event and applied it to predicting the power level of the High Flux Isotope Reactor (HFIR) at Oak Ridge National Laboratory. In this study, the reactor startup exhibits variability between occurrences due to natural variability in environmental conditions and operational procedures. Using a combination of analysis techniques, a similitude assessment was performed on data collected from HFIR to isolate clusters that were optimal for training a predictive model. Concepts such as dynamic time warping and Jaccard similarity were used in conjunction with clustering analysis. In order to validate this approach, the model was trained on every combination of unique training events and the predictive performance was compared to the performance using a subset of the training data selected by isolated clusters found through the similitude assessment.

Iyer, A↗

High-throughput genetics enables identification of nutrient utilization and accessory energy metabolism genes in a model methanogen

Archaea are widespread in the environment and play fundamental roles in diverse ecosystems; however, characterization of their unique biology requires advanced tools. This is particularly challenging when characterizing gene function. Here, we generate randomly barcoded transposon libraries in the model methanogenic archaeon Methanococcus maripaludis and use high-throughput growth methods to conduct fitness assays (RB-TnSeq) across over 100 unique growth conditions. Using our approach, we identified new genes involved in nutrient utilization and response to oxidative stress. We identified novel genes for the usage of diverse nitrogen sources in M. maripaludis including a putative regulator of alanine deamination and molybdate transporters important for nitrogen fixation. Furthermore, leveraging the fitness data, we inferred that M. maripaludis can utilize additional nitrogen sources including $\tiny{L}$-glutamine, $\tiny{D}$-glucuronamide, and adenosine. Under autotrophic growth conditions, we identified a gene encoding a domain of unknown function (DUF166) that is important for fitness and hypothesize that it has an accessory role in carbon dioxide assimilation. Finally, comparing fitness costs of oxygen versus sulfite stress, we identified a previously uncharacterized class of dissimilatory sulfite reductase-like proteins (Dsr-LP; group IIId) that is important during growth in the presence of sulfite. When overexpressed, Dsr-LP conferred sulfite resistance and enabled use of sulfite as the sole sulfur source. The high-throughput approach employed here allowed for generation of a large-scale data set that can be used as a resource to further understand gene function and metabolism in the archaeal domain.

59 BASIC BIOLOGICAL SCIENCES↗

Regulation of bacterial stringent response by an evolutionarily conserved ribosomal protein L11 methylation

Lysine and arginine methylation is an important regulator of enzyme activity and transcription in eukaryotes. However, little is known about this covalent modification in bacteria. In this work, we investigated the role of methylation in bacteria. By reanalyzing a large phyloproteomics data set from 48 bacterial strains representing six phyla, we found that almost a quarter of the bacterial proteome is methylated. Many of these methylated proteins are conserved across diverse bacterial lineages, including those involved in central carbon metabolism and translation. Among the proteins with the most conserved methylation sites is ribosomal protein L11 (bL11). bL11 methylation has been a mystery for five decades, as the deletion of its methyltransferase PrmA causes no cell growth defects. Comparative proteomics analysis combined with inorganic polyphosphate and guanosine tetra/pentaphosphate assays of the ΔprmA mutant in Escherichia coli revealed that bL11 methylation is important for stringent response signaling. In the stationary phase, we found that the ΔprmA mutant has impaired guanosine tetra/pentaphosphate production. This leads to a reduction in inorganic polyphosphate levels, accumulation of RNA and ribosomal proteins, and an abnormal polysome profile. Overall, our investigation demonstrates that the evolutionarily conserved bL11 methylation is important for stringent response signaling and ribosomal activity regulation and turnover.

59 BASIC BIOLOGICAL SCIENCES↗

Streaming Compression of Scientific Data via Weak-SINDy

Here, in this paper, a streaming weak-SINDy algorithm is developed specifically for compressing streaming scientific data. The production of scientific data, either via simulation or experiments, is undergoing a stage of exponential growth, which makes data compression important and often necessary for storing and utilizing large scientific data sets. As opposed to classical “offline” compression algorithms that perform compression on a readily available data set, streaming compression algorithms compress data “online” while the data generated from simulation or experiments is still flowing through the system. This feature makes streaming compression algorithms well suited for scientific data compression, where storing the full data set offline is often infeasible. This work proposes a new streaming compression algorithm, streaming weak-SINDy, which takes advantage of the underlying data characteristics during compression. The streaming weak-SINDy algorithm constructs feature matrices and target vectors in the online stage via a streaming integration method in a memory efficient manner. The feature matrices and target vectors are then used in the offline stage to build a model through a regression process that aims to recover equations that govern the evolution of the data. For compressing high-dimensional streaming data, we adopt a streaming proper orthogonal decomposition (POD) process to reduce the data dimension and then use the streaming weak-SINDy algorithm to compress the temporal data of the POD expansion. We propose modifications to the streaming weak-SINDy algorithm to accommodate the dynamically updated POD basis. By combining the built model from the streaming weak-SINDy algorithm and a small amount of data samples, the full data flow could be reconstructed accurately at a low memory cost, as shown in the numerical tests.

97 MATHEMATICS AND COMPUTING↗

Cluster Expansion Analysis of Atomic Order in Li-Ion Battery Cathode Material LiCo y Ni 1-y O 2

A modified cluster-expansion treatment was developed recently to analyze atomic order in LiCo y Ni 1-y O 2 , a model cathode material for Li-ion batteries. In this treatment, referred to as a “spin-atom” cluster expansion, the occupant of a lattice site is identified by its spin state as well as its atomic species. Further, Effective Cluster Interaction (ECI) coefficients are derived from a large training data set (i.e., the set of atomic arrangements for which DFT calculations are performed) which is filtered by an anomaly detection algorithm to eliminate poorly converged DFT calculations. The cluster expansion incorporates Li-Ni (LN) exchange as well as intralayer Co-Ni (CN) exchange. Monte Carlo simulations based on the cluster expansion were applied to the Ni-rich part of the phase diagram. The simulations predict a miscibility gap between y = 0.05 and y = 0.65.

25 ENERGY STORAGE↗

DSS-SimPy-RL (Open-DSS and SimPy based Cyber-Physical RL environment) [SWR-23-29]

Recently, numerous data-driven approaches to control an electric grid using machine learning techniques have been investigated. With the advancement of reinforcement learning (RL) based techniques, gradually the conventional optimization based solvers are being replaced with RL approach where there is uncertainty in the environment such as renewable generation or cyber system emulation. However, to train an agent efficiently, it requires numerous interactions with an environment to learn the best policies. There are numerous RL environments for the power systems based on some well-known simulators, similarly there are environment for communication domains. While majority of the cyber emulators are based in an UNIX environment, the power simulators are based in the Windows-based operating system, the generation of cyber-physical mixed domain RL environment has been challenging. Existing co-simulation methods are efficient but resource and time intensive to generate large scale data set for training RL agents. Hence, this software focuses on development and validation of a mixed domain RL environment using Open DSS for the physical side and leverages a discrete event simulator python package, SimPy, for cyber-side emulation which is Operating Systems agnostic. Further utilizing this software co-simulation and training RL agents for re-routing based resilient control for network reconfiguration and volt-var control in power distribution feeder are performed.

Sahu, Abhijeet↗