Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “large data sets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Machine learning inversion from small-angle scattering for charged polymers

We develop Monte Carlo simulations for uniformly charged polymers and a machine learning algorithm to interpret the intra-polymer structure factor of the charged polymer system, which can be obtained from small-angle scattering experiments. The polymer is modeled as a chain of fixed-length bonds, where the connected bonds are subject to bending energy, and there is also a screened Coulomb potential for charge interaction between all joints. The bending energy is determined by the intrinsic bending stiffness, and the charge interaction depends on the interaction strength and screening length. All three contribute to the stiffness of the polymer chain and lead to longer and larger polymer conformations. The screening length also introduces a second length scale for the polymer besides the bending persistence length. To obtain the inverse mapping from the structure factor to these polymer conformation and energy-related parameters, we generate a large data set of structure factors by running simulations for a wide range of polymer energy parameters. We use principal component analysis to investigate the intra-polymer structure factors and determine the feasibility of the inversion using the nearest neighbor distance. We employ Gaussian process regression to achieve the inverse mapping and extract the characteristic parameters of polymers from the structure factor with low relative error.

36 MATERIALS SCIENCE↗

Estimates of the wavenumber wavelet power spectrum of magnetic fluctuations during magnetic reconnection

Fluctuation analyses of experimental observations generally lack high temporal resolution and are in frequency-space f, contrary to theoretical efforts in wavenumber-space k. This is due to the inherent limits of the Fourier transform, though it is prominent due to the ease of diagnostic implementation. Advances in wavelet-based analysis have provided relief due to its temporal resolution, but in its common use, is still hard to compare to theoretical models. By using the two-point correlation technique in conjunction with large data sets, a wavelet power spectrum in wavenumber-space can be created. Dubbed the wavenumber wavelet power spectrum, this spectrum relates wavenumber to power in time. Further, this analysis technique more closely connects characterizations of experimentally observed fluctuations with other system parameters and theoretical predictions. In this article, we develop the wavenumber wavelet power spectrum using magnetic fluctuations caused by tearing instability driven magnetic reconnection in reproducible, high temperature laboratory plasmas. These dynamic magnetic fluctuations generated in reversed field pinch plasmas are broadband, ranging from the low frequency, 10's of kHz, up to the ion gyroradii frequencies, 100's of kHz. The dominant fluctuations have poloidal and toroidal mode numbers (m,n)=(1,6−10) and can grow to 2%–3% of the mean magnetic field. During these reconnection events, ions, and electrons are energized, magnetic fluctuation amplitudes increase, plasma flow is halted, and the toroidal magnetic flux increases, all on a semi-periodic basis. The newly developed spectrum provides better temporal resolution of spectrum characteristics to correlate with these particle energization phenomena.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Identification of low-momentum muons in the CMS detector using multivariate techniques in proton-proton collisions at $\sqrt{s}$ = 13.6 TeV

“Soft” muons with a transverse momentum below 10 GeV are featured in many processes studied by the CMS experiment, such as decays of heavy-flavor hadrons or rare tau lepton decays. Maximizing the selection efficiency for these muons, while simultaneously suppressing backgrounds from long-lived light-flavor hadron decays, is therefore important for the success of the CMS physics program. Multivariate techniques have been shown to deliver better muon identification performance than traditional selection techniques. To take full advantage of the large data set currently being collected during Run 3 of the CERN LHC, a new multivariate classifier based on a gradient-boosted decision tree has been developed. It offers a significantly improved separation of signal and background muons compared to a similar classifier used for the analysis of the Run 2 data. The performance of the new classifier is evaluated on a data set collected with the CMS detector in 2022 and 2023, corresponding to an integrated luminosity of 62 fb -1 .

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Genetic and Functional Diversity Help Explain Pathogenic, Weakly Pathogenic, and Commensal Lifestyles in the Genus Xanthomonas

The genus Xanthomonas has been primarily studied for pathogenic interactions with plants. However, besides host and tissue-specific pathogenic strains, this genus also comprises nonpathogenic strains isolated from a broad range of hosts, sometimes in association with pathogenic strains, and other environments, including rainwater. Based on their incapacity or limited capacity to cause symptoms on the host of isolation, nonpathogenic xanthomonads can be further characterized as commensal and weakly pathogenic. This study aimed to understand the diversity and evolution of nonpathogenic xanthomonads compared to their pathogenic counterparts based on their cooccurrence and phylogenetic relationship and to identify genomic traits that form the basis of a life history framework that groups xanthomonads by ecological strategies. We sequenced genomes of 83 strains spanning the genus phylogeny and identified eight novel species, indicating unexplored diversity. While some nonpathogenic species have experienced a recent loss of a type III secretion system, specifically the hrp2 cluster, we observed an apparent lack of association of the hrp2 cluster with lifestyles of diverse species. We performed association analysis on a large data set of 337 Xanthomonas strains to explain how xanthomonads may have established association with the plants across the continuum of lifestyles from commensals to weak pathogens to pathogens. Presence of distinct transcriptional regulators, distinct nutrient utilization and assimilation genes, transcriptional regulators, and chemotaxis genes may explain lifestyle-specific adaptations of xanthomonads.

59 BASIC BIOLOGICAL SCIENCES↗

Moment tensor event identification for collapses

SUMMARY We introduce a seismic identification method for collapse events using moment tensors (MTs). We start by computing full (six-element) MT solutions for 43 identified collapse events from around the world, and statistically characterizing the population on the MT hypersphere. We then test a large data set of over 1000 full MTs for the western U.S. against the distribution of collapses using a MT-based identification method similarly as used for testing explosions. Known collapses and explosions are readily identified, along with other anomalous events in the Geysers and central California coast. Misidentification rates are determined for various screening angles with optimal misidentification rates between earthquakes and collapses on the order of 3 per cent. The method is demonstrated to be very effective at identifying non-earthquake sources with a 97–98 per cent accuracy. It is likely to be transportable to other regions, and can be used for event identification anywhere full MT solutions are routinely calculated.

58 GEOSCIENCES↗

A unique, ring-like radio source with quadrilateral structure detected with machine learning

ABSTRACT We report the discovery of a unique object in the MeerKAT Galaxy Cluster Legacy Survey (MGCLS) using the machine learning anomaly detection framework astronomaly. This strange, ring-like source is 30′ from the MGCLS field centred on Abell 209, and is not readily explained by simple physical models. With an assumed host galaxy at redshift 0.55, the luminosity (1025 W Hz−1) is comparable to powerful radio galaxies. The source consists of a ring of emission 175 kpc across, quadrilateral enhanced brightness regions bearing resemblance to radio jets, two ‘ears’ separated by 368 kpc, and a diffuse envelope. All of the structures appear spectrally steep, ranging from −1.0 to −1.5. The ring has high polarization (25 per cent) except on the bright patches (<10 per cent). We compare this source to the Odd Radio Circles recently discovered in ASKAP data and discuss several possible physical models, including a termination shock from starburst activity, an end-on radio galaxy, and a supermassive black hole merger event. No simple model can easily explain the observed structure of the source. This work, as well as other recent discoveries, demonstrates the power of unsupervised machine learning in mining large data sets for scientifically interesting sources.

Astronomy & Astrophysics↗

Machine Learning in Infectious Disease for Risk Factor Identification and Hypothesis Generation: Proof of Concept Using Invasive Candidiasis

Machine learning (ML) models can handle large data sets without assuming underlying relationships and can be useful for evaluating disease characteristics, yet they are more commonly used for predicting individual disease risk than for identifying factors at the population level. We offer a proof of concept applying random forest (RF) algorithms to Candida-positive hospital encounters in an electronic health record database of patients in the United States. Candida-positive encounters were extracted from the Cerner HealthFacts database; invasive infections were laboratory-positive sterile site Candida infections. Features included demographics, admission source, care setting, physician specialty, diagnostic and procedure codes, and medications received before the first positive Candida culture. We used RF to assess risk factors for 3 outcomes: any invasive candidiasis (IC) vs non-IC, within-species IC vs non-IC (eg, invasive C. glabrata vs noninvasive C. glabrata), and between-species IC (eg, invasive C. glabrata vs all other IC). Fourteen of 169 (8%) variables were consistently identified as important features in the ML models. When evaluating within-species IC, for example, invasive C. glabrata vs non-invasive C. glabrata, we identified known features like central venous catheters, intensive care unit stay, and gastrointestinal operations. In contrast, important variables for invasive C. glabrata vs all other IC included renal disease and medications like diabetes therapeutics, cholesterol medications, and antiarrhythmics. Known and novel risk factors for IC were identified using ML, demonstrating the hypothesis-generating utility of this approach for infectious disease conditions about which less is known, specifically at the species level or for rarer diseases.

60 APPLIED LIFE SCIENCES↗

Detailed study of quark-hadron duality in spin structure functions of the proton and neutron

Background: The response of hadrons, the bound states of the strong force (QCD), to external probes can be described in two different, complementary frameworks: as direct interactions with their fundamental constituents, quarks and gluons, or alternatively as elastic or inelastic coherent scattering that leaves the hadrons in their ground state or in one of their excited (resonance) states. The former picture emerges most clearly in hard processes with high momentum transfer, where the hadron response can be described by the perturbative expansion of QCD, while at lower energy and momentum transfers, the resonant excitations of the hadrons dominate the cross section. The overlap region between these two pictures, where both yield similar predictions, is referred to as quark-hadron duality and has been extensively studied in reactions involving unpolarized hadrons. Some limited information on this phenomenon also exists for polarized protons, deuterons, and 3 He nuclei, but not yet for neutrons. Purpose: In this paper, we present comprehensive and detailed results on the correspondence between the extrapolated deep inelastic structure function g 1 of both the proton and the neutron with the same quantity measured in the nucleon resonance region. Thanks to the fine binning and high precision of our data, and using a well-controlled perturbative QCD (pQCD) fit for the partonic prediction, we can make quantitative statements about the kinematic range of applicability of both local duality and global duality. Method: We use the most updated QCD global analysis results at high x from the Jefferson Lab Angular Momentum Collaboration to extrapolate the spin structure function g 1 into the nucleon resonance region and then integrate over various intervals in the scaling variable x. We compare the results with the large data set collected in the quark-hadron transition region by the CLAS Collaboration, including, for the first time, deconvoluted neutron data, integrated over the same intervals. We present this comparison as a function of the momentum transfer Q 2 . Results: We find that, depending on the integration interval and the minimum momentum transfer chosen, a clear transition to quark-hadron duality can be observed in both nucleon species. Furthermore, we show, for the first time, the approach to scaling behavior for g 1 measured in the resonance region at sufficiently high momentum transfer. Conclusions: Here, our results can be used to quantify the deviations from the applicability of pQCD for data taken at moderate energies and can help with extraction of quark distribution functions from such data.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Machine-learning-informed scattering correlation analysis of sheared colloids

We have carried out theoretical analysis, Monte Carlo simulations and machine-learning analysis to quantify microscopic rearrangements of dilute dispersions of spherical colloidal particles from coherent scattering intensity. Both monodisperse and polydisperse dispersions of colloids were created and underwent a rearrangement consisting of an affine simple shear and non-affine rearrangement using the Monte Carlo method. We calculated the coherent scattering intensity of the dispersions and the correlation function of intensity before and after the rearrangement and generated a large data set of angular correlation functions for varying system parameters, including number density, polydispersity, shear strain and non-affine rearrangement. Singular value decomposition of the data set shows the feasibility of machine-learning inversion from the correlation function for the polydispersity, shear strain and non-affine rearrangement using only three parameters. A Gaussian process regressor is then trained on the data set and can retrieve the affine shear strain, non-affine rearrangement and polydispersity with relative errors of 3%, 1% and 6%, respectively. Altogether, our model provides a framework for quantitative studies of both steady and non-steady microscopic dynamics of colloidal dispersions using coherent scattering methods.

Gaussian process regression↗

Fast Gaussian Process Estimation for Large-Scale In Situ Inference using Convolutional Neural Networks

Exascale computing will bring with it significant I/O limitations. One foreseeable consequence of such restrictions is that the user can save only a small fraction of complex simulation data to disk for subsequent analysis. An alternative is to fit statistical models to data in situ, that is, inside the simulation as it runs. This option requires extremely fast statistical estimation to avoid slowing down the simulation. Gaussian processes (GPs) have state-of-the-art predictive performance for modeling spatial data. However, standard estimation methods for GPs scale quite poorly to large data sets as parameter estimation requires inverting a covariance matrix to the size of the data set. In the presented work, we use a convolutional neural network (CNN) to predict the GP parameters for a spatial data set, from a simulation or otherwise, rather than optimize the parameters directly. Here, our presented case study models spatial data from E3SM, the Department of Energy’s Exascale climate model. The CNN is trained on synthetic data simulated from GP models with known parameters and then applied to data from the climate simulation. In the presented examples, the neural network scheme produces parameter estimates that compare well with standard methods such as maximum likelihood estimation in predictive performance but is obtained four orders of magnitude faster.

big data↗

On the Use of Smart Meter Data to Estimate the Voltage Magnitude on the Primary Side of Distribution Service Transformers

This paper develops a novel method to estimate the voltage magnitude on the primary side of distribution service transformers. The proposed method relies exclusively on smart meters, and therefore it is fully data-driven. This is an important feature because electric utilities have detailed models of only the primary network - that is, the network between the distribution substation and the primary side of service transformers that are installed closer to end-customer sites. The network that connects the secondary side of service transformers to end-customer sites, referred to as the secondary network, is simply represented by a lumped load. For each secondary network, the proposed method uses data acquired from only 2 smart meters: the closest and the farthest-in the sense of electrical distance - from the service transformer. As a reference to this feature, the proposed method is named SM2Vp. To our knowledge, this is the first time a method is shown to provide actionable information for realtime operation and control of power distribution grids using only two smart meters per secondary network. This is important because utilities have experienced barriers in managing and using large data sets for real-time operation and control. SM2Vp is primarily intended to provide pseudo-measurements for distribution system state estimation, but it can also be used directly for voltage control schemes. The performance of SM2Vp is demonstrated by numerical simulations carried out on three secondary network synthetic models and by using field data provided by a utility partner serving customers in southwestern California. A maximum relative error of approximately 3.9% or less is observed for the primary voltage magnitude estimates in all numerical experiments.

distribution service transformer↗

Characterization of DER Momentary Cessation and Rate-of-Change-of-Frequency Response

Momentary cessation (MC) and response to rate-of-change-of-frequency (ROCOF) are inverter responses that can have serious impact on the stability of the grid during abnormal conditions. Though IEEE Std 1547-2018 provides fairly well defined expected responses from inverter-based DERs during abnormal grid conditions, a significant portion of currently installed DERs in the distribution network do not have a well defined response to abnormal grid conditions. Any analysis of the grid involving inverter MC and ROCOF response must consider the specific characteristics of these responses for the installed DERs to have better confidence in the analysis results, especially for grids with high levels of DERs. However, there is lack of information on MC and ROCOF response for the inverters already installed in the field. In this paper, we examine a large data set of installed inverters to map the dominant population of installed inverters. This information is then used to characterize the most dominant MC and ROCOF responses of installed inverters using experimental data.

DER↗

Heat Exchanger Design Optimization for Mitigation of Diesel Engine Exhaust Fouling

Abstract The design optimization of a diesel exhaust-coupled heat and mass exchanger that drives a 2.71 kW cooling capacity absorption heat pump is presented in this study. Fouling layer thermal resistance and pressure drops from single-tube experiments are used to develop a thermodynamic, heat transfer, and pressure drop model for the exhaust-coupled desorber. A parametric study is performed to select a desorber design that meets system performance while minimizing footprint. Experimental heat duties and pressure drops are within 10% and 3%, respectively, of the model predictions. Thus, large data sets from single-tube experiments with representative geometries are successful in accounting for fouling effects at the component level. Desorber design optimization based on this approach ensures continued heat pump performance after fouling. This study, along with the single-tube experiments, presents a systematic approach to design exhaust-coupled heat exchangers while considering the effects of fouling. These results are applicable for a wide range of waste-heat recovery applications, and this method can be extended to different geometries and operating conditions.

Energy & Fuels↗

Metagenomic clustering links specific metabolic functions to globally relevant ecosystems

ABSTRACT Metagenomic sequencing has advanced our understanding of biogeochemical processes by providing an unprecedented view into the microbial composition of different ecosystems. While the amount of metagenomic data has grown rapidly, simple-to-use methods to analyze and compare across studies have lagged behind. Thus, tools expressing the metabolic traits of a community are needed to broaden the utility of existing data. Gene abundance profiles are a relatively low-dimensional embedding of a metagenome’s functional potential and are, thus, tractable for comparison across many samples. Here, we compare the abundance of KEGG Ortholog Groups (KOs) from 6,539 metagenomes from the Joint Genome Institute’s Integrated Microbial Genomes and Metagenomes (JGI IMG/M) database. We find that samples cluster into terrestrial, aquatic, and anaerobic ecosystems with marker KOs reflecting adaptations to these environments. For instance, functional clusters were differentiated by the metabolism of antibiotics, photosynthesis, methanogenesis, and surprisingly GC content. Using this functional gene approach, we reveal the broad-scale patterns shaping microbial communities and demonstrate the utility of ortholog abundance profiles for representing a rapidly expanding body of metagenomic data. IMPORTANCE Metagenomics, or the sequencing of DNA from complex microbiomes, provides a view into the microbial composition of different environments. Metagenome databases were created to compile sequencing data across studies, but it remains challenging to compare and gain insight from these large data sets. Consequently, there is a need to develop accessible approaches to extract knowledge across metagenomes. The abundance of different orthologs (i.e., genes that perform a similar function across species) provides a simplified representation of a metagenome’s metabolic potential that can easily be compared with others. In this study, we cluster the ortholog abundance profiles of thousands of metagenomes from diverse environments and uncover the traits that distinguish them. This work provides a simple to use framework for functional comparison and advances our understanding of how the environment shapes microbial communities.

54 ENVIRONMENTAL SCIENCES↗

Lithium-Ion Battery Life Model with Electrode Cracking and Early-Life Break-in Processes

This paper develops a physically justified reduced-order capacity fade model from accelerated calendar- and cycle-aging data for 32 lithium-ion (Li-ion) graphite/nickel-manganese-cobalt (NMC) cells. The large data set reveals temperature-, charge C-rate-, depth-of-discharge-, and state of charge (SOC)-dependent degradation patterns that would be unobserved in a smaller test matrix. Model structure is informed by incremental capacity analysis that shows loss of lithium inventory and cathode-material loss as the dominant capacity fade mechanisms. The model includes terms attributable to solid-electrolyte interface (SEI) growth, electrode cracking, cycling-driven acceleration of SEI growth, and "break-in" mechanisms that slightly decrease or increase available Li inventory early in life. The study explores what mathematical couplings of these mechanisms best describe calendar aging, cycle aging, and mixed calendar/cycle aging. Various approaches are discussed for extracting relevant stress factors from complex cycling profiles to predict lifetime during real-world battery loads using models trained on constant-current laboratory test results. The complexity of the present human-driven model identification process motivates future work in machine learning to more widely search and statistically discern the optimal model that correctly extrapolates capacity fade based on physical knowledge.

25 ENERGY STORAGE↗

PS-Samp (Phase-Space Sampling)

PS-Samp is a software that reduces a given dataset by pruning data. Data is downselected such that it uniformly spans the phase-space, thereby ensuring that the full dataset is correctly represented. The algorithm relies on estimating the probability map of the data and using it to construct an acceptance probability. The software is designed to handle very large data sets.

Hassanaly, Malik↗

BLDAP Intro to Python/Data Science Curriculum v1

The Github repository contains the Jupyter notebooks for the intro to Python / Data Science course for Berkeley Lab Director's Apprenticeship Program (BLDAP). This course is designed for students with little to no experience in coding to learn skills in Python necessary for data science. Students utilize Jupyter notebooks throughout the course. The overall goal is for students to learn how to use Python to clean, analyze, and visualize large data sets in order to communicate effectively their conclusions about the data set. Students apply the skills they learned on actual data sets provided by researchers in Berkeley Lab.

Hales, Laurel [Lawrence Berkeley National Laborato↗