Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “large data sets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Machine Learning in Infectious Disease for Risk Factor Identification and Hypothesis Generation: Proof of Concept Using Invasive Candidiasis

Machine learning (ML) models can handle large data sets without assuming underlying relationships and can be useful for evaluating disease characteristics, yet they are more commonly used for predicting individual disease risk than for identifying factors at the population level. We offer a proof of concept applying random forest (RF) algorithms to Candida-positive hospital encounters in an electronic health record database of patients in the United States. Candida-positive encounters were extracted from the Cerner HealthFacts database; invasive infections were laboratory-positive sterile site Candida infections. Features included demographics, admission source, care setting, physician specialty, diagnostic and procedure codes, and medications received before the first positive Candida culture. We used RF to assess risk factors for 3 outcomes: any invasive candidiasis (IC) vs non-IC, within-species IC vs non-IC (eg, invasive C. glabrata vs noninvasive C. glabrata), and between-species IC (eg, invasive C. glabrata vs all other IC). Fourteen of 169 (8%) variables were consistently identified as important features in the ML models. When evaluating within-species IC, for example, invasive C. glabrata vs non-invasive C. glabrata, we identified known features like central venous catheters, intensive care unit stay, and gastrointestinal operations. In contrast, important variables for invasive C. glabrata vs all other IC included renal disease and medications like diabetes therapeutics, cholesterol medications, and antiarrhythmics. Known and novel risk factors for IC were identified using ML, demonstrating the hypothesis-generating utility of this approach for infectious disease conditions about which less is known, specifically at the species level or for rarer diseases.

60 APPLIED LIFE SCIENCES↗

Detailed study of quark-hadron duality in spin structure functions of the proton and neutron

Background: The response of hadrons, the bound states of the strong force (QCD), to external probes can be described in two different, complementary frameworks: as direct interactions with their fundamental constituents, quarks and gluons, or alternatively as elastic or inelastic coherent scattering that leaves the hadrons in their ground state or in one of their excited (resonance) states. The former picture emerges most clearly in hard processes with high momentum transfer, where the hadron response can be described by the perturbative expansion of QCD, while at lower energy and momentum transfers, the resonant excitations of the hadrons dominate the cross section. The overlap region between these two pictures, where both yield similar predictions, is referred to as quark-hadron duality and has been extensively studied in reactions involving unpolarized hadrons. Some limited information on this phenomenon also exists for polarized protons, deuterons, and 3 He nuclei, but not yet for neutrons. Purpose: In this paper, we present comprehensive and detailed results on the correspondence between the extrapolated deep inelastic structure function g 1 of both the proton and the neutron with the same quantity measured in the nucleon resonance region. Thanks to the fine binning and high precision of our data, and using a well-controlled perturbative QCD (pQCD) fit for the partonic prediction, we can make quantitative statements about the kinematic range of applicability of both local duality and global duality. Method: We use the most updated QCD global analysis results at high x from the Jefferson Lab Angular Momentum Collaboration to extrapolate the spin structure function g 1 into the nucleon resonance region and then integrate over various intervals in the scaling variable x. We compare the results with the large data set collected in the quark-hadron transition region by the CLAS Collaboration, including, for the first time, deconvoluted neutron data, integrated over the same intervals. We present this comparison as a function of the momentum transfer Q 2 . Results: We find that, depending on the integration interval and the minimum momentum transfer chosen, a clear transition to quark-hadron duality can be observed in both nucleon species. Furthermore, we show, for the first time, the approach to scaling behavior for g 1 measured in the resonance region at sufficiently high momentum transfer. Conclusions: Here, our results can be used to quantify the deviations from the applicability of pQCD for data taken at moderate energies and can help with extraction of quark distribution functions from such data.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Machine-learning-informed scattering correlation analysis of sheared colloids

We have carried out theoretical analysis, Monte Carlo simulations and machine-learning analysis to quantify microscopic rearrangements of dilute dispersions of spherical colloidal particles from coherent scattering intensity. Both monodisperse and polydisperse dispersions of colloids were created and underwent a rearrangement consisting of an affine simple shear and non-affine rearrangement using the Monte Carlo method. We calculated the coherent scattering intensity of the dispersions and the correlation function of intensity before and after the rearrangement and generated a large data set of angular correlation functions for varying system parameters, including number density, polydispersity, shear strain and non-affine rearrangement. Singular value decomposition of the data set shows the feasibility of machine-learning inversion from the correlation function for the polydispersity, shear strain and non-affine rearrangement using only three parameters. A Gaussian process regressor is then trained on the data set and can retrieve the affine shear strain, non-affine rearrangement and polydispersity with relative errors of 3%, 1% and 6%, respectively. Altogether, our model provides a framework for quantitative studies of both steady and non-steady microscopic dynamics of colloidal dispersions using coherent scattering methods.

Gaussian process regression↗

Fast Gaussian Process Estimation for Large-Scale In Situ Inference using Convolutional Neural Networks

Exascale computing will bring with it significant I/O limitations. One foreseeable consequence of such restrictions is that the user can save only a small fraction of complex simulation data to disk for subsequent analysis. An alternative is to fit statistical models to data in situ, that is, inside the simulation as it runs. This option requires extremely fast statistical estimation to avoid slowing down the simulation. Gaussian processes (GPs) have state-of-the-art predictive performance for modeling spatial data. However, standard estimation methods for GPs scale quite poorly to large data sets as parameter estimation requires inverting a covariance matrix to the size of the data set. In the presented work, we use a convolutional neural network (CNN) to predict the GP parameters for a spatial data set, from a simulation or otherwise, rather than optimize the parameters directly. Here, our presented case study models spatial data from E3SM, the Department of Energy’s Exascale climate model. The CNN is trained on synthetic data simulated from GP models with known parameters and then applied to data from the climate simulation. In the presented examples, the neural network scheme produces parameter estimates that compare well with standard methods such as maximum likelihood estimation in predictive performance but is obtained four orders of magnitude faster.

big data↗

Considerations for AMI-Based Operations for Distribution Feeders

More than $5 billion in investments in advanced metering infrastructure (AMI) technologies, AMI deployments, as pervasive secondary network voltage monitoring systems, provide opportunities for utility operations and controls. This paper focuses on the considerations for AMI-based tools and techniques as the industry moves toward operationalizing such large data sets. Phase identification is a first such tool. Numerous distribution network analysis, monitoring, and control applications - including volt/volt-ampere reactive control, state estimation, and distribution automation - require accurate phase connectivity information in the system models. The phase connectivity database maintained by utilities is inaccurate because of a significant amount of missing data, restoration activities, and network reconfiguration. Existing phase identification techniques that estimate phase connectivity work well in distribution feeders that have low or no photovoltaic (PV) generation; however, they fail to identify the phases accurately when considerable PV generation is present. This work addresses the phase identification problem in the presence of high PV generation using statistical analysis methods. Further, insights into the AMI data requirements for this application in terms of data window length and resolution are provided using sensitivity analysis performed on an actual distribution feeder model of San Diego Gas & Electric Company. The results of this study show that the phase connectivity, even in the presence of high PV generation, can be accurately identified using statistical analysis of AMI data of 1 day.

24 POWER TRANSMISSION AND DISTRIBUTION↗

On the Use of Smart Meter Data to Estimate the Voltage Magnitude on the Primary Side of Distribution Service Transformers

This paper develops a novel method to estimate the voltage magnitude on the primary side of distribution service transformers. The proposed method relies exclusively on smart meters, and therefore it is fully data-driven. This is an important feature because electric utilities have detailed models of only the primary network - that is, the network between the distribution substation and the primary side of service transformers that are installed closer to end-customer sites. The network that connects the secondary side of service transformers to end-customer sites, referred to as the secondary network, is simply represented by a lumped load. For each secondary network, the proposed method uses data acquired from only 2 smart meters: the closest and the farthest-in the sense of electrical distance - from the service transformer. As a reference to this feature, the proposed method is named SM2Vp. To our knowledge, this is the first time a method is shown to provide actionable information for realtime operation and control of power distribution grids using only two smart meters per secondary network. This is important because utilities have experienced barriers in managing and using large data sets for real-time operation and control. SM2Vp is primarily intended to provide pseudo-measurements for distribution system state estimation, but it can also be used directly for voltage control schemes. The performance of SM2Vp is demonstrated by numerical simulations carried out on three secondary network synthetic models and by using field data provided by a utility partner serving customers in southwestern California. A maximum relative error of approximately 3.9% or less is observed for the primary voltage magnitude estimates in all numerical experiments.

distribution service transformer↗

Characterization of DER Momentary Cessation and Rate-of-Change-of-Frequency Response

Momentary cessation (MC) and response to rate-of-change-of-frequency (ROCOF) are inverter responses that can have serious impact on the stability of the grid during abnormal conditions. Though IEEE Std 1547-2018 provides fairly well defined expected responses from inverter-based DERs during abnormal grid conditions, a significant portion of currently installed DERs in the distribution network do not have a well defined response to abnormal grid conditions. Any analysis of the grid involving inverter MC and ROCOF response must consider the specific characteristics of these responses for the installed DERs to have better confidence in the analysis results, especially for grids with high levels of DERs. However, there is lack of information on MC and ROCOF response for the inverters already installed in the field. In this paper, we examine a large data set of installed inverters to map the dominant population of installed inverters. This information is then used to characterize the most dominant MC and ROCOF responses of installed inverters using experimental data.

DER↗

Heat Exchanger Design Optimization for Mitigation of Diesel Engine Exhaust Fouling

Abstract The design optimization of a diesel exhaust-coupled heat and mass exchanger that drives a 2.71 kW cooling capacity absorption heat pump is presented in this study. Fouling layer thermal resistance and pressure drops from single-tube experiments are used to develop a thermodynamic, heat transfer, and pressure drop model for the exhaust-coupled desorber. A parametric study is performed to select a desorber design that meets system performance while minimizing footprint. Experimental heat duties and pressure drops are within 10% and 3%, respectively, of the model predictions. Thus, large data sets from single-tube experiments with representative geometries are successful in accounting for fouling effects at the component level. Desorber design optimization based on this approach ensures continued heat pump performance after fouling. This study, along with the single-tube experiments, presents a systematic approach to design exhaust-coupled heat exchangers while considering the effects of fouling. These results are applicable for a wide range of waste-heat recovery applications, and this method can be extended to different geometries and operating conditions.

Energy & Fuels↗

Metagenomic clustering links specific metabolic functions to globally relevant ecosystems

ABSTRACT Metagenomic sequencing has advanced our understanding of biogeochemical processes by providing an unprecedented view into the microbial composition of different ecosystems. While the amount of metagenomic data has grown rapidly, simple-to-use methods to analyze and compare across studies have lagged behind. Thus, tools expressing the metabolic traits of a community are needed to broaden the utility of existing data. Gene abundance profiles are a relatively low-dimensional embedding of a metagenome’s functional potential and are, thus, tractable for comparison across many samples. Here, we compare the abundance of KEGG Ortholog Groups (KOs) from 6,539 metagenomes from the Joint Genome Institute’s Integrated Microbial Genomes and Metagenomes (JGI IMG/M) database. We find that samples cluster into terrestrial, aquatic, and anaerobic ecosystems with marker KOs reflecting adaptations to these environments. For instance, functional clusters were differentiated by the metabolism of antibiotics, photosynthesis, methanogenesis, and surprisingly GC content. Using this functional gene approach, we reveal the broad-scale patterns shaping microbial communities and demonstrate the utility of ortholog abundance profiles for representing a rapidly expanding body of metagenomic data. IMPORTANCE Metagenomics, or the sequencing of DNA from complex microbiomes, provides a view into the microbial composition of different environments. Metagenome databases were created to compile sequencing data across studies, but it remains challenging to compare and gain insight from these large data sets. Consequently, there is a need to develop accessible approaches to extract knowledge across metagenomes. The abundance of different orthologs (i.e., genes that perform a similar function across species) provides a simplified representation of a metagenome’s metabolic potential that can easily be compared with others. In this study, we cluster the ortholog abundance profiles of thousands of metagenomes from diverse environments and uncover the traits that distinguish them. This work provides a simple to use framework for functional comparison and advances our understanding of how the environment shapes microbial communities.

54 ENVIRONMENTAL SCIENCES↗

Lithium-Ion Battery Life Model with Electrode Cracking and Early-Life Break-in Processes

This paper develops a physically justified reduced-order capacity fade model from accelerated calendar- and cycle-aging data for 32 lithium-ion (Li-ion) graphite/nickel-manganese-cobalt (NMC) cells. The large data set reveals temperature-, charge C-rate-, depth-of-discharge-, and state of charge (SOC)-dependent degradation patterns that would be unobserved in a smaller test matrix. Model structure is informed by incremental capacity analysis that shows loss of lithium inventory and cathode-material loss as the dominant capacity fade mechanisms. The model includes terms attributable to solid-electrolyte interface (SEI) growth, electrode cracking, cycling-driven acceleration of SEI growth, and "break-in" mechanisms that slightly decrease or increase available Li inventory early in life. The study explores what mathematical couplings of these mechanisms best describe calendar aging, cycle aging, and mixed calendar/cycle aging. Various approaches are discussed for extracting relevant stress factors from complex cycling profiles to predict lifetime during real-world battery loads using models trained on constant-current laboratory test results. The complexity of the present human-driven model identification process motivates future work in machine learning to more widely search and statistically discern the optimal model that correctly extrapolates capacity fade based on physical knowledge.

25 ENERGY STORAGE↗

PS-Samp (Phase-Space Sampling)

PS-Samp is a software that reduces a given dataset by pruning data. Data is downselected such that it uniformly spans the phase-space, thereby ensuring that the full dataset is correctly represented. The algorithm relies on estimating the probability map of the data and using it to construct an acceptance probability. The software is designed to handle very large data sets.

Hassanaly, Malik↗

BLDAP Intro to Python/Data Science Curriculum v1

The Github repository contains the Jupyter notebooks for the intro to Python / Data Science course for Berkeley Lab Director's Apprenticeship Program (BLDAP). This course is designed for students with little to no experience in coding to learn skills in Python necessary for data science. Students utilize Jupyter notebooks throughout the course. The overall goal is for students to learn how to use Python to clean, analyze, and visualize large data sets in order to communicate effectively their conclusions about the data set. Students apply the skills they learned on actual data sets provided by researchers in Berkeley Lab.

Hales, Laurel [Lawrence Berkeley National Laborato↗

An open-access simulated earthquake ground-motion database for an M7 Hayward Fault earthquake in the San Francisco Bay Region

Comprehensive understanding of earthquake ground motions, particularly in the near-fault region of large-magnitude events, is limited by gaps in strong-motion data. This challenge is prominent in areas with high seismic hazard but infrequent large earthquakes where data is sparse and difficult to interpret. These data limitations lead to uncertainties in the development of site-specific ground motions, which are crucial for engineering risk assessments. To address these challenges, physics-based regional-scale ground-motion simulations have been developed. With the emergence of exaflop-scale computing ecosystems, it is now possible to simulate regional earthquake processes at unprecedented fidelity and generate the large number of fault rupture realizations necessary to characterize both intra- and inter-event ground-motion variability. This article introduces a new database of simulated earthquake ground motions, created for applications in earthquake engineering, earthquake planning, and emergency response. The inaugural version of the database features simulated ground motions for a magnitude 7 Hayward Fault earthquake in the San Francisco Bay Region (SFBR), using the EarthQuake SIMulation (EQSIM) simulation framework and the Graves–Pitarka kinematic rupture model. The aim is to provide high-fidelity, spatially dense, three-component motions generated on the Department of Energy’s (DOE) newest generation of graphics processing unit (GPU)-accelerated supercomputers. These motions are being made openly available to the engineering, scientific, and disaster planning communities. In addition, this work develops protocols for the efficient dissemination of these large data sets and emphasizes community engagement to build confidence in their application. This article discusses the methodology behind the data, underlying software verification and validation, scalable data management, and a user interface for data access. The goal is to facilitate widespread use and elicit expert feedback to maximize the utility and exploitation of simulated motions. While the initial focus is on the San Francisco Region, simulations for additional regions will be added as the DOE program progresses.

Simulated ground-motion database↗

Automatic Waveform Quality Control for Surface Waves Using Machine Learning

Surface-wave seismograms are widely used by researchers to study Earth’s interior and earthquakes. To extract information reliably and robustly from a suite of surface waveforms, the signals require quality control screening to reduce artifacts from signal complexity and noise. This process has usually been completed by human experts labeling each waveform visually, which is time consuming and tedious for large data sets. We explore automated approaches to improve the efficiency of waveform quality control processing by investigating logistic regression, support vector machines, K-nearest neighbors, random forests (RF), and artificial neural networks (ANN) algorithms. To speed up signal quality assessment, we trained these five machine learning (ML) methods using nearly 400,000 human-labeled waveforms. The ANN and RF models outperformed other algorithms and achieved a test accuracy of 92%. We evaluated these two best-performing models using seismic events from geographic regions not used for training. The results show that the two trained models agree with labels from human analysts but required only 0.4% of the time. Although the original (human) quality assignments assessed general waveform signal-to-noise, the ANN or RF labels can help facilitate detailed waveform analysis. Our investigations demonstrate the capability of the automated processing using these two ML models to reduce outliers in surface-wave-related measurements without human quality control screening.

58 GEOSCIENCES↗

TEACHING AN OLD ACCELERATOR NEW TRICKS

The Argonne Tandem Linac Accelerator System (ATLAS) has been a National User Facility since 1985. In that time, many of the systems that help operators retrieve, modify, and store beamline parameters have not kept pace with the advancement of technology. Development of a new method of storing and retrieving beamline parameters resulted in the testing and installation of a time-series database as a potential replacement for the traditional relational database. InfluxDB was selected due to its self-hosted Open-Source version availability as well as the simplicity of installation and setup. A program was written to periodically gather all accelerator parameters in the control system and store them in the time-series database. This resulted in over 13,000 distinct data points, captured at 5-minute intervals. A second test captured 35 channels on a 1-minute cadence. Graphing of the captured data is being done on Grafana, an Open-Source version is available that co-exists well with InfluxDB as the back-end. Grafana made visualizing the data simple and flexible. The testing has allowed for the use of modern graphing tools to generate new insights into operating the accelerator, as well as opened the door to building large data sets suitable for Artificial Intelligence and Machine Learning applications.

Novak, D.↗

Mass agnostic jet taggers

Searching for new physics in large data sets needs a balance between two competing effects—signal identification vs background distortion. In this work, we perform a systematic study of both single variable and multivariate jet tagging methods that aim for this balance. The methods preserve the shape of the background distribution by either augmenting the training procedure or the data itself. Multiple quantitative metrics to compare the methods are considered, for tagging 2-, 3-, or 4-prong jets from the QCD background. This is the first study to show that the data augmentation techniques of Planing and PCA based scaling deliver similar performance as the augmented training techniques of Adversarial NN and uBoost, but are both easier to implement and computationally cheaper.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Water resource recovery modelling 2021 (WRRmod2021 conference)

Our society is transitioning fast into the digital age, spurred by development of cheap and new sensing technology, breakthroughs in computing, and development of efficient algorithms for optimization. This transition is also visible in the field of wastewater treatment and is driving new model developments, especially by exploiting the large data sets available with many utilities. Not surprisingly, WRRmod2021 had featured a strong session on ‘data-driven models and digitalization’ focused on this hot topic. At the same time, engineering practice calls for more robust models for performance evaluation and optimization of both conventional facilities and innovative processes. As a result, the WRRmod2021 program also exhibited sessions on modelling of new process units (e.g., aerobic granular sludge), modelling of the nitrogen cycle, and integrated/plant-wide modelling.

54 ENVIRONMENTAL SCIENCES↗

Nuclear Physics Network Requirements Review Report

The Energy Sciences Network (ESnet) is the Office of Science’s high-performance network user facility, delivering highly reliable data transport capabilities optimized for the requirements of data-intensive science. In essence, ESnet is the circulatory system that enables the U.S. Department of Energy (DOE) science mission by connecting each and every DOE lab and its user facilities. ESnet is funded and stewarded by the Advanced Scientific Computing Research (ASCR) Program and managed and operated by the Scientific Networking Division at Lawrence Berkeley National Laboratory (LBNL). ESnet is widely regarded as a global leader in the research and education networking community. ESnet connects DOE national laboratories, user facilities, and major experiments so scientists can use remote instruments and computing resources as well as share data with collaborators, transfer large data sets, and access distributed data repositories. While ESnet provides network connectivity, it cannot be characterized as an internet service provider as it is specifically built to provide a range of network services that are tailored to meet the unique requirements of DOE’s data-intensive science.

97 MATHEMATICS AND COMPUTING↗