Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “distributed databases”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

GriddingMachine, a database and software for Earth system modeling at global and regional scales

Land and Earth system modeling is moving towards more explicit biophysical representations, requiring increasing variety of datasets for initialization and benchmarking. However, researchers often have difficulties in identifying and integrating non-standardized datasets from various sources. We aim towards a standardized database and one-stop distribution method of global datasets. Here, we present the GriddingMachine as (1) a database of global-scale datasets commonly used to parameterize or benchmark the models, from plant traits to vegetation indices and geophysical information and (2) a cross-platform open source software to download and request a subset of datasets with only a few lines of code. The GriddingMachine datasets can be accessed either manually through traditional HTTP, or automatically using modern programming languages including Julia, Matlab, Octave, Python, and R. The GriddingMachine collections can be used for any land and Earth modeling framework and ecological research at the regional and global scales, and the number of datasets will continue to grow to meet the increasing needs of research communities.

58 GEOSCIENCES↗

Analysis of Spectral Lines in Large Databases of Synthetic Spectra for Massive Stars

In this paper, we describe a program that identifies in the optical spectrum the main parameters of a spectral line, namely the initial and final wavelengths, and the line depth. Moreover, using numerical calculations, it identifies and removes adjacent lines. Next, the program calculates the equivalent width and the FWHM. The software was tested in a sample of 300 lines in two databases of synthetic spectra generated by the CMFGEN and PoWR codes, and 300 lines in observed spectra from the IACOB database, showing a Gaussian distribution of relative errors, from which it is inferred that 80% of the measured lines have errors less than 17% and only 5% of the lines have errors greater than 26%. The program was also run on the entire database of 45,000 CMFGEN and 202 POWR synthetic spectra, generating a library of H i, He i, and He ii lines necessary to feed the FITspec code for the derivation of stellar parameters: effective temperature, surface gravity, and luminosity.

47 OTHER INSTRUMENTATION↗

Phase Identification in Real Distribution Networks with High PV Penetration Using Advanced Metering Infrastructure Data

Many distribution network monitoring and control applications - including state estimation, volt/VAR optimization, and network reconfiguration - rely on accurate network models; however, the network models maintained by utilities can become outdated because of restoration activities, network reconfiguration, and missing data. With the widespread deployment of advanced metering infrastructure (AMI), abundant measurement data from low-voltage secondary networks are available. The AMI measurement data can be used for phase identification to improve the network models. Although the existing phase identification techniques work well in passive distribution feeders that do not have photovoltaic (PV) generation, they can fail to accurately identify the phases in the presence of PV. This paper proposes a robust phase identification algorithm based on supervised machine learning that accurately identifies the AMI meter phase connectivity in the presence of significant PV generation. The proposed algorithm does not require network topology information or feeder head measurement data. The algorithm is validated using the AMI measurement data collected in the field and the field-validated phase connectivity database on two real distribution feeders from San Diego Gas & Electric Company that have significant PV generation.

advanced metering infrastructure↗

Phase Identification in Real Distribution Networks with High PV Penetration Using Advanced Metering Infrastructure Data: Preprint

Many distribution network monitoring and control applications - including state estimation, volt/VAR optimization, and network reconfiguration - rely on accurate network models; however, the network models maintained by utilities can become outdated because of restoration activities, network reconfiguration, and missing data. With the widespread deployment of advanced metering infrastructure (AMI), abundant measurement data from low-voltage secondary networks are available. The AMI measurement data can be used for phase identification to improve the network models. Although the existing phase identification techniques work well in passive distribution feeders that do not have photovoltaic (PV) generation, they can fail to accurately identify the phases in the presence of PV. This paper proposes a robust phase identification algorithm based on supervised machine learning that accurately identifies the AMI meter phase connectivity in the presence of significant PV generation. The proposed algorithm does not require network topology information or feeder head measurement data. The algorithm is validated using the AMI measurement data collected in the field and the field-validated phase connectivity database on two real distribution feeders from San Diego Gas & Electric Company that have significant PV generation.

advanced metering infrastructure↗

Phase Identification in Real Distribution Networks with High PV Penetration Using Advanced Metering Infrastructure Data

Many distribution network monitoring and control applications - including state estimation, Volt/VAr optimization, and network reconfiguration - rely on accurate network models; however, the network models maintained by utilities can become outdated because of restoration activities, network reconfiguration, and missing data. With the widespread deployment of advanced metering infrastructure (AMI), abundant measurement data from low-voltage secondary networks are available. The AMI measurement data can be used for phase identification to improve the network models. Although the existing phase identification techniques work well in passive distribution feeders that do not have photovoltaic (PV) generation, they can fail to accurately identify the phases in the presence of PV. This paper proposes a robust phase identification algorithm based on supervised machine learning that accurately identifies the AMI meter phase connectivity in the presence of significant PV generation. The proposed algorithm does not require network topology information or feeder-head measurement data. The algorithm is validated using the AMI measurement data collected in the field and the field-validated phase connectivity database on two real distribution feeders from San Diego Gas & Electric Company that have significant PV generation.

advanced metering infrastructure (AMI)↗

New Architecture to Support Integration and Processing of Seismic Data from Heterogeneous Sources

The Geophysical Monitoring Program (GMP) at Lawrence Livermore National Lab (LLNL) maintains a database and supporting infrastructure for geophysical data used in support of the Nuclear Detonation Detection mission. This database includes data from multiple sources, many of which do not distribute data to the public or for which there is no automated means of access. For example, Figure 1 shows (left) the distribution of waveform data in our database by source. The Incorporated Research Institutions for Seismology Data Management Center (IRISDMC) is our major source of waveform data and those data may be retrieved at will using the Federated Digital Seismograph Networks FDSN web Application Programming Interface (API). However, the next 6 most important sources of waveform data have no or only limited automated access to waveforms. As Figure 1 (right) shows, it is very common for waveform records associate with an event in our database to come from two or more sources, and in some cases data come from 10 sources. This diversity of data sources drives our need for efficient and correct integration of metadata, parametric data, and waveform data.

58 GEOSCIENCES↗

Enabling Modular Autonomous Feedback‐Loops in Materials Science through Hierarchical Experimental Laboratory Automation and Orchestration

Abstract Materials acceleration platforms (MAPs) operate on the paradigm of integrating combinatorial synthesis, high‐throughput characterization, automatic analysis, and machine learning. Within a MAP, one or multiple autonomous feedback loops may aim to optimize materials for certain functional properties or to generate new insights. The scope of a given experiment campaign is defined by the range of experiment and analysis actions that are integrated into the experiment framework. Herein, the authors present a method for integrating many actions within a hierarchical experimental laboratory automation and orchestration (HELAO) framework. They demonstrate the capability of orchestrating distributed research instruments that can incorporate data from experiments, simulations, and databases. HELAO interfaces laboratory hardware and software distributed across several computers and operating systems for executing experiments, data analysis, provenance tracking, and autonomous planning. Parallelization is an effective approach for accelerating knowledge generation provided that multiple instruments can be effectively coordinated, which the authors demonstrate with parallel electrochemistry experiments orchestrated by HELAO. Efficient implementation of autonomous research strategies requires device sharing, asynchronous multithreading, and full integration of data management in experimental orchestration, which to the best of the authors’ knowledge, is demonstrated for the first time herein.

36 MATERIALS SCIENCE↗

North American Lithium-Ion Battery Supply Chain Database Development - Phase II

Lithium-ion batteries (LIBs) are used in a wide range of applications, including cell phones, laptops, power tools, electric vehicles, and grid storage, and are essential for economic growth and addressing climate change. However, the significant demand for LIBs has led to supply chain issues for the United States, as China dominates the processing of battery materials and battery production. To address this concern, NAATBatt International, a trade association of North American battery companies, supported the National Renewable Energy Laboratory in developing a database of companies that mine, process, manufacture, reuse, and recycle batteries in North America. The purpose of this database was to identify strengths and gaps in the supply chain, so that private-government partnerships could develop strategies to create a competitive LIB supply chain in the US. NREL published the first version of this database in 2021 and the second version in 2022. The database includes companies that have a manufacturing facility in North America and are engaged in materials, cells, packs, end-of-life management, as well as those involved in LIB battery modeling, distribution, service and repair, and R&D. In this presentation, we will discuss our approach to collecting data and categorizing various segments and products. We will also provide a summary of the data and present various maps to illustrate the distribution of companies in the database.

ADVANCED PROPULSION SYSTEMS,ENERGY STORAGE↗

Modeling Nanoconfinement Effects Using Active Learning

Predicting the spatial configuration of gas in nanopores of is relevant in applications such as fluid flow forecasting and hydrocarbon reserves estimation. For example, shale reservoirs have suffered from computationally intractable multiscale problems, since fluid properties such as viscosity, density, and adsorption must be calculated by using expensive molecular dynamics (MD) simulations within each nanopore, whereas flow through these connected nanopores must be simulated at the micrometer scale. We utilize machine learning techniques to quickly and accurately model nanoscale confinement effects as an important step toward bridging the nano and micro scales. Our workflow is based on building and training physics-based deep-neural-networks models by learning from a database of MD calculations. The model accounts for the adsorption phenomenon by predicting the statistical distribution of gas inside nanopores. Because large databases of MD calculations are expensive to create, we investigate active learning (AL) as a data set construction strategy. In this workflow, new data are selected based on the model uncertainty via the query-by-committee approach. We show that our workflow obtains accurate models that generalize to real scanning electron microscopy geometries with 1/10th of the number of MD calculations required vs random data set generation. Our method enables the possibility of modeling nanoconfinement effects at the mesoscale, where complex connected sets of nanopores affect flow.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

The Materials Provenance Store

Abstract We present a database resulting from high throughput experimentation, primarily on metal oxide solid state materials. The central relational database, the Materials Provenance Store (MPS), manages the metadata and experimental provenance from acquisition of raw materials, through synthesis, to a broad range of materials characterization techniques. Given the primary research goal of materials discovery of solar fuels materials, many of the characterization experiments involve electrochemistry, along with optical, structural, and compositional characterizations. The MPS is populated with all information required for executing common data queries, which typically do not involve direct query of raw data. The result is a database file that can be distributed to users so that they can independently execute queries and subsequently download the data of interest. We propose this strategy as an approach to manage the highly heterogeneous and distributed data that arises from materials science experiments, as demonstrated by the management of over 30 million experiments run on over 12 million samples in the present MPS release.

36 MATERIALS SCIENCE↗

Crowd-based spatial risk assessment of urban flooding: Results from a municipal flood hotline in Detroit, MI

Climate change is increasing the frequency and intensity of extreme precipitation events, raising the risk of urban flood disasters. This study uses a crowd-sourced municipal call database to characterize the spatial distribution of flood risk in Detroit, MI. Call data including dates and addresses were obtained from the City of Detroit Department of Public Works for 2021. Calls were mapped and aggregated to census tract counts and merged with neighborhood-level data. Associations of predictors with flood calls were tested using spatial regression models. Flooding calls were located throughout the city but were concentrated in specific areas. Multivariate models of census tract level call counts indicated that increased poverty and Black, immigrant, and older residents were positively associated with flood calls, while increased elevation was associated with protective effects. Longer distances from waste water interceptors were associated with higher risk for calls. Crowd-sourced flood hotline call data can be used for effective spatial flood risk assessment. Though flooding occurs throughout the city of Detroit, infrastructural, neighborhood, and household factors influence flooding extent. Limitations included the self-reported nature of calls. Future modeling efforts might include input from local stakeholders to improve spatial risk assessment.

54 ENVIRONMENTAL SCIENCES↗

Software and computing for Run 3 of the ATLAS experiment at the LHC

The ATLAS experiment has developed extensive software and distributed computing systems for Run 3 of the LHC. These systems are described in detail, including software infrastructure and workflows, distributed data and workload management, database infrastructure, and validation. The use of these systems to prepare the data for physics analysis and assess its quality are described, along with the software tools used for data analysis itself. An outlook for the development of these projects towards Run 4 is also provided.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

A functional microbiome catalogue crowdsourced from North American rivers

Predicting elemental cycles and maintaining water quality under increasing anthropogenic influence requires knowledge of the spatial drivers of river microbiomes. However, understanding of the core microbial processes governing river biogeochemistry is hindered by a lack of genome-resolved functional insights and sampling across multiple rivers. Here we used a community science effort to accelerate the sampling, sequencing and genome-resolved analyses of river microbiomes to create the Genome Resolved Open Watersheds database (GROWdb). GROWdb profiles the identity, distribution, function and expression of microbial genomes across river surface waters covering 90% of United States watersheds. Specifically, GROWdb encompasses microbial lineages from 27 phyla, including novel members from 10 families and 128 genera, and defines the core river microbiome at the genome level. GROWdb analyses coupled to extensive geospatial information reveals local and regional drivers of microbial community structuring, while also presenting foundational hypotheses about ecosystem function. Building on the previously conceived River Continuum Concept, we layer on microbial functional trait expression, which suggests that the structure and function of river microbiomes is predictable. We make GROWdb available through various collaborative cyberinfrastructures, so that it can be widely accessed across disciplines for watershed predictive modelling and microbiome-based management practices.

59 BASIC BIOLOGICAL SCIENCES↗

EQSIM—A multidisciplinary framework for fault-to-structure earthquake simulations on exascale computers, part II: Regional simulations of building response

The existing observational database of the regional-scale distribution of strong ground motions and measured building response for major earthquakes continues to be quite sparse. As a result, details of the regional variability and spatial distribution of ground motions, and the corresponding distribution of risk to buildings and other infrastructure, are not comprehensively understood. Utilizing high-performance computing platforms, emerging high-resolution, physics-based ground motion simulations can now resolve frequencies of engineering interest and provide detailed synthetic ground motions at high spatial density. This provides an opportunity for new insight into the distribution of infrastructure seismic demands and risk. In the work presented herein, the EQSIM fault-to-structure computational framework described in a companion paper, McCallen et al., is employed to investigate the regional-scale response of buildings to large earthquakes. A representative M = 7.0 strike-slip event is used to explore the distribution and amplitude of building demand, and comparisons are made between building response computed with fault-to-structure simulations and building response computed with existing measured near-fault earthquake records. New information on the distribution and variability of building response from high-performance parallel simulations is described and analyzed, and favorable first comparisons between building response predicted with both fault-to-structure simulations and real ground motions records are presented.

58 GEOSCIENCES↗

From Data to Knowledge: A Graph-Based Reliability Approach to Assess System Health

With the goal of maximizing plant reliability and availability, complex systems such as nuclear power plants continuously monitor and record the performance and the health status of many components, assets, and systems. Such data may take the form of online monitoring data, condition reports, and maintenance reports and it carries the potential to provide system engineers with insights into anomalous behaviors or degradation trends as well as the possible causes behind them and to predict their direct consequences. The analysis of such data poses however few challenges. While some of these challenges are technical in nature (i.e., data are often distributed over several physical servers or databases), others are conceptual in nature (i.e., data elements come in different formats, numeric or textual), and measured values have different scales (e.g., vibration spectra and oil temperature). This paper directly tackles these challenges, and it focuses on the integration of all these data elements in order to assist plant system engineers in analyzing component, assets, and systems performances and optimize maintenance activities. This is performed by 1) extracting knowledge from textual data via technical language processing methods, and 2) quantifying system, asset, and component health from numeric condition-based data. We rely on model-based system engineering (MBSE) models of systems and assets to identify their architecture and functional (i.e., cause and effect) relations. Numeric and textual data elements are then associated with an MBSE graph element, based on their nature. This bonding of MBSE models and data elements constitutes a first-of-its-kind knowledge graph of a nuclear power plants system, with data elements being organized in a structured manner that enables system engineers to identify cause-effect trends in data elements and carry out appropriate actions in response.

97 MATHEMATICS AND COMPUTING↗

A Knowledge Graph Approach to Analyze Systems and Assets Health

Nuclear power plants collect large amounts of equipment reliability data elements that contain information on the statuses of component, assets, and systems. All these data elements precisely record asset and system performance and health throughout the lifecycle of those assets and systems. However, several challenges have proved to be roadblocks to this process. While some of these challenges are technical in nature (i.e., data are often distributed over several physical servers or databases), others are conceptual in nature (i.e., data elements come in different formats, numeric or textual), and measured values have different scales (e.g., vibration spectra and oil temperature). This paper directly focuses on the integration of numeric and textual data elements in order to assist plant system engineers in analyzing equipment reliability data. This task begins with preprocessing the data by extracting knowledge from textual data via natural language processing methods and quantifying system, asset, and component health based on numeric data. We then employed model-based system engineering (MBSE) models of systems and assets to identify their architecture and functional (i.e., cause and effect) relations. Data elements were then associated with a single MBSE graph element, based on their nature. This bonding of MBSE models and data elements constitutes a first-of-its-kind knowledge graph of a nuclear power plants system, with data elements being organized in a structured manner that enables system engineers to identify cause-effect trends in data elements and carry out appropriate actions in response.

97 - MATHEMATICS AND COMPUTING↗

New approaches to an old problem: addressing spatial gaps in the World Stress Map

Abstract A well-recognized characteristic of the World Stress Map (WSM) database is the continued presence of large spatial gaps in the distribution of the data records despite the more than 40-year development history of the database. The current release has more than 30 000 high-quality (A–C) data records (often referred to as ‘stress indicators’), but while some continental areas (such as Australia) have seen a significant increase in spatial converge with the latest release, other continental regions (Africa, central Asia, most of South America) remain markedly sparse. In this contribution we (1) review the current state of the spatial distribution of stress indicators in the continental regions (above sea-level); (2) quantify the clustering of the stress indicators in the latest WSM release using the Hopkins statistic as a way to explore the current spatial distribution of the indicators and assess future WSM releases; and (3) present three approaches (joint inversion, seismic anisotropy and InSAR) that provide a way to fill in the gaps (both in the S Hmax orientation and principal stress magnitudes) in regions that lack active seismicity and where borehole drilling is cost prohibitive. These three approaches have the potential to guide procedures for improving a priori estimates of the ambient stress field in the Earth's crust and reduce the uncertainty in predicting both the magnitude and orientation of the principal tectonic stresses.

58 GEOSCIENCES↗