Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data Distribution”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Computing Angular Distributions from Simulation Data

The essential idea of this algorithm is to compute the angular distribution of a vector quantity, then create an informative image. In our example, we will compute the angular distribution of linear momentum from an xRage simulation of an exploding shaped charge. We will then explore one possible method for adding information to the resulting image.

97 MATHEMATICS AND COMPUTING↗

Geospatial Capabilities to Couple Hazard and Social Vulnerability Data in Water Distribution Criticality Analysis

A resilience analysis of a water distribution system is greatly enhanced by the integration of up-to-date geospatial data describing the water system, hazards, and surrounding community. The Water Network Tool for Resilience (WNTR), an open-source Python package designed to simulate and analyze the resilience of water distribution systems, was recently updated to incorporate geographic information system (GIS) data into the resilience analysis. This paper describes the GIS capabilities and includes a case study using the drinking water distribution system model for a large city in Pennsylvania. The case study focuses on potential pipe damage from landslides and on pipes that are particularly difficult to repair. The analysis couples data on hazards, social vulnerability, and the location of emergency services to identify and prioritize high-impact critical infrastructure for mitigation. Results demonstrate that pipes can be prioritized for mitigation based on water shortage and vulnerable populations that are affected. In conclusion, the methods can be adopted for general use and are available as part of the WNTR software.

GIS, landslide↗

Data-Driven Chance-Constrained Design of Voltage Droop Control for Distribution Networks: Preprint

This paper addresses the design of local control methods for voltage control in distribution networks with high level of distributed energy resources (DERs). The designed control methods adapt the active and reactive power output of distributed energy resources proportional to the deviation of the local measured voltage magnitudes from a reference voltage, which is referred to as droop control. Thus, the design focuses on determining the droop characteristics which satisfy network-wide voltage magnitude constraints. The uncertainty and variability of DERs renders the design of optimal droop controls very challenging. Hence, this paper proposes chance constraints to limit the risk from intermittent DERs, by designing droop control coefficients that guarantee the satisfaction of network operational constraints with a specific probability. In addition, the proposed approach relies entirely on historical data rather than assuming knowledge of the probability distributions that characterize the uncertainty of DERs. The efficacy of the proposed method is demonstrated on a 37-bus distribution feeder.

chance-constrained optimization↗

Deep Generative Models that Solve PDEs: Distributed Computing for Training Large Data-Free Models

Recent progress in scientific machine learning (SciML) has opened up the possibility of training novel neural network architectures that solve complex partial differential equations (PDEs). Several (nearly data free) approaches have been recently reported that successfully solve PDEs, with examples including deep feed forward networks, generative networks, and deep encoder-decoder networks. However, practical adoption of these approaches is limited by the difficulty in training these models, especially to make predictions at large output resolutions (≥1024×1024). Here we report on a software framework for data parallel distributed deep learning that resolves the twin challenges of training these large SciML models - training in reasonable time as well as distributing the storage requirements. Our framework provides several out of the box functionality including (a) loss integrity independent of number of processes, (b) synchronized batch normalization, and (c) distributed higher-order optimization methods. We show excellent scalability of this framework on both cloud as well as HPC clusters, and report on the interplay between bandwidth, network topology and bare metal vs cloud. We deploy this approach to train generative models of sizes hitherto not possible, showing that neural PDE solvers can be viably trained for practical applications. We also demonstrate that distributed higher-order optimization methods are 2-3× faster than stochastic gradient-based methods and provide minimal convergence drift with higher batch-size.

PDEs↗

Autonomous semantic data discovery for distributed networked systems

Systems, methods, techniques and apparatuses for managing distributed applications of networked intelligent agents are disclosed. The agents are operably to autonomously discover semantic profiles and associated data of other agents in a networked system participating in a given application. The agents need not be in direct communication with or known to all the other agents in the networked system.

Brissette, Alexander↗

Knowledge Beacons: Web services for data harvesting of distributed biomedical knowledge

The continually expanding distributed global compendium of biomedical knowledge is diffuse, heterogeneous and huge, posing a serious challenge for biomedical researchers in knowledge harvesting: accessing, compiling, integrating and interpreting data, information and knowledge. In order to accelerate research towards effective medical treatments and optimizing health, it is critical that efficient and automated tools for identifying key research concepts and their experimentally discovered interrelationships are developed. As an activity within the feasibility phase of a project called “Translator” (https://ncats.nih.gov/translator) funded by the National Center for Advancing Translational Sciences (NCATS) to develop a biomedical science knowledge management platform, we designed a Representational State Transfer (REST) web services Application Programming Interface (API) specification, which we call a Knowledge Beacon. Knowledge Beacons provide a standardized basic API for the discovery of concepts, their relationships and associated supporting evidence from distributed online repositories of biomedical knowledge. This specification also enforces the annotation of knowledge concepts and statements to the NCATS endorsed the Biolink Model data model and semantic encoding standards (https://biolink.github.io/biolink-model/). Implementation of this API on top of diverse knowledge sources potentially enables their uniform integration behind client software which will facilitate research access and integration of biomedical knowledge.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Feature Extraction for Pipeline Defects Inspection Based Upon Distributed Acoustic Fiber Optic Sensing Data

Fiber-optic distributed acoustic sensing (DAS) is becoming an increasingly important tool for real-time monitoring of energy and civil infrastructure structural health such as pipelines. We present a systematic theoretical study of the potential for DAS to be directly coupled with guided ultrasonic waves typically used in conventional acoustic non-destructive evaluation (NDE) methods for real-time pipeline health monitoring. We are referring to this innovative new NDE technique as ultrasonic guided wave and optical fiber sensor fusion. In the practical application of DAS coupled with guided ultrasonic waves, the structural design of (1) the specific guided waves excited, (2) the physical installation of the acoustic transducers and the fiber optic sensors, and (3) the functional performance specifications (gauge length, sensitivity, Etc.) of fiber optic DAS have an important influence on overall capabilities of the monitoring system. Meanwhile, physics-based analysis of acoustic waves is still a challenge due to the complex nature of the Lamb wave when it propagates, scatters, and disperses in the presence of structural defects. In this work, we simulate carbon steel pipes relevant for oil and gas pipeline applications with diameters of approximately 6-12” and wall thickness of 0.5” as the objects to be monitored. By establishing and implementing these capabilities, we seek to pursue an in-depth study on structural parameter optimization of DAS network, measurement range, and signal processing with an ultimate goal of increasing the sensitivity and efficacy of DAS to defect identification for various modes of corrosion expected in practice. To study the characteristics of scattered acoustic waves and performance of DAS for defect identification, we simulated the response of DAS for multiple pipe structures, defect types, and DAS sensor network configuration using finite element software Ansys, then the properties of signal response are extracted to construct defect-sensitive features. The raw data simulated, and the associated features extracted can ultimately be utilized as annotated training data to benchmark various designs for DAS applications, guided acoustic excitation sources, and learning model parameters to enhance early detection of potentially problematic defects.

Pipeline Defects Inspection, Fiber-optic sensors, ↗

Out-of-Distribution Detection and Radiological Data Monitoring Using Statistical Process Control

Abstract Machine learning (ML) models often fail with data that deviates from their training distribution. This is a significant concern for ML-enabled devices as data drift may lead to unexpected performance. This work introduces a new framework for out of distribution (OOD) detection and data drift monitoring that combines ML and geometric methods with statistical process control (SPC). We investigated different design choices, including methods for extracting feature representations and drift quantification for OOD detection in individual images and as an approach for input data monitoring. We evaluated the framework for both identifying OOD images and demonstrating the ability to detect shifts in data streams over time. We demonstrated a proof-of-concept via the following tasks: 1) differentiating axial vs. non-axial CT images, 2) differentiating CXR vs. other radiographic imaging modalities, and 3) differentiating adult CXR vs. pediatric CXR. For the identification of individual OOD images, our framework achieved high sensitivity in detecting OOD inputs: 0.980 in CT, 0.984 in CXR, and 0.854 in pediatric CXR. Our framework is also adept at monitoring data streams and identifying the time a drift occurred. In our simulations tracking drift over time, it effectively detected a shift from CXR to non-CXR instantly, a transition from axial to non-axial CT within few days, and a drift from adult to pediatric CXRs within a day—all while maintaining a low false positive rate. Through additional experiments, we demonstrate the framework is modality-agnostic and independent from the underlying model structure, making it highly customizable for specific applications and broadly applicable across different imaging modalities and deployed ML models.

Zamzmi, Ghada↗

WHONDRS Surface Water Geochemistry and Organic Matter Characterization Data from Streams Distributed across Latin America

This dataset supports a broader study examining global transferability of stream biogeochemistry and was generated in collaboration with the MicroSudAqua (µSudAqua) network (https://microsudaqua.netlify.app/en/). The dataset provides surface water geochemistry (dissolved organic carbon, total dissolved nitrogen, cations) and organic matter characterization (FTICR-MS) from streams in Argentina, Brazil, Chile, and Colombia. Samples were collected across stream orders (1st to 6th order) within five basins. Related data were collected and will be published separately in collaboration with the µSudAqua network. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. In addition to this readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes all column/row headers and variable definitions. This dataset is comprised of (1) a folder of field photos; (2) a folder of surface water sample data, (3) a folder of raw Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) data; (4) file-level metadata; (5) data dictionary; (6) field metadata; (7) readme; (8) international generic sample number (IGSN) mapping file; and (9) field protocol. The sample data subfolder contains (1) dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC) data and averages; (2) total dissolved nitrogen data and averages; (3) anions and averages; (4) methods codes; (5) FTICR-MS methods; and (15) a subfolder of 9.4 Tesla (9.4T) FTICR-MS data. This folder contains the processed data and three subfolders, one containing the .xml files, one containing the water CoreMS output files, and the other containing instructions and scripts for processing the files in CoreMS (https://github.com/EMSL-Computing/CoreMS). All files are .csv, .pdf, .R, .xml, .d, .html, .Rmd, .py, .cal, .json, .jpg, .jpeg, .png, .mov, or .mp4.

Anions↗

Estimation of Aerosol Columnar Size Distribution from Spectral Extinction Data in Coastal and Maritime Environment

Aerosol columnar size distributions (SDs) are commonly provided by aerosol inversions based on measurements of both spectral extinction and sky radiance. These inversions developed for a fully clear sky offer few SDs for areas with abundant clouds. Here, we estimate SDs from spectral extinction data alone for cloudy coastal and maritime regions using aerosol refractive index (RI) obtained from chemical composition data. Our estimation involves finding volume and mean radius of lognormally distributed modes of an assumed bimodal size distribution through fitting of the spectral extinction data. We demonstrate that vertically integrated SDs obtained from aircraft measurements over a coastal site have distinct seasonal changes, and these changes are captured reasonably well by the estimated columnar SDs. We also demonstrate that similar seasonal changes occur at a maritime site, and columnar SDs retrieved from the combined extinction and sky radiance measurements are approximated quite well by their extinction only counterparts (correlation exceeds 0.9) during a 7-year period (2013–2019). The level of agreement between the estimated and retrieved SDs depends weakly on wavelength selection within a given spectral interval (roughly 0.4–1 µm). Since the extinction-based estimations can be performed frequently for partly cloudy skies, the number of periods where SDs can be found is greatly increased.

54 ENVIRONMENTAL SCIENCES↗

An efficient method to propagate model uncertainty when inverting seismic data for time domain seismic moment tensors

SUMMARY We present a computationally efficient method to approximately propagate uncertainty when linearly inverting seismic data for point source, time variable moment tensor components. The method is based on the assumption that the data residual, given by the difference between the observed seismic data and the data predicated by a linear inversion, contains the effects of both data and model uncertainty. Our method uses a distribution of data residuals, added directly to the data, in a pseudo-Monte Carlo scheme. Using the assumption that the data residual is a stochastic process, we use the well-known Karhunen–Loève (KL) theorem to construct a distribution of data residuals, where the required basis functions are constructed using Fourier series. The Fourier series are scaled by a product of a random variable and the real-valued spectral amplitudes of the original data residual’s spectrum. Thus, the Fourier series and spectral amplitudes are eigenfunction-eigenvalue pairs used in the KL-based construction of data residual distribution. Using tests with synthetic data, we show that our method compares closely with a Finite Difference Monte Carlo (FDMC) method that we presented previously. More importantly, the method presented here is computationally several orders of magnitude faster than our previous FDMC method, and requires no a priori assumptions of model and/or data uncertainty.

Poppeliers, Christian (ORCID:0000000159526849)↗

Distribution Grid Modeling Using Smart Meter Data

The knowledge of distribution grid models, including topologies and line impedances, is essential for grid monitoring, control and protection. However, such information is often unavailable, incomplete or outdated. The increasing deployment of smart meters (SMs) provides a unique opportunity to tackle this issue. This paper proposes a two-stage framework for distribution grid modeling using SM data. In the first stage, the network topology is identified by reconstructing a weighted Laplacian matrix of distribution networks. In the second stage, a least absolute deviations (LAD) regression model is developed for estimating line impedance of a single branch based on the nonlinear (inverse) power flow model, wherein a conductor library is leveraged to narrow down the solution space. The LAD regression model is originally a mixed-integer nonlinear program whose continuous relaxation is still non-convex. Furthermore, we specially address its convex relaxation and discuss the exactness. The modified regression model is then embedded within a bottom-up sweep algorithm to achieve the identification across the network in a branch-wise manner. Numerical results on the IEEE 13-bus, 37-bus and 69-bus test feeders validate the effectiveness of the proposed methods.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Asynchronous and Load-Balanced Union-Find for Distributed and Parallel Scientific Data Visualization and Analysis

We present a novel distributed union-find algorithm that features asynchronous parallelism and k-d tree based load balancing for scalable visualization and analysis of scientific data. Applications of union-find include level set extraction and critical point tracking, but distributed union-find can suffer from high synchronization costs and imbalanced workloads across parallel processes. In this study, we prove that global synchronizations in existing distributed union-find can be eliminated without changing final results, allowing overlapped communications and computations for scalable processing. We also use a k-d tree decomposition to redistribute inputs, in order to improve workload balancing. We benchmark the scalability of our algorithm with up to 1,024 processes using both synthetic and application data. Here, we demonstrate the use of our algorithm in critical point tracking and super-level set extraction with high-speed imaging experiments and fusion plasma simulations, respectively.

97 MATHEMATICS AND COMPUTING↗

Deep learning ferroelectric polarization distributions from STEM data via with and without atom finding

Over the last decade, scanning transmission electron microscopy (STEM) has emerged as a powerful tool for probing atomic structures of complex materials with picometer precision, opening the pathway toward exploring ferroelectric, ferroelastic, and chemical phenomena on the atomic scale. Analyses to date extracting a polarization signal from lattice coupled distortions in STEM imaging rely on discovery of atomic positions from intensity maxima/minima and subsequent calculation of polarization and other order parameter fields from the atomic displacements. Here, we explore the feasibility of polarization mapping directly from the analysis of STEM images using deep convolutional neural networks (DCNNs). In this approach, the DCNN is trained on the labeled part of the image (i.e., for human labelling), and the trained network is subsequently applied to other images. We explore the effects of the choice of the descriptors (centered on atomic columns and grid-based), the effects of observational bias, and whether the network trained on one composition can be applied to a different one. This analysis demonstrates the tremendous potential of the DCNN for the analysis of high-resolution STEM imaging and spectral data and highlights the associated limitations.

36 MATERIALS SCIENCE↗