Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data Distribution”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Interactively Assessing Disentanglement in GANs

Abstract Generative adversarial networks (GAN) have witnessed tremendous growth in recent years, demonstrating wide applicability in many domains. However, GANs remain notoriously difficult for people to interpret, particularly for modern GANs capable of generating photo‐realistic imagery. In this work we contribute a visual analytics approach for GAN interpretability, where we focus on the analysis and visualization of GAN disentanglement. Disentanglement is concerned with the ability to control content produced by a GAN along a small number of distinct, yet semantic, factors of variation. The goal of our approach is to shed insight on GAN disentanglement, above and beyond coarse summaries, instead permitting a deeper analysis of the data distribution modeled by a GAN. Our visualization allows one to assess a single factor of variation in terms of groupings and trends in the data distribution, where our analysis seeks to relate the learned representation space of GANs with attribute‐based semantic scoring of images produced by GANs. Through use‐cases, we show that our visualization is effective in assessing disentanglement, allowing one to quickly recognize a factor of variation and its overall quality. In addition, we show how our approach can highlight potential dataset biases learned by GANs.

Jeong, Sangwon↗

Overview of the distributed image processing infrastructure to produce the Legacy Survey of Space and Time

The Vera C. Rubin Observatory is preparing to execute the most ambitious astronomical survey ever attempted, the Legacy Survey of Space and Time (LSST). Currently the final phase of construction is under way in the Chilean Andes, with the Observatory’s ten-year science mission scheduled to begin in 2025. Rubin’s 8.4-meter telescope will nightly scan the southern hemisphere collecting imagery in the wavelength range 320–1050 nm covering the entire observable sky every 4 nights using a 3.2 gigapixel camera, the largest imaging device ever built for astronomy. Automated detection and classification of celestial objects will be performed by sophisticated algorithms on high-resolution images to progressively produce an astronomical catalog eventually composed of 20 billion galaxies and 17 billion stars and their associated physical properties. In this article we present an overview of the system currently being constructed to perform data distribution as well as the annual campaigns which reprocess the entire image dataset collected since the beginning of the survey. These processing campaigns will utilize computing and storage resources provided by three Rubin data facilities (one in the US and two in Europe). Each year a Data Release will be produced and disseminated to science collaborations for use in studies comprising four main science pillars: probing dark matter and dark energy, taking inventory of solar system objects, exploring the transient optical sky and mapping the Milky Way. Also presented is the method by which we leverage some of the common tools and best practices used for management of large-scale distributed data processing projects in the high energy physics and astronomy communities. We also demonstrate how these tools and practices are utilized within the Rubin project in order to overcome the specific challenges faced by the Observatory.

79 ASTRONOMY AND ASTROPHYSICS↗

Question-answering system extracts information on injection drug use from clinical notes

Background. Injection drug use (IDU) can increase mortality and morbidity. Therefore, identifying IDU early and initiating harm reduction interventions can benefit individuals at risk. However, extracting IDU behaviors from patients’ electronic health records (EHR) is difficult because there is no other structured data available, such as International Classification of Disease (ICD) codes, and IDU is most often documented in unstructured free-text clinical notes. Although natural language processing can efficiently extract this information from unstructured data, there are no validated tools. Methods. Here, to address this gap in clinical information, we design a question-answering (QA) framework to extract information on IDU from clinical notes for use in clinical operations. Our framework involves two main steps: (1) generating a gold-standard QA dataset and (2) developing and testing the QA model. We use 2323 clinical notes of 1145 patients curated from the US Department of Veterans Affairs (VA) Corporate Data Warehouse to construct the gold-standard dataset for developing and evaluating the QA model. We also demonstrate the QA model’s ability to extract IDU-related information from temporally out-of-distribution data. Results. Here, we show that for a strict match between gold-standard and predicted answers, the QA model achieves a 51.65% F1 score. For a relaxed match between the gold-standard and predicted answers, the QA model obtains a 78.03% F1 score, along with 85.38% Precision and 79.02% Recall scores. Moreover, the QA model demonstrates consistent performance when subjected to temporally out-of-distribution data. Conclusions. Our study introduces a QA framework designed to extract IDU information from clinical notes, aiming to enhance the accurate and efficient detection of people who inject drugs, extract relevant information, and ultimately facilitate informed patient care.

60 APPLIED LIFE SCIENCES↗

SILIA: software implementation of a multi-channel, multi-frequency lock-in amplifier for spectroscopy and imaging applications

In this work, we describe a software implementation of a multi-channel, multi-frequency Lock-in Amplifier (SILIA) to extract modulated signals from noisy data distributed over multiple channels of arbitrary number and size. This software implementation emulates the functionality of a multi-channel, multi-frequency lock-in amplifier in a post-processing step following data acquisition. Unlike most traditional lock-in amplifiers, SILIA can work with any number of input channels and is especially useful to analyze data distributed over many channels. We demonstrate the versatility and performance for extracting weak signals in spectroscopy and fluorescence microscopy. We also discuss more general applications and exhibit a method to automatically estimate error from a lock-in result.

47 OTHER INSTRUMENTATION↗

Data-Driven Distribution System Coordinated PV Inverter Control Using Deep Reinforcement Learning

The deployment of distributed solar photovoltaic (PV) systems has increased consistently over the past decades. High penetrations of PVs could cause a series of adverse grid impacts, such as voltage violations. The recent development of smart inverter technologies rises the incentives of developing PV control solutions that regulate the inverter output power and seeking the optimization on system operational objectives. This paper proposes a data-driven control solution based on deep reinforcement learning (DRL) to optimize PV inverters for voltage regulation. The proposed solution can minimize PV real power curtailment while maintaining network voltage at an acceptable range. Comparison results between the proposed DRL control algorithms with deep deterministic policy gradient (DDPG) and volt-var control on a real feeder in west Colorado highlight the advantage of the proposed framework in controlling the system voltage while minimizing the PV real power curtailment.

deep reinforcement learning↗

Influence of Alkyne Precursor Structure on Carbon Nanotube Chiral Distribution: Data-Dense Analysis Across Multiple Catalyst Types

Carbon nanotubes (CNTs) are a desirable material in the field of optoelectronics and semiconductors due to electronic properties (e.g., bandgap) that are dependent upon their chirality, defined by their diameter and lattice angle. Unfortunately, industrial-scale syntheses have yet to realize growth of a single desired chirality and instead rely on postsynthetic separation techniques to refine a chiral mixture, which increases process complexity and cost. Here, we studied the influence of precursor structure on chiral distribution, using a series of terminal alkyne precursors (acetylene, methylacetylene, vinylacetylene, 1-butyne, two enantiomers of 3-butyn-2-ol and a racemic mixture thereof) to grow CNTs across five transition-metal catalysts (Fe, FeMo, and three proportions of CoMo). Multiwavelength Raman spectroscopy on 5,145 spots (5 catalysts, 7 precursors, 3 lasers, and 49 distinct substrate locations on each) determined that acetylene grew the smallest diameter CNTs, while vinylacetylene produced fewer subnanometer CNTs. Though precursor structure did not dictate a uniform chiral shift, it was shown to broaden or narrow chiral distribution, while catalyst structure played a dominant role. In conclusion, this is consistent with metal-precursor binding occurring through unsaturated bonds in the hydrocarbons via the alkyne polymerization mechanism.

Carbon nanotubes↗

Behavioral and Population Data-Driven Distribution System Load Modeling

Distribution system residential load modeling and analysis for different geographic areas within a utility or an independent system operator territory are critical for enabling small-scale, aggregated distributed energy resources to participate in grid services under Federal Energy Regulatory Commission Order No. 2222 [1]. In this study, we develop a methodology of modeling residential load profiles in different geographic areas with a focus on human behavior impact. First, we construct a behavior-based load profile model leveraging state-of-the-art appliance models. We simulate human activity and occupancy using Markov chain Monte Carlo methods calibrated with the American Time Use Survey data set. Second, we link our model with cleaned Current Population Survey data from the U.S. Census Bureau. Finally, we populate two sets of 500 households using California and Texas census data, respectively, to perform an initial analysis of the load in different geographic areas with various group features (e.g., different income levels). To distinguish the effect of population behavior differences on aggregated load, we simulate load profiles for both sets assuming fixed physical household parameters and weather data. Analysis shows that average daily load profiles vary significantly by income and income dependency varies by locality.

American Time Use Survey↗

Visualization of Noisy and Less Noisy Computational Basis States in Quantum Computing

Quantum computing technology holds substantial promise as a reliable computational paradigm. However, current noisy intermediate scale quantum (NISQ) systems, are significantly impacted by noise originating from hardware inconsistencies. This noise causes errors and lowers output fidelity. So we must find which basis states cause errors. However, there are two main challenges in analyzing noise corresponding to basis states. First, the noise distribution data is high dimensional in nature, thereby making its analysis challenging. Second, although functional box plots have been used in the state of the art research to understand such a high dimensional data, they suffer from clutter and occlusion issues because of overplotting. In this study, we introduce an innovative visualization pipeline to address the aforementioned challenges to provide a clear depiction of noisy and less-noisy basis states. Specifically, our proposed visualization pipeline comprises three stages namely, low dimensional embedding, clustering, and violin plot visualization, to reduce visual clutter and effectively analyze high-dimensional noise distribution data. Our analysis uses quantum machine learning (QML) circuits as case study for drawing a distinction between noisy and less noisy basis states.

Senapati, Priyabrata [Kent State University]↗

EMPDF : inferring the Milky Way mass with data-driven distribution function in phase space

We introduce the emPDF (empirical distribution function), a novel dynamical modelling method that infers the gravitational potential from kinematic tracers with optimal statistical efficiency under the minimal assumption of steady state. emPDF determines the best-fitting potential by maximizing the similarity between instantaneous kinematics and the time-averaged phase-space distribution function (DF), which is empirically constructed from observation upon the theoretical foundation of oPDF (Han et al. 2016). This approach eliminates the need for presumed functional forms of DFs or orbit libraries required by conventional DF- or orbit-based methods. emPDF stands out for its flexibility, efficiency, and capability in handling observational effects, making it preferable to the popular Jeans equation or other minimal assumption methods, especially for the Milky Way (MW) outer halo where tracers often have limited sample size and poor data quality. We apply emPDF to infer the MW mass profile using Gaia DR3 data of satellite galaxies and globular clusters, obtaining enclosed masses of M (,r) = 26±8, 46±8, 90±13⁠, and 149±40 x 10 10 M ⊙ at r = 30, 50, 100⁠, and 200 kpc, respectively. These are consistent with the updated constraints from simulation-informed DF fitting (Li et al. 2020). While the simulation-informed DF offers superior precision owing to the additional information extracted from simulations, emPDF is independent of such supplementary knowledge and applicable to general tracer populations. emPDF is currently implemented for tracers with complete 6D kinematics within spherical potentials, but it can potentially be extended to address more general problems.

Astrophysics of Galaxies (astro-ph.GA)↗

Müllerian mimicry and the coloration patterns of sympatric coral snakes

Abstract Coral snakes in the genus Micrurus are venomous, aposematic organisms that signal danger to predators through vivid coloration. Previous studies found that they serve as models to several harmless species of Batesian mimics. However, the extent to which Micrurus species engage in Müllerian mimicry remains poorly understood. We integrate detailed morphological and geographical distribution data to investigate if coral snakes are Müllerian mimics. We found that coloration is spatially structured and that Micrurus species tend to be more similar where they co-occur. Though long supposed, we demonstrate for the first time that coral snakes might indeed be Müllerian mimics as they show some convergence in coloration patterns. Additionally, we found that the length of red-coloured rings in Micrurus is conserved, even at large geographic scales. This finding suggests that bright red rings may be under more substantial stabilizing selection than other aspects of coloration and probably function as a generalized signal for deterring predators.

Evolutionary Biology↗

Cross-Layered Distributed Data-Driven Framework for Enhanced Smart Grid Cyber-Physical Security

Smart Grid (SG) research and development has drawn much attention from academia, industry and government due to the great impact it will have on society, economics and the environment. Securing the SG is a considerably significant challenge due the increased dependency on communication networks to assist in physical process control, exposing them to various cyber-threats. In addition to attacks that change measurement values using False Data Injection (FDI) techniques, attacks on the communication network may disrupt the power system's real-time operation by intercepting messages, or by flooding the communication channels with unnecessary data. Addressing these attacks requires a cross-layer approach. In this paper a cross-layered strategy is presented, called Cross-Layer Ensemble CorrDet with Adaptive Statistics(CECD-AS), which integrates the detection of faulty SG measurement data as well as inconsistent network inter-arrival times and transmission delays for more reliable and accurate anomaly detection and attack interpretation. Numerical results show that CECD-AS can detect multiple False Data Injections, Denial of Service (DoS) and Man In The Middle (MITM) attacks with a high F1-score compared to current approaches that only use SG measurement data for detection such as the traditional physics-based State Estimation, Ensemble CorrDet with Adaptive Statistics strategy and other machine learning classification-based detection schemes.

cyber-physical security↗

RAPIDS: Reconciling Availability, Accuracy, and Performance in Managing Geo-Distributed Scientific Data

In modern science, big data plays an increasingly important role. Many scientific applications, such as running simulations on supercomputers or conducting experiments on advanced instruments, produce huge amount of data at unprecedented speed. Analyzing and understanding such big data is the key for scientists to make scientific breakthroughs. However, data might become unavailable for scientists to access when outages or maintenance of the storage system occur, which severely hinders scientific discovery. To improve the data availability, data duplication and erasure coding (EC) are often used. But as the scientific data gets larger, using these two methods can cause considerable storage and network overhead.In this paper, we propose RAPIDS, a hybrid approach that combines the multigrid-based error-bounded lossy compression with erasure coding, to significantly reduce the storage and network overhead required for maintaining high data availability. Our experiments show that RAPIDS reduces the storage overhead by up to 7.5x and network overhead by up to 3x to achieve the same level of availability compared to the regular EC method. We improve RAPIDS by building two models to optimize the fault tolerance configurations and data gathering strategy. We demonstrate that RAPIDS significantly improves performance when running on many CPU cores in parallel or on GPUs.

Wan, Lipeng↗

rmap: An R package to plot and compare tabular data on customizable maps across scenarios and time

`rmap` is an R package that allows users to easily plot tabular data (CSV or R data frames) on maps without any Geographic Information Systems (GIS) knowledge. Maps produced by `rmap` are `ggplot` objects and thus capitalize on the flexibility and advancements of the `ggplot2` package and all elements of each map are thus fully customizable. Additionally `rmap` automatically detects and produces comparison maps if the data has multiple scenarios or time periods as well as animations for time series data. Advanced users can load their own shapefiles if desired. `rmap` comes with a range of pre-built color palettes but users can also provide any `R` color palette or create their own as needed. Four different legend types are available to highlight different kinds of data distributions. The input spatial data can be both gridded or polygon data. `rmap` is desgined in particular for comparing spatial data across scenarios and time periods and comes preloaded with standard country, state, and basin maps as well as custom maps compatible with the Global Change Analysis Model (GCAM) spatial boundaries. `rmap` has a growing number of users and its products have been used in multiple multisector dynamics publications as well as a required dependency in other R packages such as `rfasst` and `metis`. `rmap's` automatic processing of tabular data using pre-built map selection, difference map calculations, faceting, and animations offers unique functionality which makes it a powerful and yet simple tool for users looking to explore multi-sector, multi-scenario data across space and time.

58 GEOSCIENCES↗