Data-Efficient Scientific Design Optimization with Neural Network Surrogates
Explore the source record for details and available documents.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
The overarching goal of the project is to develop an integrated open-source scientific data management system (Apache V2 license), ResonantEco, that meets the need of biological and environmental researchers and developers for data management, curation, and data processing for analyses with a wide range of scale and complexity. ResonantEco will provide web enabled data services with features such as unified data interfaces and federated views of data and metadata for heterogeneous data sources with an interactive web client for data exploration. Our use of the term fusion is taken from geospatial (GIS) domain where data fusion is often synonymous with data integration. In particular, data integration in ResonantEco involves combining data residing in different sources and providing users with a unified view of them.
With the advancement of exascale computing, the amount of scientific data is increasing day by day. Efficient data access is necessary for scientific discoveries. Unfortunately, the I/O performance is not improved, like the CPU and network speed. So, I/O operations take longer time than data generation or analysis. Asynchronous I/O has been proposed to extenuate the I/O bottleneck by overlapping I/O and computation time. However, multiple small write operations can diminish the benefits of asynchronous I/O, as the I/O time becomes significantly longer than the compute time, with little time to overlap with. To overcome these issues, we present an optimization technique to merge small contiguous write operations. We integrated our solution into the HDF5 asynchronous I/O VOL connector and demonstrated the effectiveness of merging HDF5 write operations automatically and transparently without requiring any code change from the application.
Cities, the world over, are fuelling economic growth. At the same time, rapid urbanization is a root cause of serious environmental damage. Recent WHO global air pollution guidelines highlight air pollution as a critical environmental threat along with climate change. To address these threats, smart cities and clean air programs are on a rise. In smart cities, data and Information and Communication Technologies (ICT) are major drivers of city transformations. The 4th Industrial Revolution (4IR) technologies such as the Internet of Things (IoT), big data, artificial intelligence (AI), and cloud computing have the potential to accelerate these transformations toward urban resilience. However, the success of smart cities and clean air programs depends on cohesive multi-sector stakeholder contributions. This study conducted interdisciplinary participative stakeholder analysis to understand the data, and sectorial challenges, to outline the technological opportunities to facilitate clean air programs in Indian smart cities. The research highlights gaps due to siloed stakeholder operations, lack of data calibration, non-alignment of smart city and air quality management services, non-availability of health exposure data, and difficulty in translating scientific data into implementable actions. Stakeholders expressed potential ‘fit for the purpose’ use of IoT devices, satellites, smartphones, and mobility data augmented by AI methods in bridging these gaps. In conclusion, the analysis points toward a need to develop an easily accessible and ubiquitous urban data governance ecosystem enabling seamless cross-sector data exchanges to build trusting relationships among the stakeholders across the air quality management value chain.
The ever-increasing volume of data produced by HPC simulations necessitates scalable methods for data exploration and knowledge extraction. Scientific data analysis often involves complex queries across distributed datasets, requiring manipulation of multiple primary variables and generating derived data that needs to be handled efficiently, creating challenges for applications that need to parse many large datasets. Relying on individual applications to handle all intermediate data generally leads to redundant computations across studies and unnecessary data transfers. In this paper, we investigate the performance of different approaches where applications define derived variables as quantities of interest (QoIs) and offload the computation and transfer of these QoIs to the I/O library. This significantly reduces redundancy and optimizes data movement across the distributed storage and processing infrastructure by allowing control over when and where derived variables are computed. We present a detailed analysis of the performance-storage trade-offs associated with different solutions and showcase results for our study on two large-scale datasets created from climate and combustion simulations.
The Atmospheric Radiation Measurement (ARM) User Facility provides one of the world's largest openly accessible repositories of atmospheric observations through the ARM Data Discovery platform. Although the repository contains more than three decades of measurements collected from permanent observatories, mobile facilities, aircraft campaigns, and field experiments, identifying appropriate datasets can be challenging, particularly for new users unfamiliar with ARM instrumentation and datastream organization. To improve data accessibility, the ARM Data Center developed the ARM Data Advisor (ADA), an artificial intelligence-powered assistant designed to facilitate scientific data discovery, dataset interpretation, and user guidance. This report evaluates ADA's performance as a domain-specific scientific assistant using realistic atmospheric science workflows. The evaluation examines five key capabilities: data retrieval and curation efficiency, hallucination resistance, scientific reasoning, response to ambiguous queries, and content retention and session continuity. Representative prompts were developed to simulate typical interactions between researchers and the ARM Data Discovery platform, and ADA's responses were assessed for retrieval completeness, scientific accuracy, consistency, and practical usefulness. In these representative tests, ADA reduced the complexity of discovering and accessing ARM datasets by recommending appropriate datastreams, explaining instrumentation, interpreting metadata, and assisting with data processing workflows. ADA also exhibits strong domain knowledge of atmospheric science terminology and generally resists hallucination by acknowledging unavailable datasets and requesting clarification when appropriate. Overall, the results indicate that ADA represents a promising advancement in scientific data discovery within the ARM User Facility and has considerable potential to improve researcher productivity, particularly for new users and interdisciplinary scientists seeking efficient access to ARM observations.
Data- and computationally-intensive scientific research, such as numerical simulations and inversions or the training of large neural networks in machine learning applications, that are well suited for HPC environments also often require expert insight and evaluation throughout the computation which can be greatly facilitated with the use of interactive computing tools, such as those in the Jupyter ecosystem. HPC workflows and interactive workflows are typically treated as orthogonal, however, the next generation of research will require both. The first challenge we face in this project is thus designing the right level of abstractions to allow interactive capabilities in the JupyterLab environment to allow the working scientist to flexibly explore and query their data at multiple levels, with a minimal amount of customization required of the underlying optimized codes. In addition to these questions regarding the high-level representation of data for interactive use in HPC, we tackled two additional issues that are part of the entire lifecycle of research and that become particularly acute in HPC contexts: how to improve the experience of interfacing with the HPC system's scheduling environment for a scientist focused on exploratory questions, and how can that scientist then best share the results of their work with others in a self-contained, reproducible manner.
Correction to: Scientific Data, published online 25 May 2022 The order of the column headings was incorrect in Table 1, meaning values were listed under the wrong descriptions. This has been corrected in the pdf and HTML versions of the article.
Improvements in computational and experimental capabilities are rapidly increasing the amount of scientific data that are routinely generated. In applications that are constrained by memory and computational intensity, excessively large datasets may hinder scientific discovery, making data reduction a critical component of data-driven methods. Datasets are growing in two directions: the number of data points and their dimensionality. Whereas dimension reduction typically aims at describing each data sample on lower-dimensional space, the focus here is on reducing the number of data points. A strategy is proposed to select data points such that they uniformly span the phase-space of the data. The algorithm proposed relies on estimating the probability map of the data and using it to construct an acceptance probability. An iterative method is used to accurately estimate the probability of the rare data points when only a small subset of the dataset is used to construct the probability map. Instead of binning the phase-space to estimate the probability map, its functional form is approximated with a normalizing flow. Therefore, the method naturally extends to high-dimensional datasets. The proposed framework is demonstrated as a viable pathway to enable data-efficient machine learning when abundant data are available.
Ever-increasing data size raises many challenges for scientific data analysis. Particularly in cosmological N-body simulation, finding the center of a dark matter halo suffers heavily from the large computational cost associated with the large number of particles (up to 20 million). In this work, we exploit the latent structure embed in a halo, and we propose a hierarchical approach to approximate the exact gravitational potential calculation for each particle in order to more efficiently find the halo center. Tests of our method on data from N-body simulations show that in many cases the hierarchical algorithm performs significantly faster than existing methods with a desirable accuracy.
MGARD (MultiGrid Adaptive Reduction of Data) is an algorithm for compressing and refactoring scientific data, based on the theory of multigrid methods. The core algorithm is built around stable multilevel decompositions of conforming piecewise linear $C^0$ finite element spaces, enabling accurate error control in various norms and derived quantities of interest. In this work, we extend this construction to arbitrary order Lagrange finite elements $\mathbb{Q}_p$, $p \geq 0$, and propose a reformulation of the algorithm as a lifting scheme with polynomial predictors of arbitrary order. Additionally, a new formulation using a compactly supported wavelet basis is discussed, and an explicit construction of the proposed wavelet transform for uniform dyadic grids is described.
The Adaptable I/O System (ADIOS) represents the culmination of substantial investment in Scientific Data Management, and it has demonstrated success for several important extreme-scale science cases. However, looking towards the exascale and beyond, we see the development of yet more stringent data management requirements that require new abstractions. Therefore, there is an opportunity to attempt to connect the traditional realms of HPC I/O optimization with the Database / Data Management community. As such, in this paper we offer some specific examples from our ongoing work in managing data structures, services, and performance at the extreme scale for scientific computing. Using the publish/subscribe model afforded by ADIOS, we demonstrate a set of services that connect data format, metadata, queries, data reduction, and high-performance delivery. The resulting publish/subscribe framework facilitates connection to on-line workflow systems to enable the dynamic capabilities that will be required for exascale science.
As the growth of data sizes continues to outpace computational resources, there is a pressing need for data reduction techniques that can significantly reduce the amount of data and quantify the error incurred in compression. Compressing scientific data presents many challenges for reduction techniques since it is often on non-uniform or unstructured meshes, is from a high-dimensional space, and has many Quantities of Interests (QoIs) that need to be preserved. To illustrate these challenges, we focus on data from a large scale fusion code, XGC. XGC uses a Particle-In-Cell (PIC) technique which generates hundreds of PetaBytes (PBs) of data a day, from thousands of timesteps. XGC uses an unstructured mesh, and needs to compute many QoIs from the raw data, f.One critical aspect of the reduction is that we need to ensure that QoIs derived from the data (density, temperature, flux surface averaged momentums, etc.) maintain a relative high accuracy. We show that by compressing XGC data on the high-dimensional, nonuniform grid on which the data is defined, and adaptively quantizing the decomposed coefficients based on the characteristics of the QoIs, the compression ratios at various error tolerances obtained using a multilevel compressor (MGARD) increases more than ten times. We then present how to mathematically guarantee that the accuracy of the QoIs computed from the reduced f is preserved during the compression. We show that the error in the XGC density can be kept under a user-specified tolerance over 1000 timesteps of simulation using the mathematical QoI error control theory of MGARD, whereas traditional error control on the data to be reduced does not guarantee the accuracy of the QoIs.
Here, a semisupervised machine learning method for the discovery of structure-spectrum relationships is developed and then demonstrated using the specific example of interpreting x-ray absorption near-edge structure (XANES) spectra. This method constructs a one-to-one mapping between individual structure descriptors and spectral trends. Specifically, an adversarial autoencoder is augmented with a rank constraint (RankAAE). The RankAAE methodology produces a continuous and interpretable latent space, where each dimension can track an individual structure descriptor. As a part of this process, the model provides a robust and quantitative measure of the structure-spectrum relationship by decoupling intertwined spectral contributions from multiple structural characteristics. This makes it ideal for spectral interpretation and the discovery of descriptors. The capability of this procedure is showcased by considering five local structure descriptors and a database of >50 000 simulated XANES spectra across eight first-row transition metal oxide families. The resulting structure-spectrum relationships not only reproduce known trends in the literature but also reveal unintuitive ones that are visually indiscernible in large datasets. The results suggest that the RankAAE methodology has great potential to assist researchers in interpreting complex scientific data, testing physical hypotheses, and revealing patterns that extend scientific insight.
DataFed is a scientific data management system for big data providing simple and uniform data access, organization, discovery, and sharing within and across scientific facilities - with the goals of enhancing productivity and scientific reproducibility. We hope to aid in communicating development contributions to DataFed over the last year and any planning we can provide for the future developments in the next year
Metadata management is one of three major areas and parts of functionality of scientific data management along with replica management and workflow management. Metadata is the information describing the data stored in a data item, a file or an object. It includes the data item provenance, recording conditions, format and other attributes. MetaCat is a metadata management database designed and developed for High Energy Physics experiments. As a component of a data management system, it’s main objectives are to provide efficient metadata storage and management and fast data items selection functionality. MetaCat is supposed to work on the scale of 100 million files (or objects) and beyond. The article will discuss the functionality of MetaCat and technological solutions used to implement the product.
The High Luminosity upgrade to the LHC (HL-LHC) is expected to deliver scientific data at the multi-exabyte scale. In order to address this unprecedented data storage challenge, the ATLAS experiment launched the Data Carousel project in 2018. Data Carousel is a tape-driven workflow whereby bulk production campaigns with input data resident on tape are executed by staging and promptly processing a sliding window to disk buffer such that only a small fraction of inputs are pinned on disk at any one time. Data Carousel is now in production for ATLAS in Run3. In this paper, we provide updates on recent Data Carousel R&D projects, including data-on-demand and tape smart writing. Data-on-demand removes from disk data that has not been accessed for a predefined period, when users request them, they will be either staged from tape or recreated by following the original production steps. Tape smart writing employs intelligent algorithms for file placement on tape in order to retrieve data back more efficiently, which is our long term strategy to achieve optimal tape usage in Data Carousel.
DataFlow is a web application that helps scientific data to flow from one source location to another destination location. DataFlow helps scientists easily capture scientific metadata associated with an experiment and transmit both metadata and experimental data to a designated, centralized data storage resource. This report describes the software engineering efforts and architecture of the project for the fiscal year 2021 developments. We hope it effectively communicates findings from our work, challenges we have overcome, and how we will continue our development of DataFlow in the future.