Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data Sources”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

An Inventory of AI-ready Benchmark Data for US Fires, Heatwaves, and Droughts

Extreme weather events, including fires, heatwaves, and droughts, have significant impacts on earth, environmental, and energy systems. Mechanistic and predictive understanding, as well as probabilistic risk assessment of these extreme weather events, are crucial for detecting, planning for, and responding to these extremes. Records of extreme weather events provide an important data source for understanding present and future extremes, but the existing data needs preprocessing before it can be used for analysis. Moreover, there are many nonstandard metrics defining the levels of severity or impacts of extremes. In this study, we compile a comprehensive benchmark data inventory of extreme weather events, including fires, heatwaves, and droughts. The dataset covers the period from 2001 to 2020 with a daily temporal resolution and a spatial resolution of 0.5°×0.5° (~55km×55km) over the continental United States (CONUS), and a spatial resolution of 1km × 1km over the Pacific Northwest (PNW) region, together with the co-located and relevant meteorological variables. By exploring and summarizing the spatial and temporal patterns of these extremes in various forms of marginal, conditional, and joint probability distributions, we gain a better understanding of the characteristics of climate extremes. The resulting AI/ML-ready data products can be readily applied to ML-based research, fostering and encouraging AI/ML research in the field of extreme weather. This study can contribute significantly to the advancement of extreme weather research, aiding researchers, policymakers, and practitioners in developing improved preparedness and response strategies to protect communities and ecosystems from the adverse impacts of extreme weather events. Usage Notes We presented a long term (2001-2020) and comprehensive data inventory of historical extreme events with daily temporal resolution covering the separate spatial extents of CONUS (0.5°×0.5°) and PNW(1km×1km) for various applications and studies. The dataset with 0.5°×0.5° resolution for CONUS can be used to help build more accurate climate models for the entire CONUS, which can help in understanding long-term climate trends, including changes in the frequency and intensity of extreme events, predicting future extreme events as well as understanding the implications of extreme events on society and the environment. The data can also be applied for risk accessment of the extremes. For example, ML/AI models can be developed to predict wildfire risk or forecast HWs by analyzing historical weather data, and past fires or heateave , allowing for early warnings and risk mitigation strategies. Using this dataset, AI-driven risk assessment models can also be built to identify vulnerable energy and utilities infrastructure, imrpove grid resilience and suggest adaptations to withstand extreme weather events. The high-resolution 1km×1km dataset ove PNW are advantageous for real-time, localized and detailed applications. It can enhance the accuracy of early warning systems for extreme weather events, helping authorities and communities prepare for and respond to disasters more effectively. For example, ML models can be developed to provide localized HW predictions for specific neighborhoods or cities, enabling residents and local emergency services to take targeted actions; the assessment of drought severity in specific communities or watersheds within the PNW can help local authorities manage water resources more effectively.

Lin, Xinming↗

The Illumination of Thunderclouds by Lightning: 3. Retrieving Optical Source Altitude

Abstract Optical space‐based lightning sensors such as the Geostationary Lightning Mapper (GLM) detect and geolocate lightning by recording rapid changes in cloud top illumination. While lightning locations can be determined to within a pixel on the GLM imaging array, these instruments are not individually able to natively report lightning altitude. It has previously been shown that thunderclouds are illuminated differently based on the altitude of the optical source. In this study, we examine how altitude information can be extracted from the spatial distributions of GLM energy recorded from each optical pulse. We match GLM “groups” with Lightning Mapping Array (LMA) source data that accurately report the 3‐D positions of coincident Radio‐Frequency (RF) emitters. We then use machine learning methods to predict the mean LMA source altitudes matched to GLM groups using metrics from the optical data that describe the amplitude, breadth, and texture of the group spatial energy distribution. The resulting model can predict the LMA mean source altitude from GLM group data with a median absolute error of <1.5 km, which is sufficient to determine the location of the charge layer where the optical energy originated. This model is able to capture changes to the source altitude distribution in response to convective processes in the thunderstorm, and the GLM predictions can reveal the vertical structure of individual flashes ‐ enabling 3‐D flash geolocation with GLM for the first time. Future work will account for differences in thunderstorm charge/precipitation structures and viewing angle across the GLM Field of View.

54 ENVIRONMENTAL SCIENCES↗

Study of Classifiers for U-235 Source Signatures Using Gamma Spectral Measurements

Signatures associated with low-level U-235 sources are studied from a classification analytics perspective, using NaI gamma-ray spectral measurements from detectors located at various distances from the source. Data sets collected at a shielded facility are utilized, wherein the source is introduced via a conduit into a formation of 21 NaI detectors deployed over 6 x 6 meters area in a formation of two concentric circles and one spiral. The activity levels in the spectral regions associated with potential U-235 signatures are estimated as counts at 1 second intervals, and are used as features to train classifiers for detecting the presence of the source. Eight different classifiers are trained and tested using the background and source measurements collected over multiple experimental runs. As expected, the classifier performance improved overall as measurements from the detectors closer to source are used, but also revealed unexpectedly low performance by two detectors that are identically produced and configured as others. Six of eight classifiers have an overall comparable performance, for example, three of them achieved zero training error and 99% detection at 4% false alarm rate for a detector located 1.3 meters away from the source. Also, larger training sets led to improved classification performance across all classifiers, and interestingly, the classifiers with the minimum training error did not necessarily achieve the highest classification performance on test data.

Rao, Nageswara↗

Diffuse X-ray scattering from polished silicon: application of the distorted wave Born approximation

Measured diffuse X-ray scattering data for a `smooth' as well as for a `rough' silicon sample were fit to theoretical expressions within the distorted wave Born approximation (DWBA). Data for the power spectral density (PSD) for both samples were also obtained by means of atomic force microscopy and optical interferometry. The Fourier transforms of trial correlation functions were fit to the PSD data and then applied to the DWBA formalism. The net correlation functions needed to fit the PSD data for each sample comprised the sum of two terms with different cutoff lengths and different self-affine fractal exponents. At zero distance these correlation functions added up to yield net values of σ 2 = (2) 2 and (71) 2 Å 2 for the smooth and rough samples, respectively. X-ray scattering data were obtained at beamline 1-BM of the Advanced Photon Source. Data and fits at values of q z = 0.05 and 0.10 Å −1 for the smooth sample are reported. Good fits for the smooth sample were obtained at both q z values simultaneously, that is, identical fitting parameters were applied at both values of q z . The smooth sample also exhibited weak Yoneda wings and a clear distinction between the strong specular scattering and the weak diffuse scattering. Data for the rough sample were qualitatively different and exhibited very weak scattering at the specular condition in contrast to extremely large Yoneda wings. Fits for the rough sample are reported for q z = 0.04, 0.05, and 0.06 Å −1 . Although the large Yoneda wings could be fit quite well in both position and amplitude, scattering near the specular condition could not be equally well fit by applying the same fitting parameters at all values of q z . Albeit imperfect, best-fitting results at the specular condition were obtained by invoking only diffuse scattering, that is, without including a separate theoretical expression for specular scattering.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Grafana Plugin: Chord Diagram (esnet-chord-diagram) v1

The Esnet chord diagram is a plugin to the open source Grafana platform which lets users create chord diagrams based on data queried from configured data sources. Its designed as a general purpose chord visualization, which we often use internally to show quantitative relationships between network customers. The software itself uses the d3.js library to generate visuals and its primary value is making it easy to use these within Grafana. Graphing library Here: https://github.com/d3/d3-chord ISC License Grafana https://github.com/grafana/grafana AGPL3.0

Balas, Edward↗

Integrating Applied Energy and BER Smart Data Capabilities to Develop a DOE Data Fabric for Energy-Water R&D

Focal Area(s): 1) Data acquisition and assimilation enabled by machine learning, AI, and advanced methods including experimental/network design/optimization, unsupervised learning (including deep learning), and hardware-related efforts involving AI (e.g., edge computing). Science Challenge: DOE R&D, including DOE’s Basic Energy Research (BER)’s Environmental Systems Science Division (EESSD) program and DOE’s applied energy research (AER) programs (EERE, FE, and NE) are producers and consumers of Earth systems datasets. This white paper focuses on the first topic area from the call in relation to how crosscutting resources and innovations from DOE’s EESSD and AER can be brought to bear to mutual benefit and more efficient energy-water, Earth system data resources through improved. The overarching challenge posed by this call focuses on how DOE can directly leverage artificial intelligence (AI) to engineer a substantial (paradigm-changing) improvement in Earth System Predictability? While stemming from DOE BER’s EESSD program, this is a challenge that is faced and also being addressed by DOE’s AER programs. Over the past decade plus, FE, EERE, and NE programs have made important strides towards addressing this need. These strides are in many ways highly complementary to EESSD’s MODEX efforts. Energy water systems spanning metocean to groundwater to surface water systems all are data driven whether for basic energy or applied energy. These are remote, multi-variate, complex natural, and in many cases engineered, systems. Key needs and challenges of both EESSD and AER include developing data-focused tools to enhance data search and discovery to fill in knowledge gaps (address sparse data challenge), and rapidly transform datasets, including disparate and multi-source data. Leveraging DOE on-premise computing (HPC, exascale) infrastructure supports the computing-intensive algorithms required to execute these data acquisition and transformation processes to derive enriched knowledge and data, driving AI/ML and big data analytics for these systems. The opportunity lies in combining BER and AER efforts to provide a more robust, advanced, efficient and complete computing data fabric to address energy-water data acquisition and assimilation needs which currently pose significant impediments to AI/ML predictions and research.

54 ENVIRONMENTAL SCIENCES↗

Application of the DNAS framework expansion to occupant population synthesis

Research in occupant behaviour is now using a more elaborate framework of building occupant interaction. Researchers often face challenges in collecting data, particularly for the data to meet the minimum number of required data points and the data interoperability requirements. Researchers address the first issue with the synthetic population and the latter with data ontologies. While synthetic population is commonly used to address the first issue, data ontology development is used to address the latter. The two solutions are complementary to each other. One of the known ontologies in building occupant behaviour research is the Drivers-Needs-Actions-Systems (DNAS) ontology, which has been used by building modelers to describe energy-related occupant behaviour. This paper describes the ontology-based synthetic population generation that can be used in the agent-based modeling (ABM) applications. This paper considers multiple data sources, including ASHRAE Thermal Comfort DB II and IEA Annex 66 data sets. A case study of an office building is used to present the workflow of DNAS framework expansion, synthetic population generation, and agent-based modeling.

Chandra Putra, Handi↗

Natural Language Processing-Enhanced Nuclear Industry Operating Experience Data Analysis to Support Risk Model Parameter Estimations

This set of slides has been prepared for a talk at the INL AI/ML Symposium held on September 8, 2022. Presentation outline: Background - Nuclear power plant operating experience data sources • Research focus and motivation - Analyzing free-text operating experience data: present and future • Research method - Input - Methodological steps - Output • Conclusions and next steps

99 GENERAL AND MISCELLANEOUS↗

Efficient loading of reduced data ensembles produced at ORNL SNS/HFIR neutron time-of-flight facilities

We present algorithmic improvements to the loading operations of certain reduced data ensembles produced from neutron scattering experiments at Oak Ridge National Laboratory (ORNL) facilities. Ensembles from multiple measurements are required to cover a wide range of the phase space of a sample material of interest. They are stored using the standard NeXus schema on individual HDF5 files. This makes it a scalability challenge, as the number of experiments stored increases in a single ensemble file. The present work follows up on our previous efforts on data management algorithms, to address identified input output (I/O) bottlenecks in Mantid, an open-source data analysis framework used across several neutron science facilities around the world. We reuse an in-memory binary-tree metadata index that resembles data access patterns, to provide a scalable search and extraction mechanism. In addition, several memory operations are refactored and optimized for the current common use cases, ranging most frequently from 10 to 180, and up to 360 separate measurement configurations. Results from this work show consistent speed ups in wall-clock time on the Mantid LoadMD routine, ranging from 19% to 23% on average, on ORNL production computing systems. The latter depends on the complexity of the targeted instrument-specific data and the system I/O and compute variability for the shared computational resources available to users of ORNL’s Spallation Neutron Source (SNS) and the High Flux Isotope Reactor (HFIR) instruments. Nevertheless, we continue to highlight the need for more research to address reduction challenges as experimental data volumes, user time and processing costs increase.

Godoy, William↗

Data for "Plasmon-driven exciton formation in a non-equilibrium Fermi liquid"

This repository contains source data for key plots presented in the manuscript "Plasmon-driven exciton formation in a non-equilibrium Fermi liquid." Experimental data that was analyzed in Igor Pro 8 are presented as the .pxp files used to generate individual sub-plots. Electronic spectral function calculations are provided as .txt files, in which consecutive rows refer to the meshgrid x coordinate, y coordinate, spectral function (and, where relevant, axis-projected local angular momentum). We additionally include the Wannier model and DFT-obtained bulk band structure on which the Wannier model was based. Files are named as the number of the figure in the manuscript to which they correspond, with additional details included where necessary. Details of file names: 2a_DOS_Lxz_Ek_KGM_40layer_xnum_800kpt_tot.txt: Density of states, xz-axis projected local orbital angular momentum, for 800 points along the K-Gamma-M path, for a 40-layer model. 2c_composite_y.pxp: ARPES (angle-resolved photoemission spectroscopy) spectra along the ky axis, including both a scan near the Fermi level and a scan at high kinetic energies. 2d_LCP_RCP_diff_Sect_20K.pxp: difference between ARPES constant energy cuts at T=20 K at E0 + 0.23 eV taken with left- and right-circularly polarized photons. The polarization-integrated intensity at the constant energy cut is also included. 2e_DOS_L45_E11pt79_m0pt25to0pt25_xnum_800kpt_tot.txt: Density of states, xz-projected local orbital angular momentum, and corresponding k-points in two dimensions from ab-initio electronic structure calculations for a constant-energy cut. 3a_[x]_[y]ps: ARPES cut under excitation at a fluence of x uJ/cm2, measured y ps after photoexcitation. Measurements were performed at 9 K. 3b_[x]: Energy distribution curves under excitation at a fluence x uJ/cm2 at selected delay times after photoexcitation. 4a_ImSigma_vs_temperature.pxp: Imaginary self energy (extracted from ARPES linewidths) at different energies above E0 for selected lattice temperatures. 4b_EELS_lowE.pxp: Electron energy loss spectrum over a low energy range 5b_diff_55m15.pxp: Difference between momentum-integrated Tr-ARPES traces at 55 uJ/cm2 and 15 uJ/cm2 photoexcitation. Time-dependent intensity at each energy level has been normalized to a maximum of 1 for each individual fluence prior to subtraction. 5d_invtau_at_EX_vs_fluence.pxp: decay rate at a specified energy EX for different excitation fluences, from single exponential fits. NOTE: Analyses based on the Wannier model presented here should cite both the associated Article and this dataset. For all other files in the repository, citing the dataset alone is sufficient.

Acharya, Rishi [University of Illinois] (ORCID:000↗

Entropy removal of medical diagnostics

Shannon entropy is a core concept in machine learning and information theory, particularly in decision tree modeling. To date, no studies have extensively and quantitatively applied Shannon entropy in a systematic way to quantify the entropy of clinical situations using diagnostic variables (true and false positives and negatives, respectively). Decision tree representations of medical decision-making tools can be generated using diagnostic variables found in literature and entropy removal can be calculated for these tools. This concept of clinical entropy removal has significant potential for further use to bring forth healthcare innovation, such as quantifying the impact of clinical guidelines and value of care and applications to Emergency Medicine scenarios where diagnostic accuracy in a limited time window is paramount. This analysis was done for 623 diagnostic tools and provided unique insights into their utility. For studies that provided detailed data on medical decision-making algorithms, bootstrapped datasets were generated from source data to perform comprehensive machine learning analysis on these algorithms and their constituent steps, which revealed a novel and thorough evaluation of medical diagnostic algorithms.

97 MATHEMATICS AND COMPUTING↗

Model fusion with physics-guided machine learning: Projection-based reduced-order modeling

The unprecedented amount of data generated from experiments, field observations, and large-scale numerical simulations at a wide range of spatiotemporal scales has enabled the rapid advancement of data-driven and especially deep learning models in the field of fluid mechanics. Although these methods are proven successful for many applications, there is a grand challenge of improving their generalizability. This is particularly essential when data-driven models are employed within outer-loop applications like optimization. In this work, we put forth a physics-guided machine learning (PGML) framework that leverages the interpretable physics-based model with a deep learning model. Leveraging a concatenated neural network design from multi-modal data sources, the PGML framework is capable of enhancing the generalizability of data-driven models and effectively protects against or inform about the inaccurate predictions resulting from extrapolation. We apply the PGML framework as a novel model fusion approach combining the physics-based Galerkin projection model and long- to short-term memory (LSTM) network for parametric model order reduction of fluid flows. We demonstrate the improved generalizability of the PGML framework against a purely data-driven approach through the injection of physics features into intermediate LSTM layers. Our quantitative analysis shows that the overall model uncertainty can be reduced through the PGML approach, especially for test data coming from a distribution different than the training data. Moreover, we demonstrate that our approach can be used as an inverse diagnostic tool providing a confidence score associated with models and observations. The proposed framework also allows for multi-fidelity computing by making use of low-fidelity models in the online deployment of quantified data-driven models.

42 ENGINEERING↗

Virtual Log-Structured Storage for High-Performance Streaming

Over the past decade, given the higher number of data sources (e.g., Cloud applications, Internet of things) and critical business demands, Big Data transitioned from batch-oriented to real-time analytics. Stream storage systems, such as Apache Kafka, are well known for their increasing role in real-time Big Data analytics. For scalable stream data ingestion and processing, they logically split a data stream topic into multiple partitions. Stream storage systems keep multiple data stream copies to protect against data loss while implementing a stream partition as a replicated log. This architectural choice enables simplified development while trading cluster size with performance and the number of streams optimally managed. This paper introduces a shared virtual log-structured storage approach for improving the cluster throughput when multiple producers and consumers write and consume in parallel data streams. Stream partitions are associated with shared replicated virtual logs transparently to the user, effectively separating the implementation of stream partitioning (and data ordering) from data replication (and durability). We implement the virtual log technique in the KerA stream storage system. When comparing with Apache Kafka, KerA improves the cluster ingestion throughput by up to 4x when multiple producers write over hundreds of data streams.

consistent stream ordering↗

Global analysis of the yeast knockout phenome

Genome-wide phenotypic screens in the budding yeast Saccharomyces cerevisiae, enabled by its knockout collection, have produced the largest, richest, and most systematic phenotypic description of any organism. However, integrative analyses of this rich data source have been virtually impossible because of the lack of a central data repository and consistent metadata annotations. Here, we describe the aggregation, harmonization, and analysis of ~14,500 yeast knockout screens, which we call Yeast Phenome. Using this unique dataset, we characterized two unknown genes (YHR045W and YGL117W) and showed that tryptophan starvation is a by-product of many chemical treatments. Furthermore, we uncovered an exponential relationship between phenotypic similarity and intergenic distance, which suggests that gene positions in both yeast and human genomes are optimized for function.

59 BASIC BIOLOGICAL SCIENCES↗

AMIA KDDM Working Group Collaborative Workshop: Enriching Electronic Health Records with Social Determinants of Health to Improve Outcomes and Health Equity

Prior research has demonstrated that social determinants of health (SDoH) are major drivers of health outcomes and contributors to widespread health inequities. It was estimated that, in the United States, SDoH could be responsible for up to 40% of all preventable deaths, significantly higher than the 10-15% for which better medical care is responsible. Public health interventions that target SDoH are instrumental for improving health outcomes and reducing long-standing health inequities. Currently, most mainstream EHR vendors have implemented SDoH screeners in their EHR systems. However, the utility of the screeners is low, rendering patient-level SDoH still widely unavailable in the structured fields. SDoH are sometimes mentioned in free-text clinical notes (e.g., social context section) where natural language processing (NLP) can be applied to extract relevant information. Contextual-level SDoH can be identified from multiple data sources, many of which are publicly available and spatiotemporally linked to EHR data. As such, there is an opportunity for the KDDM research community to create innovative solutions to draw meaningful insights by creating and using rich data with SDoH to improve health outcomes while reducing disparities. In this workshop organized by AMIA Knowledge Discovery and Data Mining Working Group (AMIA KDDM WG), we will invite world-leading experts from academia, national laboratories, and life science industry with varied backgrounds in biomedical informatics, epidemiology, data science, machine learning, natural language processing, and pediatric cardiology to discuss the best practice of capturing, standardizing, and using SDoH information in various applications aiming at improving outcomes and health equity.

He, Zhe↗

Hestia-SWIFL: hourly anthropogenic fossil fuel CO2 and heat on the 2km WRF grid, version 1.1

The Hestia-SWIFL version 1.1 anthropogenic heat (AH) and fossil fuel CO2 (FFCO2) emissions data product represent emissions due to the combustion of fossil fuel and cement production within the state of Arizona from 2019 to 2022. This product was developed as part of the Southwest Urban Corridor Integrated Field Laboratory (SW-IFL) project, which aims to provide new knowledge and tools that address extreme heat, air quality, climate change and related urban environmental issues by integrating high-resolution observations, modeling, and civic engagement. The emissions are generated using a bottom-up/engineering approach and are tied to results generated by the Vulcan Project version 4, an effort to quantify space/time-resolved FFCO2 & AH emissions for the entire United States landscape. A large number of data sources are combined to best estimate the emissions at fine scales such as air quality emissions data, traffic flow data, building information, sociodemographic information, and fuel statistics. The AH product provides emissions for two emissions sources (transportation and point source emissions) in units of Watts per hour per square meter (W/m2) per year (annual files) or per hour (hourly files). The FFCO2 product provides emissions from nine individual emission sectors as well as the total, and in units of tons of carbon (tC) per grid cell per year or per hour. The output made available here places the native spatial resolution of the Hestia FFCO2 & AH emissions data product (points, lines, and polygons) into a regularized 2km x 2km grid at hourly and annual temporal resolutions, and stored in netCDF files. The exact spatial extent is defined by the ASU Weather Research Forecast (WRF) simulation grid. All data are processed using R/Python pm high-performance computing system. 2-27-2026 updates: Bugs in airport hourly profile (both AH and FFCO2) and building spatial patterns (FFCO2 only) were fixed. Hourly emissions are reprocessed for all years to reflect those changes.

54 ENVIRONMENTAL SCIENCES↗

Generation and representation of synthetic smart meter data

Advanced energy algorithms running at big-data scale will be necessary to identify, realize, and verify energy savings to meet government and utility goals of building energy efficiency. Any algorithm must be well characterized and validated before it is trusted to run at these scales. Smart meter data from real buildings will ultimately be required for the development, testing, and validation of these energy algorithms and processes. However, for initial development and testing, smart meter data are difficult to work with due to privacy restrictions, noise from unknown sources, data accessibility, and other concerns which can complicate algorithm development and validation. This paper describes a new methodology to generate synthetic smart meter data of electricity use in buildings using detailed building energy modeling, which aims to capture the variability and stochastics of real energy use in buildings. The methodology can create datasets tailored to represent specific scenarios with known truth and controllable amounts of synthetic noise. Knowledge of ground truth also allows the development and validation of enhanced processes which leverage building metadata, such as building type or size (floor area), in addition to smart meter data. The methodology described in this paper includes the key influencing factors of real-world building energy use including weather data, occupant-driven loads, building operation and maintenance practices, and special events. Data formats to support workflows leveraging both synthetic meter data and associated metadata are proposed and discussed. Finally, example use cases of the synthetic meter data are described to illustrate potential applications.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗