Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “scientific data change analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Transforming Energy Through Sustainable Mobility: Expanding Low-Carbon Transportation R&D Solutions

As the nation's premier laboratory for cutting-edge transportation decarbonization research and development solutions, the National Renewable Energy Laboratory (NREL) pioneers the creation and deployment of sustainable mobility technologies and strategies, with a focus on slashing transportation sector greenhouse gas emissions and combatting climate change. Trucks, planes, cargo ships, and other difficult-to-decarbonize vehicles are part of this essential transition for the transportation sector, which is currently the nation's largest source of the greenhouse gas emissions. NREL provides the scientific building blocks needed to spur innovation through multifaceted analysis, research, and engineering. This work acts as a catalyst to help industry bring affordable, high-performance, energy-efficient, and low-emission modes of transport and related infrastructure to market sooner. Our researchers collaborate closely with academic, government, and industry partners to design better batteries, drivetrains, and engines. They develop technologies for high-power charging, thermal management, energy storage, and power electronics. They are also reimagining fuels and combustion while creating sustainable lightweight materials. Unbiased expert guidance - backed by credible data and analysis, tools, and scientific rigor - empowers partners to make informed decisions about sustainable transportation. NREL recognizes that communities with limited mobility options face reduced access to employment opportunities, health care, and education, lowering overall quality of life. Alongside partners, NREL experts are creating transportation solutions that meet community-identified needs and increase mobility equity in historically underserved and overburdened communities. Rather than providing a one-size-fits-all solution, we take an interdisciplinary approach to mobility equity that considers the needs and challenges of diverse groups, maximizing benefits at the individual, community, and societal levels.

ADVANCED PROPULSION SYSTEMS↗

Decision Science for Machine Learning (DeSciML)

The increasing use of machine learning (ML) models to support high-consequence decision making drives a need to increase the rigor of ML-based decision making. Critical problems ranging from climate change to nonproliferation monitoring rely on machine learning for aspects of their analyses. Likewise, future technologies, such as incorporation of data-driven methods into the stockpile surveillance and predictive failure analysis for weapons components, will all rely on decision-making that incorporates the output of machine learning models. In this project, our main focus was the development of decision scientific methods that combine uncertainty estimates for machine learning predictions, with a domain-specific model of error costs. Other focus areas include uncertainty measurement in ML predictions, designing decision rules using multiobjecive optimization, the value of uncertainty reduction, and decision-tailored uncertainty quantification for probability estimates. By laying foundations for rigorous decision making based on the predictions of machine learning models, these approaches are directly relevant to every national security mission that applies, or will apply, machine learning to data, most of which entail some decision context.

97 MATHEMATICS AND COMPUTING↗

Enabling Innovative Analysis on Heterogeneous Clusters through HTCdaskgateway

High energy particle (HEP) physics research is going through fundamental changes as we move to collect larger amounts of data from the Large Hadron Collider (LHC). Analysis facilities and distributed computing, through HTCs, have come together to create the next pythonic generation of analysis by utilizing HTCdaskgateway, a Dask gateway extension, allowing users to spawn workers compatible with both their analysis and heterogeneous clusters in line with authentication requirements. This is enabling physicists to engage with scientific python in ways they had not before because of domain specific C++ tools. An example of HTCdaskgateway’s use is Fermilab’s Elastic Analysis Facility.

Chavez, Elise [U. Wisconsin, Madison (main)]↗

Volumetric Rendering on Wavelet-Based Adaptive Grid

Numerical modeling of physical phenomena frequently involves processes across a wide range of spatial and temporal scales. In the last two decades, the advancements in wavelet-based numerical methodologies to solve partial differential equations, combined with the unique properties of wavelet analysis to resolve localized structures of the solution on dynamically adaptive computational meshes, make it feasible to perform large-scale numerical simulations of a variety of physical systems on a dynamically adaptive computational mesh that changes both in space and time. Volumetric visualization of the solution is an essential part of scientific computing, yet the existing volumetric visualization techniques do not take full advantage of multi-resolution wavelet analysis and are not fully tailored for visualization of a compressed solution on the wavelet-based adaptive computational mesh. Our objective is to explore the alternatives for the visualization of time-dependent data on space-time varying adaptive mesh using volume rendering while capitalizing on the available sparse data representation. Two alternative formulations are explored. The first one is based on volumetric ray casting of multi-scale datasets in wavelet space. Rather than working with the wavelets at the finest possible resolution, a partial inverse wavelet transform is performed as a preprocessing step to obtain scaling functions on a uniform grid at a user-prescribed resolution. As a result, a solution in physical space is represented by a superposition of scaling functions on a coarse regular grid and wavelets on an adaptive mesh. An efficient and accurate ray casting algorithm is based just on these coarse scaling functions. Additional details are added during the ray tracing by taking an appropriate number of wavelets into account based on support overlap with the interpolation point, wavelet coefficient magnitude, and other characteristics, such as opacity accumulation (front to back ordering) and deviation from frontal viewing direction. The second approach is based on complementing of wavelet-based adaptive mesh to the traditional Adaptive Mesh Refinement (AMR) mesh. Both algorithms are illustrated and compared to the existing volume visualization software for Rayleigh-Benard thermal convection and electron density data sets in terms of rendering time and visual quality for different data compression of both wavelet-based and AMR adaptive meshes.

Vezolainen, Alexei V.↗

TR-XPS Realtime Analysis Tool (ArroyoXPS) v0.1

The ALS has developed a Time-Resolved X-ray Photoelectron Spectroscopy (TR-XPS) technique, which involves applying a specific pattern of voltage curves to a sample while measuring XPS peaks. This pattern is repeated over multiple cycles, and changes in the material's response provide valuable scientific insights. Traditionally, file-based analysis workflows have been used: scans are run for a predetermined time, and after one or more scans are complete, calculations are made. ArroyoXPS changes this by offering in-experiment scan and analysis, allowing researchers to gain insights before a scan is finished. This enables them to adjust experimental parameters quickly, potentially saving valuable beamtime. ArroyoXPS includes tools for integrating with beamline control systems, performing analysis, and visualizing scan data in a web browser.

McReynolds, Dylan [Lawrence Berkeley National Labo↗

A Cast of Thousands: How the IDEAS Productivity Project Has Advanced Software Productivity and Sustainability

Computational and data-enabled science and engineering are revolutionizing advances throughout science and society, at all scales of computing. For example, teams in the U.S. Department of Energy’s Exascale Computing Project have been tackling new frontiers in modeling, simulation, and analysis by exploiting unprecedented exascale computing capabilities—building an advanced software ecosystem that supports next-generation applications and addresses disruptive changes in computer architectures. However, concerns are growing about the productivity of the developers of scientific software. Members of the Interoperable Design of Extreme-scale Application Software project serve as catalysts to address these challenges through fostering software communities, incubating and curating methodologies and resources, and disseminating knowledge to advance developer productivity and software sustainability. This article discusses how these synergistic activities are advancing scientific discovery—mitigating technical risks by building a firmer foundation for reproducible, sustainable science at all scales of computing, from laptops to clusters to exascale and beyond.

97 MATHEMATICS AND COMPUTING↗

ThunderSecure: deploying real-time intrusion detection for 100G research networks by leveraging stream-based features and one-class classification network

Nowadays, data generated by large-scale scientific experiments are on the scale of petabytes per month. These data are transferred through dedicated high-bandwidth networks (40/100G) across distributed sites for processing, storage, and analysis. Like general purpose networks, research networks experience intrusions. However, monitoring anomalies in such high-speed network traffics is challenging given current cyber-infrastructure. Moreover, traditional network intrusion detection systems (NIDS) are signature based. However, anomaly patterns are difficult to define and that rulesets are often not updated frequently enough to reflect the changes of attack behaviors. We present ThunderSecure, a high-throughput, unsupervised learning-based intrusions detection system for 100G research networks. ThunderSecure implements an efficient packet processing and detection pipeline using multi-cores and GPUs. It extracts statistical and temporal features from real-time network data streams and feeds them to a one-class anomaly detection network. A baseline of normal distribution will be created based on the training observation. Testing traffic deviated from the learned profile will be marked as anomalies. We trained ThunderSecure on hundreds of billions of science data packets mirrored from two 100G network connections at Fermi National Accelerator Laboratory. The detection performance was evaluated on traffic captured from the same research network days and weeks after the training with different types of attack flows injected. Results show that ThunderSecure can recognize science data traffic captured long after the training and made nearly certain detection on the segment of the streams where anomalous flows were injected.

100G research network↗

Imaging Bragg Edge Analysis TooLs for Engineering Structures (iBeatles)

The Spallation Neutron Source (SNS) at Oak Ridge National Laboratory (ORNL) provides pulsed neutrons with energies varying from epithermal to cold. In preparation for VENUS, the neutron imaging beamline to be located at beam port 10, we have performed a series of experiments focused on wavelength-dependent radiography and computed tomography for a broad range of applications, from materials science to biological tissues.One of the time-of-flight (TOF) techniques that is of interest to the scientific community is the 2-dimensional mapping of phases and average crystalline plane orientation in samples both ex-situ and during applied stresses such as tensile loading and heating. This technique is known as Bragg edgeimaging and relies on the identification of changes of transmission values, fitting of the edge to measure its displacement, and thus identify the shift in lattice parameter due to stresses. One of the challenges of TOF imaging measurements is the amount of data and the inability to observe Bragg edge shifts in real time during an experiment. Thus, we have been focusing on creating a Python-based interface that allows fast data processing and instantaneous mapping and fitting of the Bragg edges, and their evolution through time. Python libraries and Jupyter notebooks have been implemented to facilitate decision making during an experiment. The advantage of the notebooks is the possibility to guide an experiment as they can quickly process and display Bragg edge data. These notebooks can be used independently, or can be combined in a Python Graphical User Interface (GUI) tool called iBeatles. This interface permits visualization and fitting of the Bragg edges, and ultimately back-projects the fitting results onto the radiographs to display a strain map. Assuming data collection has sufficient statistics, the strain mapping analysis can be performed on a pixel-by-pixel basis. This development is a step forward toward a better user experience at the future VENUS beamline in terms of live feedback and productivity. Analysis that used to take days of switching between different applications can now be done in minutes within the

Bilheux, JeanChristophe [Oak Ridge National Labora↗

Optimization of distributed compute resources utilization in the CMS Global Pool

The CMS Submission Infrastructure is the primary system for managing computing resources for CMS workflows, including data processing, simulation, and analysis. It integrates geographically distributed resources from Grid, HPC, and cloud providers into federated pools managed by HTCondor and Glidein- WMS, for a total of around 500k CPU cores. This system dynamically manages workloads based on priorities defined by the collaboration. Additionally, CMS scheduling strategies must be flexible to handle multiple concurrent workloads while considering changing processing demands and resource availability from various providers.Efficient utilization of vast amounts of distributed compute resources is a key element for the success of the scientific programs of the LHC experiments. Optimizing the system is essential to maximize resource efficiency and fully utilize the distributed computing power. The CMS Submission Infrastructure team thus systematically investigates sources of inefficiency in workload scheduling to reduce their impact. In addition, a strategy of pilot overloading has been introduced to compensate for other inefficiency sources, thereby optimizing resource utilization and enhancing computational throughput.

Mascheroni, Marco [UC, San Diego (main)]↗

Validation of LOCA2 and STAR-ESDM Statistically Downscaled Products

The National Climate Assessment (NCA) is the preeminent national report examining current and future risks posed by climate change. Countless agencies, policymakers, stakeholders and other end-users rely upon guidance from the NCA to plan for an uncertain future. These groups all depend on modern curated data, provided alongside the NCA, to quantify the impact of climate change on metrics of relevance for their decision processes. In its fifth iteration (NCA5), two statistically downscaled ensemble products, each providing data at grid spacing of approximately 5km over the contiguous United States, were selected to accompany the report. These include LOCalized Analogs version 2 (LOCA2) and Seasonal Trends and Analysis of Residuals Empirical-Statistical Downscaling Model (STAR-ESDM). Both data products are produced through a process known as statistical downscaling, where relatively coarse Global Climate Model (GCM) data is refined to locally relevant scales through the application of scientifically-supported empirical and algorithmic relationships. In support of the NCA effort, this report provides an independent validation of these two products against historical observations, with a focus on precipitation and near-surface temperature variables. Based on the results of this validation, several recommendations are provided related to the use of these data products. The structure of this report is as follows: In section 2, we review three gridded observational products that are used as part of our intercomparison. In section 3, we describe the two statistical downscaling techniques and their corresponding datasets that are the focus of this study. In section 4, the methodology we employ for validation is described. Section 5 provides results of the validation, which in turn motivate our recommendations on the use of these data products. A brief summary is provided in section 6.

54 ENVIRONMENTAL SCIENCES↗

Examining Nuisance Aerosol Detections in Light of the Origin of the Screening Process (November 2021)

The evolution of philosophy and computations in the International Data Center (IDC) related to aerosol samples have had profound impacts on the number of recorded detections in the network since routine operations began in 2000. Key decisions from policymakers have been the list of triggering radionuclides, the scheme for categorizing these into interest levels 1-5, and an algorithm for determining when an anthropogenic isotope is seen so often that it is no longer interesting, known as the Exponential Weighted Moving Average (EWMA). These are described in the Operations Manual of the IDC. Key parameters that are controlled by the IDC but for which the IDC receives occasional input from policymakers include the constants in EWMA and the peak significance threshold for individual gamma rays, the latter of which directly leads to determination of the presence or absence of a radionuclide in a sample. There are also changes in computations which the IDC makes and informs policy makers about, such as changes in how background is computed, which could also affect the ease of detecting a peak – real or false. Rather than focus on the quantitative changes due to computation changes, this work records some thinking on how isotopes and peak significance levels were chosen, and the resulting detections seen over 18 years during the buildup of the International Monitoring System (IMS). These detections are considered on a global scale to try to determine the relative impact on monitoring, and in some cases, the nature of their existence. Repeated detections of 131I and 133I are the most troublesome, but they are not so frequent to be a major problem for the Verification Regime. These detections could probably be handled adequately using scientific methods currently under development for xenon backgrounds. It is also somewhat problematic that top-level analysis of aerosol backgrounds has not been reported previous to this. The steep increase in the rate of detections after 2016 are a concern, either in the actual backgrounds or from changes in the calculations methods used to generate the Reviewed Radionuclide Report (RRR.) Final conclusions of the authors are that the computational stability of the RRR is very important. With computational stability, changes can be usefully analyzed as being due to changes in radioactivity in Earth’s atmosphere This report is a distillation into text of a talk given in the Radionuclide Experts Group (RNEG) in Vienna during Working Group B (WGB) in February of 2019. This report does not directly contain any IDC data, only summaries by year, or by isotope, or by location. No specific IDC detection by time, location, or isotope is included.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Geographic Shift and Environment Change of U.S. Tornado Activities in a Warming Climate

Even with ever-increasing societal interest in tornado activities engendering catastrophes of loss of life and property damage, the long-term change in the geographic location and environment of tornado activity centers over the last six decades (1954–2018), and its relationship with climate warming in the U.S., is still unknown or not robustly proved scientifically. Utilizing discriminant analysis, we show a statistically significant geographic shift of U.S. tornado activity center (i.e., Tornado Alley) under warming conditions, and we identify five major areas of tornado activity in the new Tornado Alley that were not identified previously. By contrasting warm versus cold years, we demonstrate that the shift of relative warm centers is coupled with the shifts in low pressure and tornado activity centers. The warm and moist air carried by low-level flow from the Gulf of Mexico combined with upward motion acts to fuel convection over the tornado activity centers. Employing composite analyses using high resolution reanalysis data, we further demonstrate that high tornado activities in the U.S. are associated with stronger cyclonic circulation and baroclinicity than low tornado activities, and the high tornado activities are coupled with stronger low-level wind shear, stronger upward motion, and higher convective available potential energy (CAPE) than low tornado activities. The composite differences between high-event and low-event years of tornado activity are identified for the first time in terms of wind shear, upward motion, CAPE, cyclonic circulation and baroclinicity, although some of these environmental variables favorable for tornado development have been discussed in previous studies.

54 ENVIRONMENTAL SCIENCES↗

Commitment to Active Allyship Is Required to Address the Lack of Hispanic and Latinx Representation in the Earth and Atmospheric Sciences

In 2021, people of Hispanic and Latinx origin made up 6% of the atmospheric and Earth sciences workforce of the United States, yet they represent 20% of the population. Motivated by this disparity in Hispanic and Latinx representation in the atmospheric and Earth science workforce, this manuscript documents the lack of representation through existing limited demographic data. The analysis presents a clear gap in participation by Hispanic and Latinx people in academic settings, with a widening gap through each education and career stage. Several factors and challenges impacting the representation disparity include the lack of funding for and collaboration with Hispanic-serving institutions, limited opportunities due to immigration status, and limited support for international research collaborations. We highlight the need for actionable steps to address the lack of representation and provide targeted recommendations to federal funding agencies, educational institutions, faculty, and potential employers. While we wait for systemic cultural change from our scientific institutions, grassroots initiatives like those proudly led by the AMS Committee for Hispanic and Latinx Advancement will emerge to address the needs of the Hispanic and Latinx scientific and broader community. We briefly highlight some of those achievements. Lasting cultural change can only happen if our leaders are active allies in the creation of a more diverse, equitable, and inclusive future. Alongside our active allies we will continue to champion for change in our weather, water, and climate enterprise.

99 GENERAL AND MISCELLANEOUS↗

MODE: A Web Application for Interactive Visualization and Exploration of Omics Data

Studies generating transcriptomics, proteomics, lipidomics, and metabolomics (colloquially referred to as “omics”) data allow researchers to find biomarkers or molecular targets, or understand complex biological structures and functions by identifying changes in biomolecule abundance and expression between experimental conditions. Omics data is multi-dimensional and oftentimes summarization techniques such as principal component analysis (PCA) are used to identify high-level patterns in data. Though useful, these summaries don’t allow exploration of detailed patterns in omics data that may have biological relevance. The use of interactive HTML displays with plots allows researchers to interact with omics data at a detailed level, but building these displays requires significant coding expertise. To overcome this barrier, the software MODE was built to empower users to build their own interactive HTML displays to support scientific discovery. These displays are easily shareable, do not depend on a specific operating system, and allow users to effortlessly sort and filter plots by categorical or numerical variables. MODE allows users to build and share these displays with several options for plot design and meta selection. In conclusion, the MODE web application and its capabilities are presented and then demonstrated on lipidomics data from a leaf wounding study.

lipidomics↗

Advanced Research Directions on AI for Science, Energy, and Security: Report on Summer 2022 Workshops

Over the past decade, fundamental changes in artificial intelligence (AI)—from foundational to applied—have delivered dramatic insights across a wide breadth of U.S. Department of Energy (DOE) mission space. AI is helping to augment and improve scientific and engineering workflows (e.g., for control, design, and dramatic performance gains through surrogate models) in national security, the Office of Science, and DOE’s applied energy programs. The progress and potential for AI in DOE science was captured in the 2020 “AI for Science” report from the DOE laboratory community in collaboration with academia and industry. Specific scientific areas ready to further leverage the power of AI ranged from the scale and performance of computational models to data analysis to creating new classes of observations using computer vision. Since that report, the scale and scope of scientific AI have accelerated, revealing new, emergent properties that yield insights that go beyond enabling opportunities to being potentially transformative in the way that scientific problems are posed and solved. Thus, under the guidance of both the Office of Science (SC) and the National Nuclear Security Administration (NNSA), the DOE national laboratories organized a series of workshops in 2022 to gather input on new and rapidly emerging opportunities and challenges of scientific AI. This 2023 report is a synthesis of those workshops. The scientific community believes AI can have a foundational impact on a broad range of DOE missions, including science, energy, and national security. Further, DOE has unique capabilities that enable the community to drive progress in scientific use of AI, building on long-standing DOE strengths and investments in computation, data, and communications infrastructure, spanning the Energy Sciences Network (ESnet), the Exascale Computing Project (ECP), and integrative programs such as the NNSA Office of Defense Programs Advanced Simulation and Computing (ASC) and the SC Scientific Discovery through Advanced Computing (SciDAC) programs.

97 MATHEMATICS AND COMPUTING↗

Adaptive elasticity policies for staging-based in situ visualization

In situ processing aims to alleviate the growing gap between computation and I/O capabilities by performing data processing close to the data source. In situ processing is widely used to process data generated by multiple data sources, including observation data from edge devices or scientific observational facilities and the simulation data generated by scientific computation on a high-performance computing (HPC) platform. For a scientific workflow that is run on an HPC platform and composed of a simulation program and an in situ data analytics or visualization (abbreviated as ana/vis) task, there is an implicit assumption that the computing resources assigned to the workflow keep static during the workflow execution. However, with the converging trend between the HPC and cloud computing platform, running the in situ ana/vis task in an elastic way is promising to decrease its overhead and improve its resource utilization rate. Resource elasticity represents the ability to change resource configurations such as the number of computing nodes/processes during workflow execution. An elastic job may dynamically adjust resource configurations; it may use a few resources at the beginning and more resources toward the end of the job when interesting data appear. However, it is hard to predict a priori how many computing nodes/processes need to be added/removed during the workflow execution to adapt to changing workflow needs. How to efficiently guide elasticity operations, such as growing or shrinking the number of processes used for in situ analysis during workflow execution, is an open-ended research question. In this article, we present adaptive elasticity policies that adopt workflow runtime information collected during workflow execution to predict how to trigger the addition/removal of processes in order to minimize in situ processing overhead. Taking in situ visualization tasks as an example, we integrate the presented elasticity policies into a staging-based elastic workflow and evaluate its efficiency in multiple elasticity scenarios. Compared with the situation without elasticity or with a static elasticity policy that uses a fixed number of processes for each rescaling operation, the adaptive elasticity policy can save overhead in finding a proper resource configuration and improve resource utilization efficiency. Furthermore, one experiment illustrates that the adaptive elasticity policy saves 41% of core-hours compared with the situation without the resource elasticity.

97 MATHEMATICS AND COMPUTING↗

A modERN resource: identification of Drosophila transcription factor candidate target genes using RNAi

Transcription factors (TFs) play a key role in development and in cellular responses to the environment by activating or repressing the transcription of target genes in precise spatial and temporal patterns. In order to develop a catalog of target genes of Drosophila melanogaster TFs, the modERN consortium systematically knocked down the expression of TFs using RNAi in whole embryos followed by RNA-seq. We generated data for 45 TFs which have 18 different DNA-binding domains and are expressed in 15 of the 16 organ systems. The range of inactivation of the targeted TFs by RNAi ranged from log2fold change -3.52 to +0.49. The TFs also showed remarkable heterogeneity in the numbers of candidate target genes identified, with some generating thousands of candidates and others only tens. We present detailed analysis from five experiments, including those for three TFs that have been the focus of previous functional studies (ERR, sens, and zfh2) and two previously uncharacterized TFs (sens-2 and CG32006), as well as short vignettes for selected additional experiments to illustrate the utility of this resource. The RNA-seq datasets are available through the ENCODE DCC (http://encodeproject.org) and the Sequence Read Archive (SRA). TF and target gene expression patterns can be found here: https://insitu.fruitfly.org. These studies provide data that facilitate scientific inquiries into the functions of individual TFs in key developmental, metabolic, defensive, and homeostatic regulatory pathways, as well as provide a broader perspective on how individual TFs work together in local networks during embryogenesis.

59 BASIC BIOLOGICAL SCIENCES↗

Efficient Asynchronous I/O with Request Merging

With the advancement of exascale computing, the amount of scientific data is increasing day by day. Efficient data access is necessary for scientific discoveries. Unfortunately, the I/O performance is not improved, like the CPU and network speed. So, I/O operations take longer time than data generation or analysis. Asynchronous I/O has been proposed to extenuate the I/O bottleneck by overlapping I/O and computation time. However, multiple small write operations can diminish the benefits of asynchronous I/O, as the I/O time becomes significantly longer than the compute time, with little time to overlap with. To overcome these issues, we present an optimization technique to merge small contiguous write operations. We integrated our solution into the HDF5 asynchronous I/O VOL connector and demonstrated the effectiveness of merging HDF5 write operations automatically and transparently without requiring any code change from the application.

Chowdhury, Kamal Hossain↗