Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data shift”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

A biology-informed similarity metric for simulated patches of human cell membrane

Complex scientific inquiries rely increasingly upon large and autonomous multiscale simulation campaigns, which fundamentally require similarity metrics to quantify ‘sufficient’ changes among data and/or configurations. However, subject matter experts are often unable to articulate similarity precisely or in terms of well-formulated definitions, especially when new hypotheses are to be explored, making it challenging to design a meaningful metric. Furthermore, the key to practical usefulness of such metrics to enable autonomous simulations lies in in situ inference, which requires generalization to possibly substantial distributional shifts in unseen, future data. Here, we address these challenges in a cancer biology application and develop a meaningful similarity metric for ‘patches’—regions of simulated human cell membrane that express interactions between certain proteins of interest and relevant lipids. In the absence of well-defined conditions for similarity, we leverage several biology-informed notions about data and the underlying simulations to impose inductive biases on our metric learning framework, resulting in a suitable similarity metric that also generalizes well to significant distributional shifts encountered during the deployment. We combine these intuitions to organize the learned embedding space in a multiscale manner, which makes the metric robust to incomplete and even contradictory intuitions. Our approach delivers a metric that not only performs well on the conditions used for its development and other relevant criteria, but also learns key spatiotemporal relationships without ever being exposed to any such information during training.

97 MATHEMATICS AND COMPUTING↗

Exploring OpenSNAPI Use Cases and Evolving Requirements [Slides]

Emerging system architectures are rapidly transforming in order to meet shifting requirements. Motivated by expanding data volumes, energy efficiency concerns, and the omnipresent need to improve performance, architectures are increasingly adopting a data-centric approach. At the core of this concept is the goal of minimizing data motion and instead processing data in-situ to the greatest degree possible. Therefore, data-centric designs, in contrast to conventional CPU-centric models, typically distribute compute capabilities throughout the architecture. As part of this paradigm shift, a novel class of devices known as data processing units (DPUs), alongside CPUs and GPUs, are quickly forming a third pillar of data-centric systems. These devices, which include smart network adapters and switches, seek to offload computation on data at the network edge as well as in-flight within the network fabric. The Open Smart Network API (OpenSNAPI) project seeks to develop a unified API for DPU devices. In our previous talks, we introduced the OpenSNAPI project and detailed our investigations regarding the viability of offloading compute intensive kernels to BlueField DPUs. In contrast, in this talk we detail our efforts to offload application-level file I/O to the DPU. We also discuss plans and early efforts to explore in-network compute capabilities. Finally, we describe our observations with respect to the evolving design of OpenSNAPI.

97 MATHEMATICS AND COMPUTING↗

Boundary-Aware Adversarial Learning Domain Adaption and Active Learning for Cross-Sensor Building Extraction

The use of convolutional neural networks (CNNs) for building extraction from remote sensing images has been widely studied and many public datasets have been made available for accelerating development of these CNN models. Yet adapting pretrained models at scale in real-world scenarios remains a challenging task. The main barrier is that certain new labels are still needed to compensate for domain shifting between the labeled data and new images that potentially cover new geographic locations or that are from a different sensor. In this article, we propose to add informatively labeled samples from a new image pool under the paradigm of active learning. To select the most useful samples based on model uncertainty, we first tackle the problem of uncalibrated uncertainty estimation due to distribution shifting by adapting feature extractors with boundary-based adversarial learning. Calibrated uncertainty is used as the query criterion in the active learning process, where the most uncertain samples are selected for annotation and included for model retraining. The proposed workflow was tested with three data pairs in which each workflow represents a scenario often encountered in real-world applications, including adapting pretrained models to new images collected with different sensors or to new geographic areas where appearances and types of buildings are very different. Compared to several baselines, including random sampling, temperature scaling (a well-known uncertainty calibration technique), different query strategies, and active domain adaptation methods, the proposed workflow shows that strategically querying a smaller set of samples for labeling achieves comparable or better building extraction performance. The proposed method reduces the number of labeled samples required to achieve sufficient model accuracy, thus significantly reducing hundreds of person-hours for labeled data creation. In addition, we include a few considerations when deploying this workflow in a GPU cluster that can be easily adapted to achieve operational building extraction model retraining.

97 MATHEMATICS AND COMPUTING↗

Integration of sensors through additive manufacturing leading to increased efficiencies of gas turbines for power generation and propulsion

To realize the full capability of additively manufactured components in complex energy systems, it is imperative to minimize early component failures during development phases and during operation. Traditional field feedback timelines and offline inspection protocols significantly reduce the design-manufacturing iteration times. To address this specific question, the project developed and demonstrated a method for the integration of sensors into complex components through additive manufacturing. The team used gas turbine engines as a platform, which meets the need of both power generation and propulsion and offer opportunities for cost reductions and efficiency increases. The innovation of this intelligent integration of sensors into complex components uniquely customized to address questions of integrity and durability for additively manufactured components. With real-time sensing data from additively manufactured components, turbine manufacturers will realize higher efficiencies, reduced component failures, and a 30-50% acceleration in product deployment of high efficiency gas turbine components due to a faster reduction in component risk assessment under actual operating conditions. This is a transformative shift towards a data-driven design and qualification of additively manufactured gas turbine components. To directly integrate sensors into additively manufactured components with all the complexities of actual hardware, powder bed fusion (direct metal laser sintering) and laser metal deposition technologies was developed. Validation took take place in two university laboratories both of which contain actual engine hardware and closely simulate a gas turbine prior to demonstrating the technology in a turbine development test. Indeed, two major technologies from this research cold impact turbine systems in the near future: (1) higher efficiency materials and designs enabled by additive manufacturing with 50% faster design to manufacturing cycle time, to enable faster time-to-market targets; and (2) integration of sensors into additively manufactured components enabling broad health and condition based prognostics for faster component and engine risk reduction.

33 ADVANCED PROPULSION SYSTEMS↗

Design of an 8-channel 40 GS/s 20 mW/Ch waveform sampling ASIC in 65 nm CMOS

One picosecond timing resolution is the entry point to signature based searches relying on secondary/tertiary vertices and particle identification. We describe PSEC5, an 8-channel 40 GS/s waveform-sampling ASIC in TSMC 65 nm process targetting one picosecond resolution at 20 mW power per channel. Each channel consists of four fast and one slow switched capacitor arrays (SCA), allowing for picosecond time resolution combined with a long effective buffer. Each fast SCA is 1.6 ns long and has a nominal sampling rate of 40 GS/s. The slow SCA is 204.8 ns long and samples at 5 GS/s. Recording of the analog data for each channel is triggered by a fast discriminator capable of multiple triggering during the window of the slow SCA. To achieve a large dynamic range, low leakage, and high bandwidth, the SCA sampling switches are implemented as 2.5 V nMOSFETs controlled by 1.2 V shift registers. Stored analog data are digitized by an external ADC at 10 bits or better. Specifications on operational parameters include a 4 GHz analog bandwidth and a dead time of 20 microseconds, corresponding to a 50 kHz readout rate, determined by the choice of the external ADC. PSEC5 has been submitted for fabrication.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Translytics: A Novel Approach for Runtime Selection of Database Layout Based on User’s Context

Currently, organizations have to maintain separate systems for transactions and analytics inside the company. Notably, different vendors provide the capabilities for either of these tasks that require specialized hardware or software. Data engineers are required to retrieve data from one source and transform it into another format to obtain the maximum benefit in a minimum time. Organizations strive for a competitive advantage that is achieved by fetching data from their customers and getting insights earliest for timely decision making. Present practices do not permit the view of the latest data for analytics since the first data have to be fetched from the source, transformed, and loaded to other systems to be utilized for analysis by relevant teams. This paper introduces a single system for both transactions and analytics. Our proposed solution would permit companies to seamlessly adapt our solution without the need to shift all of their data to newer systems and allow all the teams. It would grant all the teams to have a view of the latest available data without extra expertise and budget.

Tanvir, Muhammad Makhshif↗

A Data-Agnostic, Continuous Machine Learning Framework for Application in High Energy Physics and Beyond: Phase 1 Final Scientific/Technical Report

This Phase 1 effort has focused on the development of continual learning frameworks for use in machine learning, specifically in the applied context of High Energy Physics (HEP). Machine learning (ML) is a transformative technology by which computers, typically through the use of neural networks, are able to perform tasks with proficiency that rivals or surpasses that of human users. Model Degradation & Catastrophic Forgetting are two undesired phenomena which can occur in ML where the performance of a model degrades when either deployed on novel data streams, or trained on novel data which are sufficiently different than the data the models were initially trained on. A natural example where these sorts of effects can be observed is in the performance of detectors in harsh environments, where the detector signature may change over the lifetime of the detector as it ages and deteriorates — precisely what occurs in the experiments conducted in HEP. Real world HEP data is therefore an excellent test-ground and use-case for Continual Learning paradigms, which are techniques used in ML to counteract these problems. Ensemble learning is one such technique, where multiple smaller models are trained on subsets of the overall data and are ensembled together during inference. The intuition behind this technique is that, although there are shifts in the distributions which govern the incoming data streams, these shifts are not expected to be homogeneous or global. If a sufficient diversity in solutions within the various sub-models has been achieved, then at least one sub-model is expected to retain its performance within the overall ensemble. One further strength of this approach is that the architectures of the various models do not need to be identical, and in fact even different modalities of data can naturally be combined in this way. This work focused on applying ensemble learning techniques to derive results using two main datasets, anomaly detection in HEP data & time-series forecasting in semiconductor manufacturing data. Semiconductor manufacturing involves data with surprising similarity to that of HEP (e.g. wafer maps look very similar to digi-occupancy maps) and Cerium Lab’s prominence within the semiconductor industry makes semiconductor manufacturing a natural opportunity for commercialization of this work. Our efforts have led to two strong results. The first is that we evaluated the proposed ensembling techniques using previously proposed machine learning architectures for use in anomaly detection, namely AutoEncoder based models and their derivatives. We also developed new architectures which have not been evaluated in this context before. In fact, this work marks the first use of Vision Transformers for anomaly detection in HEP. Second, we demonstrated that ensemble learning significantly improves model performance in scenarios prone to degradation, validating its effectiveness across both HEP and semiconductor datasets. These results further support ensemble learning as a powerful strategy for mitigating catastrophic forgetting and maintaining robust performance in evolving data environments.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

High dimensional binary classification under label shift: phase transition and regularization

Label Shift has been widely believed to be harmful to the generalization performance of machine learning models. Researchers have proposed many approaches to mitigate the impact of the label shift, e.g., balancing the training data. However, these methods often consider the underparametrized regime, where the sample size is much larger than the data dimension. The research under the overparametrized regime is very limited. Here, to bridge this gap, we propose a new asymptotic analysis of the Fisher Linear Discriminant classifier for binary classification with label shift. Specifically, we prove that there exists a phase transition phenomenon: Under certain overparametrized regime, the classifier trained using imbalanced data outperforms the counterpart with reduced balanced data. Moreover, we investigate the impact of regularization to the label shift: The aforementioned phase transition vanishes as the regularization becomes strong.

binary classification↗

Assessing the ASME Section III, Division 5, Class A Primary Load Design Rules Against Creep Notch Effects

This report assesses the ASME Section III, Division 5, Subsection HB, Subpart B rules covering the design and construction of high temperature Class A nuclear reactor components for their robustness against creep notch effects. The creep notch effect combines the effect of multiaxial stresses on material creep deformation and damage. This report considers both effects, including the potential for creep mechanism shifts affecting the extrapolation of design data from high stress, short-time experimental data to low stress, long time operating component conditions. The report summarizes the history of and literature on multiaxial creep and surveys the current ASME design rules dealing with multiaxial effects. The report then describes the results of two dedicated numerical studies, one assessing the robustness of the ASME rules against uncertainty in extrapolating from uniaxial creep test data to multiaxial component conditions and a second study examining the potential effects of a mechanism shift from high stress, dislocation-mediated creep to diffusion-dominated creep at lower stresses. The final conclusion of the report is that the ASME rules adequately guard against multiaxial creep failure, though there are several aspects of the Code that could be optimized to provide less over-conservative design predictions or to provide a more consistent design margin as a function of temperature and stress. In a few areas, particularly the potential for mechanism shifts outside the currently-available experimental data, the Code rules should be revaluated as additional experimental data and new modeling and simulation results become available

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Landmark-embedded Gaussian process with applications for functional data modeling

In practice, we often need to infer the value of a target variable from functional observation data. A challenge in this task is that the relationship between the functional data and the target variable is very complex: the target variable not only influences the shape but also the location of the functional data. In addition, due to the uncertainties in the environment, the relationship is probabilistic, that is, for a given fixed target variable value, we still see variations in the shape and location of the functional data. To address this challenge, we present a landmark-embedded Gaussian process model that describes the relationship between the functional data and the target variable. A unique feature of the model is that landmark information is embedded in the Gaussian process model so that both the shape and location information of the functional data are considered simultaneously in a unified manner. Gibbs-Metropolis-Hasting algorithm is used for model parameters estimation and target variable inference. The performance of the proposed framework is evaluated by extensive numerical studies and a case study of nano-sensor calibration.

42 ENGINEERING↗

Superallowed 0 + → 0 + nuclear β decays: 2020 critical survey, with implications for V ud and CKM unitarity

A new critical survey of all half-life, decay-energy and branching-ratio measurements related to 23 superallowed 0 + → 0 + β decays is presented. Included are 222 individual measurements of comparable precision obtained from 174 published references. Compared with our last survey in 2015, we have added results from 28 new publications and eliminated an approximately equal number whose results have been superseded by much more precise modern data. We obtain world-average ft values for each of the 21 transitions that have a complete set of data, then apply radiative and isospin-symmetry-breaking corrections to extract “corrected” Ft values. Fifteen of these Ft values now have a precision of 0.3% or better and all take the same value within statistics, as expected from conservation of the vector current. Their average, Ft¯, when combined with the muon lifetime, yields the up-down quark-mixing element of the Cabibbo-Kobayashi-Maskawa matrix, V ud = 0.97373 ± 0.00031. This is lower than our 2015 result by one standard deviation and its uncertainty is increased by 50%. This is a consequence, not of any shifts in the experimental data, but of new calculations for the radiative corrections. The lower V ud value now leads to greater tension in the top-row test of unitarity in the CKM matrix. Updates in experimental data have independently led to a factor-of-two tighter limit being set on the possible existence of a scalar interaction. In conclusion, the new limit on Fierz interference is b F ≤ 0.0033 at the 90% confidence level.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Fiducial-cosmology-dependent systematics for the DESI 2024 full-shape analysis

We assess the impact of the fiducial cosmology choice on cosmological inference from full-shape (FS) fits of the galaxy power spectrum in the DESI 2024 Data Release 1 (DR1). Using a suite of AbacusSummit DR1 mock catalogues based on the Planck 2018 best-fit cosmology, we quantify potential systematic shifts introduced by analysing the data under five secondary cosmologies — featuring variations in matter density, thawing dark energy, higher effective number of neutrino species, reduced clustering amplitude, and the DESI DR1 BAO best-fit w 0 w a CDM cosmology — relative to DESI's baseline Planck 2018 cosmology. We investigate two complementary FS analysis approaches: full-modelling (FM) and ShapeFit (SF), each with distinct sensitivities to the assumed fiducial model. Across all tracers, we find for FM that systematic shifts induced by fiducial cosmology mismatches remain well below the DESI DR1 statistical uncertainties, with maximum deviations of 0.22σ DR1 in ΛCDM scenarios and 0.12σ DR1+SN when including SN Ia mock data in extended w 0 w a CDM fits. For SF, the shifts in the compressed parameters remain below 0.45σ DR1 for all tracers and cosmologies.

dark energy experiments↗

Machine Learning for Optimized Polarization at Jefferson Lab

Polarized cryo-targets and polarized photon beams are widely used in experiments at Jefferson Lab. Traditional methods for maintaining the optimal polarization involve manual adjustments throughout data taking by human shift takers. This may introduce some level of inconsistency simply due to the wide variety of experience and expertise of the shift takers themselves. Implementing machine learning-based control systems can improve the stability of the polarization without relying on human intervention. The cryo-target polarization is influenced by temperature, microwave energy, the distribution of paramagnetic radicals, as well as operational conditions including the radiation dose. Diamond radiators are used to generate linearly polarized photons from a primary electron beam. The energy spectrum of these photons can drift over time due to changes in the primary electron beam conditions and diamond degradation. As a first step towards automating the continuous optimization and control processes, uncertainty aware surrogate models have been developed to predict the polarization based on historical data. This talk will provide an overview of the use cases and models developed, highlighting the collaboration between data scientists and physicists at Jefferson Lab.

Jeske, Torri [Thomas Jefferson National Accelerato↗

Regulators of early maize leaf development inferred from transcriptomes of laser capture microdissection (LCM)-isolated embryonic leaf cells

The superior photosynthetic efficiency of C 4 leaves over C 3 leaves is owing to their unique Kranz anatomy, in which the vein is surrounded by one layer of bundle sheath (BS) cells and one layer of mesophyll (M) cells. Kranz anatomy development starts from three contiguous ground meristem (GM) cells, but its regulators and underlying molecular mechanism are largely unknown. To identify the regulators, we obtained the transcriptomes of 11 maize embryonic leaf cell types from five stages of pre-Kranz cells starting from median GM cells and six stages of pre-M cells starting from undifferentiated cells. Principal component and clustering analyses of transcriptomic data revealed rapid pre-Kranz cell differentiation in the first two stages but slow differentiation in the last three stages, suggesting early Kranz cell fate determination. In contrast, pre-M cells exhibit a more prolonged transcriptional differentiation process. Differential gene expression and coexpression analyses identified gene coexpression modules, one of which included 3 auxin transporter and 18 transcription factor (TF) genes, including known regulators of Kranz anatomy and/or vascular development. In situ hybridization of 11 TF genes validated their expression in early Kranz development. We determined the binding motifs of 15 TFs, predicted TF target gene relationships among the 18 TF and 3 auxin transporter genes, and validated 67 predictions by electrophoresis mobility shift assay. From these data, we constructed a gene regulatory network for Kranz development. Our study sheds light on the regulation of early maize leaf development and provides candidate leaf development regulators for future study.

59 BASIC BIOLOGICAL SCIENCES↗

Multi-fidelity information fusion with concatenated neural networks

Recently, computational modeling has shifted towards the use of statistical inference, deep learning, and other data-driven modeling frameworks. Although this shift in modeling holds promise in many applications like design optimization and real-time control by lowering the computational burden, training deep learning models needs a huge amount of data. This big data is not always available for scientific problems and leads to poorly generalizable data-driven models. This gap can be furnished by leveraging information from physics-based models. Exploiting prior knowledge about the problem at hand, this study puts forth a physics-guided machine learning (PGML) approach to build more tailored, effective, and efficient surrogate models. For our analysis, without losing its generalizability and modularity, we focus on the development of predictive models for laminar and turbulent boundary layer flows. In particular, we combine the self-similarity solution and power-law velocity profile (low-fidelity models) with the noisy data obtained either from experiments or computational fluid dynamics simulations (high-fidelity models) through a concatenated neural network. We illustrate how the knowledge from these simplified models results in reducing uncertainties associated with deep learning models applied to boundary layer flow prediction problems. The proposed multi-fidelity information fusion framework produces physically consistent models that attempt to achieve better generalization than data-driven models obtained purely based on data. While we demonstrate our framework for a problem relevant to fluid mechanics, its workflow and principles can be adopted for many scientific problems where empirical, analytical, or simplified models are prevalent. In line with grand demands in novel PGML principles, this work builds a bridge between extensive physics-based theories and data-driven modeling paradigms and paves the way for using hybrid physics and machine learning modeling approaches for next-generation digital twin technologies.

42 ENGINEERING↗

Downscaled precipitation and mean air temperature datasets; East-Taylor subbasin; 2008-2019; daily temporal resolution; 400 m spatial resolution

This dataset provides gridded meteorological forcing data (specifically, daily precipitation and daily mean air temperature). The dataset has been generated by downscaling the Parameter-elevation Regressions on Independent Slopes Model (PRISM) dataset from a spatial resolution of 800 m to 400 m. The time period is 2008-2019, and the mapped area is East Taylor subbasin in Upper Colorado. The data are in the form of NetCDF files, arranged by year. Compared to the PRISM dataset, the dates of the downscaled dataset are shifted backwards by one (e.g., downscaled data for May 25 corresponds to PRISM data for May 26). This temporal shifting makes the meteorological forcing correspond more closely to the prescribed date. The NetCDF format is a standard raster format that can be read using any Geographic Information System (GIS) software; plenty of modules exist in popular scripting languages (such as Python, R and Matlab) that can also be used to read NetCDF files.East_Taylor_Tavg_PRISM400 corresponds to mean air temperature.East_Taylor_Precip_PRISM400.zip corresponds to precipitation.The proprietary PRISM data (800 m resolution) were purchased with funding from the Watershed Function Scientific Focus Area supported by U.S. Department of Energy, Office of Science, Office of Biological and Environmental Research under award no. DE-AC02-05CH11231.Dataset update on March 9, 2022:Uploaded new data files that are identical to the earlier files but are now in the NetCDF format, arranged by year. Earlier, the files were in the GeoTiff format, arranged by date.

54 ENVIRONMENTAL SCIENCES↗

A Data-Driven Voltage Control Strategy for Distribution Grids With Distributed Energy Resources

Traditionally, distribution system control approaches have been model-based. The deployment of advanced metering infrastructure has provided electric utilities with the capability of data-driven control with real-time measurements. The shift from model-based to data-driven control represents a significant advancement in the management of distribution systems, offering a more adaptive approach to system control because of the ability to dynamically adapt to changing conditions without the need for system modeling. Here, in this paper, a behavioral data-driven control method is developed to provide voltage regulation to an actual distribution system by controlling the legacy devices and distributed energy resource (DER) assets. The studied distribution system has a load tap changer and three capacitor banks as the legacy devices and photovoltaic systems as the DERs. The performance of the proposed control algorithm is validated using a laboratory test bed setup considering multiple scenarios. The results show that the proposed control achieved 99% voltage regulation.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Validation of time-dependent shift using the pulsed sphere benchmarks

The detailed behavior of neutrons in a rapidly changing time-dependent physical system is a challenging computational physics problem, particularly when using Monte Carlo methods on heterogeneous high-performance computing architectures. A small number of algorithms and code implementations have been shown to be performant for time-independent (fixed source and k-eigenvalue) Monte Carlo, and there are existing simulation tools that successfully solve the time-dependent Monte Carlo problem on smaller computing platforms. To bridge this gap, a time-dependent version of ORNL’s Shift code has been recently developed. Shift’s history-based algorithm on CPUs, and its event-based algorithm on GPUs, have both been observed to scale well to very large numbers of processors, which motivated the extension of this code to solve time-dependent problems. The validation of this new capability requires a comparison with time-dependent neutron experiments. Lawrence Livermore National Laboratory’s (LLNL) pulsed sphere benchmark experiments were simulated in Shift to validate both the time-independent as well as new time-dependent features recently incorporated into Shift. A suite of pulsed-sphere models was simulated using Shift and compared to the available experimental data and simulations with MCNP. Overall results indicate that Shift accurately simulates the pulsed sphere benchmarks, and that the new time-dependent modifications of Shift are working as intended. Validated exascale neutron transport codes are essential for a wide variety of future multiphysics applications.

Palmer, Camille J.↗