Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data shift”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Simple new methods for deducing lifetimes in recoil distance Doppler-shift measurements

In this work, new approaches for lifetime determination using data from recoil distance Doppler-shift experiments are presented based on the fundamental properties of the functions describing the time evolution of the population of excited nuclear states. To some extent, one of them represents a contraction of the well-known Differential decay-curve method (DDCM) by using the most reliable data point [the maximum of the $n_i(t)$ function describing the population of level $\textit{i}$ in time] and a purely numerical procedure avoiding any fitting of decay curves. The combination with the standard DDCM analysis is promising for improving the reliability and the precision of the results for the lifetimes obtained. The novel part of the approach consists of using a chain of equations at the consecutive maxima of the ni(t) functions, which allow us to precisely determine the ratio of the lifetimes of two consecutive levels and, in the case where one of these lifetimes is known, to determine the unknown one. In addition, a simple integral derivation of the lifetime is presented involving the peak areas measured at different distances, and an application of the first moments (expectation values and centroids in time) of the $n_i(t)$ functions for determining lifetimes is also demonstrated to be useful.

47 OTHER INSTRUMENTATION↗

Potential of Data Center Controls in Grid Services

The rapid proliferation of large data centers brings both challenges and opportunities for grid reliability. The data center resources and their potential flexibility have the potential to contribute resources to grid operations. Through capabilities like energy shifting and resource coordination, data centers can help reduce their net demand on the transmission network, as well as provide additional grid services to support reliable operation on the grid. While transient and long-term grid planning and operations are the scenarios that draw most attention, the quasi-steady state timeseries (QSTS) operation of data centers and grid bring interesting scenarios that can help evaluate the data center controls to aid grid services. This work is focused on modeling data centers for QSTS applications – incorporating the AI data center load profiles and building on the PNNL digital twin model for the thermal management loads to enable simulation studies to reveal the impact of data center controls on grid performance. This includes the integration of a QSTS battery and natural gas generator model to incorporate local resource impacts to the system. The simulation study is performed with a modified IEEE 24-Bus transmission system. Scenarios are focused on evaluating the data center load impacts on the transmission system and leveraging both data center and local generation controls to mitigate those impacts and provide additional grid services. The data center controls revealed the ability to contribute to two main kinds of grid services: preventing congestion on a weak grid by coordinating the data center resources with the collocated BESS and onsite generation; and the ability to help the grid operations during stressed times of operation like during a contingency. Leveraging these and other capabilities has the potential to help data centers become grid responsive assets, aiding in both their integration into the power system and grid reliability.

power grid simulation↗

EVT 16s Data and Large Supplementary Files

Soil microorganisms often interact to carry out decomposition of complex organic carbon and nitrogen compounds, such as chitin, but the high diversity and complexity of the soil microbiome and habitat has posed a challenge to elucidating such interactions between soil microorganisms. Here, we seek to address this challenge through analysis of a model soil consortium (MSC-2) of eight soil bacterial species. Our aim was to elucidate specific roles of the member species during chitin metabolism. Samples were collected from MSC-2 incubated in chitin-enriched soil over three months. Multi-omics was used to understand how the community composition, transcripts, proteins and chitin decomposition shifted over time. The data clearly and consistently revealed a temporal shift during chitin decomposition with defined contributions by individual species. A Streptomyces genus member (sp001905665) was a key player in early steps of chitin decomposition, with other MSC-2 members being central in carrying out later steps. These results illustrate how multi-omics applied to a defined consortium untangles interactions between soil microorganisms.

McClure, Ryan [Pacific Northwest National Laborato↗

The Importance of Being Adaptable: An Exploration of the Power and Limitations of Domain Adaptation for Simulation-Based Inference with Galaxy Clusters

The application of deep machine learning methods in astronomy has exploded in the last decade, with new models showing remarkably improved performance on benchmark tasks. Not nearly enough attention is given to understanding the models' robustness, especially when the test data are systematically different from the training data, or "out of domain." Domain shift poses a significant challenge for simulation-based inference, where models are trained on simulated data but applied to real observational data. In this paper, we explore domain shift and test domain adaptation methods for a specific scientific case: simulation-based inference for estimating galaxy cluster masses from X-ray profiles. We build datasets to mimic simulation-based inference: a training set from the Magneticum simulation, a scatter-augmented training set to capture uncertainties in scaling relations, and a test set derived from the IllustrisTNG simulation. We demonstrate that the Test Set is out of domain in subtle ways that would be difficult to detect without careful analysis. We apply three deep learning methods: a standard neural network (NN), a neural network trained on the scatter-augmented input catalogs, and a Deep Reconstruction-Regression Network (DRRN), a semi-supervised deep model engineered to address domain shift. Although the NN improves results by 17% in the Training Data, it performs 40% worse on the out-of-domain Test Set. Surprisingly, the Scatter-Augmented Neural Network (SANN) performs similarly. While the DRRN is successful in mapping the training and Test Data onto the same latent space, it consistently underperforms compared to a straightforward Yx scaling relation. These results serve as a warning that simulation-based inference must be handled with extreme care, as subtle differences between training simulations and observational data can lead to unforeseen biases creeping into the results.

Ntampaka, Michelle [Baltimore, Space Telescope Sci↗

Roadmap on data-centric materials science

Science is and always has been based on data, but the terms ‘data-centric’ and the ‘4th paradigm’ of materials research indicate a radical change in how information is retrieved, handled and research is performed. It signifies a transformative shift towards managing vast data collections, digital repositories, and innovative data analytics methods. The integration of artificial intelligence and its subset machine learning, has become pivotal in addressing all these challenges. This Roadmap on Data-Centric Materials Science explores fundamental concepts and methodologies, illustrating diverse applications in electronic-structure theory, soft matter theory, microstructure research, and experimental techniques like photoemission, atom probe tomography, and electron microscopy. While the roadmap delves into specific areas within the broad interdisciplinary field of materials science, the provided examples elucidate key concepts applicable to a wider range of topics. The discussed instances offer insights into addressing the multifaceted challenges encountered in contemporary materials research.

36 MATERIALS SCIENCE↗

AI Data Quality Monitoring with Hydra

Hydra is an extensible framework for training and managing AI for near real time monitoring that aims to replace the tedious and repetitive data quality monitoring activities the shift crew and online monitoring coordinator typically perform. It continuously scans incoming data in the form of monitoring plots for signs of problems, flagging them for human review. A web app was developed such that experts can efficiently label images for training. Labels are stored in a database for use in training and model validation. Backed up by a comprehensive database, it utilizes an additional web based front-end for viewing the current monitoring status from anywhere in the world. The system has been in production use for the GlueX experiment at Jefferson Lab for more than 2 years with new features still under active development.

Britton, Thomas↗

Propagating Uncertainties in the SALT3 Model-training Process to Cosmological Constraints

Type Ia supernovae (SNe Ia) are standardizable candles that must be modeled empirically to yield cosmological constraints. To understand the robustness of this modeling to variations in the model-training procedure, we build an end-to-end pipeline to test the recently developed SALT3 model. We explore the consequences of removing pre-2000s low-z or poorly calibrated U-band data, adjusting the amount and fidelity of SN Ia spectra, and using a model-independent framework to simulate the training data. We find that the SALT3 model surfaces are improved by having additional spectra and U-band data, and can be shifted by ~5% if host-galaxy contamination is not sufficiently removed from SN spectra. We find that resulting measurements of w are consistent to within 2.5% for all of the training variants explored in this work, with the largest shifts coming from variants that add color-dependent calibration offsets or host-galaxy contamination to the training spectra and those that remove pre-2000s low-z data. These results demonstrate that the SALT3 model-training procedure is largely robust to reasonable variations in the training data, but that additional attention must be paid to the treatment of spectroscopic data in the training process. We also find that the training procedure is sensitive to the color distributions of the input data—the resulting w measurement can be biased by ~2% if the color distribution is not sufficiently wide. Future low-z data, particularly u-band observations and high signal-to-noise ratio SN Ia spectra, will help to significantly improve SN Ia modeling in the coming years.

79 ASTRONOMY AND ASTROPHYSICS↗

A biology-informed similarity metric for simulated patches of human cell membrane

Complex scientific inquiries rely increasingly upon large and autonomous multiscale simulation campaigns, which fundamentally require similarity metrics to quantify ‘sufficient’ changes among data and/or configurations. However, subject matter experts are often unable to articulate similarity precisely or in terms of well-formulated definitions, especially when new hypotheses are to be explored, making it challenging to design a meaningful metric. Furthermore, the key to practical usefulness of such metrics to enable autonomous simulations lies in in situ inference, which requires generalization to possibly substantial distributional shifts in unseen, future data. Here, we address these challenges in a cancer biology application and develop a meaningful similarity metric for ‘patches’—regions of simulated human cell membrane that express interactions between certain proteins of interest and relevant lipids. In the absence of well-defined conditions for similarity, we leverage several biology-informed notions about data and the underlying simulations to impose inductive biases on our metric learning framework, resulting in a suitable similarity metric that also generalizes well to significant distributional shifts encountered during the deployment. We combine these intuitions to organize the learned embedding space in a multiscale manner, which makes the metric robust to incomplete and even contradictory intuitions. Our approach delivers a metric that not only performs well on the conditions used for its development and other relevant criteria, but also learns key spatiotemporal relationships without ever being exposed to any such information during training.

97 MATHEMATICS AND COMPUTING↗

Exploring OpenSNAPI Use Cases and Evolving Requirements [Slides]

Emerging system architectures are rapidly transforming in order to meet shifting requirements. Motivated by expanding data volumes, energy efficiency concerns, and the omnipresent need to improve performance, architectures are increasingly adopting a data-centric approach. At the core of this concept is the goal of minimizing data motion and instead processing data in-situ to the greatest degree possible. Therefore, data-centric designs, in contrast to conventional CPU-centric models, typically distribute compute capabilities throughout the architecture. As part of this paradigm shift, a novel class of devices known as data processing units (DPUs), alongside CPUs and GPUs, are quickly forming a third pillar of data-centric systems. These devices, which include smart network adapters and switches, seek to offload computation on data at the network edge as well as in-flight within the network fabric. The Open Smart Network API (OpenSNAPI) project seeks to develop a unified API for DPU devices. In our previous talks, we introduced the OpenSNAPI project and detailed our investigations regarding the viability of offloading compute intensive kernels to BlueField DPUs. In contrast, in this talk we detail our efforts to offload application-level file I/O to the DPU. We also discuss plans and early efforts to explore in-network compute capabilities. Finally, we describe our observations with respect to the evolving design of OpenSNAPI.

97 MATHEMATICS AND COMPUTING↗

Boundary-Aware Adversarial Learning Domain Adaption and Active Learning for Cross-Sensor Building Extraction

The use of convolutional neural networks (CNNs) for building extraction from remote sensing images has been widely studied and many public datasets have been made available for accelerating development of these CNN models. Yet adapting pretrained models at scale in real-world scenarios remains a challenging task. The main barrier is that certain new labels are still needed to compensate for domain shifting between the labeled data and new images that potentially cover new geographic locations or that are from a different sensor. In this article, we propose to add informatively labeled samples from a new image pool under the paradigm of active learning. To select the most useful samples based on model uncertainty, we first tackle the problem of uncalibrated uncertainty estimation due to distribution shifting by adapting feature extractors with boundary-based adversarial learning. Calibrated uncertainty is used as the query criterion in the active learning process, where the most uncertain samples are selected for annotation and included for model retraining. The proposed workflow was tested with three data pairs in which each workflow represents a scenario often encountered in real-world applications, including adapting pretrained models to new images collected with different sensors or to new geographic areas where appearances and types of buildings are very different. Compared to several baselines, including random sampling, temperature scaling (a well-known uncertainty calibration technique), different query strategies, and active domain adaptation methods, the proposed workflow shows that strategically querying a smaller set of samples for labeling achieves comparable or better building extraction performance. The proposed method reduces the number of labeled samples required to achieve sufficient model accuracy, thus significantly reducing hundreds of person-hours for labeled data creation. In addition, we include a few considerations when deploying this workflow in a GPU cluster that can be easily adapted to achieve operational building extraction model retraining.

97 MATHEMATICS AND COMPUTING↗

Integration of sensors through additive manufacturing leading to increased efficiencies of gas turbines for power generation and propulsion

To realize the full capability of additively manufactured components in complex energy systems, it is imperative to minimize early component failures during development phases and during operation. Traditional field feedback timelines and offline inspection protocols significantly reduce the design-manufacturing iteration times. To address this specific question, the project developed and demonstrated a method for the integration of sensors into complex components through additive manufacturing. The team used gas turbine engines as a platform, which meets the need of both power generation and propulsion and offer opportunities for cost reductions and efficiency increases. The innovation of this intelligent integration of sensors into complex components uniquely customized to address questions of integrity and durability for additively manufactured components. With real-time sensing data from additively manufactured components, turbine manufacturers will realize higher efficiencies, reduced component failures, and a 30-50% acceleration in product deployment of high efficiency gas turbine components due to a faster reduction in component risk assessment under actual operating conditions. This is a transformative shift towards a data-driven design and qualification of additively manufactured gas turbine components. To directly integrate sensors into additively manufactured components with all the complexities of actual hardware, powder bed fusion (direct metal laser sintering) and laser metal deposition technologies was developed. Validation took take place in two university laboratories both of which contain actual engine hardware and closely simulate a gas turbine prior to demonstrating the technology in a turbine development test. Indeed, two major technologies from this research cold impact turbine systems in the near future: (1) higher efficiency materials and designs enabled by additive manufacturing with 50% faster design to manufacturing cycle time, to enable faster time-to-market targets; and (2) integration of sensors into additively manufactured components enabling broad health and condition based prognostics for faster component and engine risk reduction.

33 ADVANCED PROPULSION SYSTEMS↗

Design of an 8-channel 40 GS/s 20 mW/Ch waveform sampling ASIC in 65 nm CMOS

One picosecond timing resolution is the entry point to signature based searches relying on secondary/tertiary vertices and particle identification. We describe PSEC5, an 8-channel 40 GS/s waveform-sampling ASIC in TSMC 65 nm process targetting one picosecond resolution at 20 mW power per channel. Each channel consists of four fast and one slow switched capacitor arrays (SCA), allowing for picosecond time resolution combined with a long effective buffer. Each fast SCA is 1.6 ns long and has a nominal sampling rate of 40 GS/s. The slow SCA is 204.8 ns long and samples at 5 GS/s. Recording of the analog data for each channel is triggered by a fast discriminator capable of multiple triggering during the window of the slow SCA. To achieve a large dynamic range, low leakage, and high bandwidth, the SCA sampling switches are implemented as 2.5 V nMOSFETs controlled by 1.2 V shift registers. Stored analog data are digitized by an external ADC at 10 bits or better. Specifications on operational parameters include a 4 GHz analog bandwidth and a dead time of 20 microseconds, corresponding to a 50 kHz readout rate, determined by the choice of the external ADC. PSEC5 has been submitted for fabrication.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Translytics: A Novel Approach for Runtime Selection of Database Layout Based on User’s Context

Currently, organizations have to maintain separate systems for transactions and analytics inside the company. Notably, different vendors provide the capabilities for either of these tasks that require specialized hardware or software. Data engineers are required to retrieve data from one source and transform it into another format to obtain the maximum benefit in a minimum time. Organizations strive for a competitive advantage that is achieved by fetching data from their customers and getting insights earliest for timely decision making. Present practices do not permit the view of the latest data for analytics since the first data have to be fetched from the source, transformed, and loaded to other systems to be utilized for analysis by relevant teams. This paper introduces a single system for both transactions and analytics. Our proposed solution would permit companies to seamlessly adapt our solution without the need to shift all of their data to newer systems and allow all the teams. It would grant all the teams to have a view of the latest available data without extra expertise and budget.

Tanvir, Muhammad Makhshif↗

A Data-Agnostic, Continuous Machine Learning Framework for Application in High Energy Physics and Beyond: Phase 1 Final Scientific/Technical Report

This Phase 1 effort has focused on the development of continual learning frameworks for use in machine learning, specifically in the applied context of High Energy Physics (HEP). Machine learning (ML) is a transformative technology by which computers, typically through the use of neural networks, are able to perform tasks with proficiency that rivals or surpasses that of human users. Model Degradation & Catastrophic Forgetting are two undesired phenomena which can occur in ML where the performance of a model degrades when either deployed on novel data streams, or trained on novel data which are sufficiently different than the data the models were initially trained on. A natural example where these sorts of effects can be observed is in the performance of detectors in harsh environments, where the detector signature may change over the lifetime of the detector as it ages and deteriorates — precisely what occurs in the experiments conducted in HEP. Real world HEP data is therefore an excellent test-ground and use-case for Continual Learning paradigms, which are techniques used in ML to counteract these problems. Ensemble learning is one such technique, where multiple smaller models are trained on subsets of the overall data and are ensembled together during inference. The intuition behind this technique is that, although there are shifts in the distributions which govern the incoming data streams, these shifts are not expected to be homogeneous or global. If a sufficient diversity in solutions within the various sub-models has been achieved, then at least one sub-model is expected to retain its performance within the overall ensemble. One further strength of this approach is that the architectures of the various models do not need to be identical, and in fact even different modalities of data can naturally be combined in this way. This work focused on applying ensemble learning techniques to derive results using two main datasets, anomaly detection in HEP data & time-series forecasting in semiconductor manufacturing data. Semiconductor manufacturing involves data with surprising similarity to that of HEP (e.g. wafer maps look very similar to digi-occupancy maps) and Cerium Lab’s prominence within the semiconductor industry makes semiconductor manufacturing a natural opportunity for commercialization of this work. Our efforts have led to two strong results. The first is that we evaluated the proposed ensembling techniques using previously proposed machine learning architectures for use in anomaly detection, namely AutoEncoder based models and their derivatives. We also developed new architectures which have not been evaluated in this context before. In fact, this work marks the first use of Vision Transformers for anomaly detection in HEP. Second, we demonstrated that ensemble learning significantly improves model performance in scenarios prone to degradation, validating its effectiveness across both HEP and semiconductor datasets. These results further support ensemble learning as a powerful strategy for mitigating catastrophic forgetting and maintaining robust performance in evolving data environments.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

High dimensional binary classification under label shift: phase transition and regularization

Label Shift has been widely believed to be harmful to the generalization performance of machine learning models. Researchers have proposed many approaches to mitigate the impact of the label shift, e.g., balancing the training data. However, these methods often consider the underparametrized regime, where the sample size is much larger than the data dimension. The research under the overparametrized regime is very limited. Here, to bridge this gap, we propose a new asymptotic analysis of the Fisher Linear Discriminant classifier for binary classification with label shift. Specifically, we prove that there exists a phase transition phenomenon: Under certain overparametrized regime, the classifier trained using imbalanced data outperforms the counterpart with reduced balanced data. Moreover, we investigate the impact of regularization to the label shift: The aforementioned phase transition vanishes as the regularization becomes strong.

binary classification↗

Assessing the ASME Section III, Division 5, Class A Primary Load Design Rules Against Creep Notch Effects

This report assesses the ASME Section III, Division 5, Subsection HB, Subpart B rules covering the design and construction of high temperature Class A nuclear reactor components for their robustness against creep notch effects. The creep notch effect combines the effect of multiaxial stresses on material creep deformation and damage. This report considers both effects, including the potential for creep mechanism shifts affecting the extrapolation of design data from high stress, short-time experimental data to low stress, long time operating component conditions. The report summarizes the history of and literature on multiaxial creep and surveys the current ASME design rules dealing with multiaxial effects. The report then describes the results of two dedicated numerical studies, one assessing the robustness of the ASME rules against uncertainty in extrapolating from uniaxial creep test data to multiaxial component conditions and a second study examining the potential effects of a mechanism shift from high stress, dislocation-mediated creep to diffusion-dominated creep at lower stresses. The final conclusion of the report is that the ASME rules adequately guard against multiaxial creep failure, though there are several aspects of the Code that could be optimized to provide less over-conservative design predictions or to provide a more consistent design margin as a function of temperature and stress. In a few areas, particularly the potential for mechanism shifts outside the currently-available experimental data, the Code rules should be revaluated as additional experimental data and new modeling and simulation results become available

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Landmark-embedded Gaussian process with applications for functional data modeling

In practice, we often need to infer the value of a target variable from functional observation data. A challenge in this task is that the relationship between the functional data and the target variable is very complex: the target variable not only influences the shape but also the location of the functional data. In addition, due to the uncertainties in the environment, the relationship is probabilistic, that is, for a given fixed target variable value, we still see variations in the shape and location of the functional data. To address this challenge, we present a landmark-embedded Gaussian process model that describes the relationship between the functional data and the target variable. A unique feature of the model is that landmark information is embedded in the Gaussian process model so that both the shape and location information of the functional data are considered simultaneously in a unified manner. Gibbs-Metropolis-Hasting algorithm is used for model parameters estimation and target variable inference. The performance of the proposed framework is evaluated by extensive numerical studies and a case study of nano-sensor calibration.

42 ENGINEERING↗

Fiducial-cosmology-dependent systematics for the DESI 2024 full-shape analysis

We assess the impact of the fiducial cosmology choice on cosmological inference from full-shape (FS) fits of the galaxy power spectrum in the DESI 2024 Data Release 1 (DR1). Using a suite of AbacusSummit DR1 mock catalogues based on the Planck 2018 best-fit cosmology, we quantify potential systematic shifts introduced by analysing the data under five secondary cosmologies — featuring variations in matter density, thawing dark energy, higher effective number of neutrino species, reduced clustering amplitude, and the DESI DR1 BAO best-fit w 0 w a CDM cosmology — relative to DESI's baseline Planck 2018 cosmology. We investigate two complementary FS analysis approaches: full-modelling (FM) and ShapeFit (SF), each with distinct sensitivities to the assumed fiducial model. Across all tracers, we find for FM that systematic shifts induced by fiducial cosmology mismatches remain well below the DESI DR1 statistical uncertainties, with maximum deviations of 0.22σ DR1 in ΛCDM scenarios and 0.12σ DR1+SN when including SN Ia mock data in extended w 0 w a CDM fits. For SF, the shifts in the compressed parameters remain below 0.45σ DR1 for all tracers and cosmologies.

dark energy experiments↗