Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data shift”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

A biology-informed similarity metric for simulated patches of human cell membrane

Complex scientific inquiries rely increasingly upon large and autonomous multiscale simulation campaigns, which fundamentally require similarity metrics to quantify ‘sufficient’ changes among data and/or configurations. However, subject matter experts are often unable to articulate similarity precisely or in terms of well-formulated definitions, especially when new hypotheses are to be explored, making it challenging to design a meaningful metric. Furthermore, the key to practical usefulness of such metrics to enable autonomous simulations lies in in situ inference, which requires generalization to possibly substantial distributional shifts in unseen, future data. Here, we address these challenges in a cancer biology application and develop a meaningful similarity metric for ‘patches’—regions of simulated human cell membrane that express interactions between certain proteins of interest and relevant lipids. In the absence of well-defined conditions for similarity, we leverage several biology-informed notions about data and the underlying simulations to impose inductive biases on our metric learning framework, resulting in a suitable similarity metric that also generalizes well to significant distributional shifts encountered during the deployment. We combine these intuitions to organize the learned embedding space in a multiscale manner, which makes the metric robust to incomplete and even contradictory intuitions. Our approach delivers a metric that not only performs well on the conditions used for its development and other relevant criteria, but also learns key spatiotemporal relationships without ever being exposed to any such information during training.

97 MATHEMATICS AND COMPUTING↗

Exploring OpenSNAPI Use Cases and Evolving Requirements [Slides]

Emerging system architectures are rapidly transforming in order to meet shifting requirements. Motivated by expanding data volumes, energy efficiency concerns, and the omnipresent need to improve performance, architectures are increasingly adopting a data-centric approach. At the core of this concept is the goal of minimizing data motion and instead processing data in-situ to the greatest degree possible. Therefore, data-centric designs, in contrast to conventional CPU-centric models, typically distribute compute capabilities throughout the architecture. As part of this paradigm shift, a novel class of devices known as data processing units (DPUs), alongside CPUs and GPUs, are quickly forming a third pillar of data-centric systems. These devices, which include smart network adapters and switches, seek to offload computation on data at the network edge as well as in-flight within the network fabric. The Open Smart Network API (OpenSNAPI) project seeks to develop a unified API for DPU devices. In our previous talks, we introduced the OpenSNAPI project and detailed our investigations regarding the viability of offloading compute intensive kernels to BlueField DPUs. In contrast, in this talk we detail our efforts to offload application-level file I/O to the DPU. We also discuss plans and early efforts to explore in-network compute capabilities. Finally, we describe our observations with respect to the evolving design of OpenSNAPI.

97 MATHEMATICS AND COMPUTING↗

Boundary-Aware Adversarial Learning Domain Adaption and Active Learning for Cross-Sensor Building Extraction

The use of convolutional neural networks (CNNs) for building extraction from remote sensing images has been widely studied and many public datasets have been made available for accelerating development of these CNN models. Yet adapting pretrained models at scale in real-world scenarios remains a challenging task. The main barrier is that certain new labels are still needed to compensate for domain shifting between the labeled data and new images that potentially cover new geographic locations or that are from a different sensor. In this article, we propose to add informatively labeled samples from a new image pool under the paradigm of active learning. To select the most useful samples based on model uncertainty, we first tackle the problem of uncalibrated uncertainty estimation due to distribution shifting by adapting feature extractors with boundary-based adversarial learning. Calibrated uncertainty is used as the query criterion in the active learning process, where the most uncertain samples are selected for annotation and included for model retraining. The proposed workflow was tested with three data pairs in which each workflow represents a scenario often encountered in real-world applications, including adapting pretrained models to new images collected with different sensors or to new geographic areas where appearances and types of buildings are very different. Compared to several baselines, including random sampling, temperature scaling (a well-known uncertainty calibration technique), different query strategies, and active domain adaptation methods, the proposed workflow shows that strategically querying a smaller set of samples for labeling achieves comparable or better building extraction performance. The proposed method reduces the number of labeled samples required to achieve sufficient model accuracy, thus significantly reducing hundreds of person-hours for labeled data creation. In addition, we include a few considerations when deploying this workflow in a GPU cluster that can be easily adapted to achieve operational building extraction model retraining.

97 MATHEMATICS AND COMPUTING↗

Integration of sensors through additive manufacturing leading to increased efficiencies of gas turbines for power generation and propulsion

To realize the full capability of additively manufactured components in complex energy systems, it is imperative to minimize early component failures during development phases and during operation. Traditional field feedback timelines and offline inspection protocols significantly reduce the design-manufacturing iteration times. To address this specific question, the project developed and demonstrated a method for the integration of sensors into complex components through additive manufacturing. The team used gas turbine engines as a platform, which meets the need of both power generation and propulsion and offer opportunities for cost reductions and efficiency increases. The innovation of this intelligent integration of sensors into complex components uniquely customized to address questions of integrity and durability for additively manufactured components. With real-time sensing data from additively manufactured components, turbine manufacturers will realize higher efficiencies, reduced component failures, and a 30-50% acceleration in product deployment of high efficiency gas turbine components due to a faster reduction in component risk assessment under actual operating conditions. This is a transformative shift towards a data-driven design and qualification of additively manufactured gas turbine components. To directly integrate sensors into additively manufactured components with all the complexities of actual hardware, powder bed fusion (direct metal laser sintering) and laser metal deposition technologies was developed. Validation took take place in two university laboratories both of which contain actual engine hardware and closely simulate a gas turbine prior to demonstrating the technology in a turbine development test. Indeed, two major technologies from this research cold impact turbine systems in the near future: (1) higher efficiency materials and designs enabled by additive manufacturing with 50% faster design to manufacturing cycle time, to enable faster time-to-market targets; and (2) integration of sensors into additively manufactured components enabling broad health and condition based prognostics for faster component and engine risk reduction.

33 ADVANCED PROPULSION SYSTEMS↗

The Single Event Upset (SEU) response to 590 MeV protons

The presence of high-energy protons in cosmic rays, solar flares, and trapped radiation belts around Jupiter poses a threat to the Galileo project. Results of a test of 10 device types (including 1K RAM, 4-bit microP sequencer, 4-bit slice, 9-bit data register, 4-bit shift register, octal flip-flop, and 4-bit counter) exposed to 590 MeV protons at the Swiss Institute of Nuclear Research are presented to clarify the picture of SEU response to the high-energy proton environment of Jupiter. It is concluded that the data obtained should remove the concern that nuclear reaction products generated by protons external to the device can cause significant alteration in the device SEU response. The data also show only modest increases in SEU cross section as proton energies are increased up to the upper limits of energy for both the terrestrial and Jovian trapped proton belts.

Nichols, D. K.↗

Balloon-aircraft ranging, data, and voice experiment.

The test facilities used in the experiment consisted of a ground station, a balloon platform, radar tracking stations, and a test aircraft. As a direct result of the experiment, several modifications have been incorporated into the equipment. The two most important modifications were the introduction of a 10-sec delay into the search mode and the use of differentially coded phase shift keying for the data channel.

Wishna, S.↗

Design of an 8-channel 40 GS/s 20 mW/Ch waveform sampling ASIC in 65 nm CMOS

One picosecond timing resolution is the entry point to signature based searches relying on secondary/tertiary vertices and particle identification. We describe PSEC5, an 8-channel 40 GS/s waveform-sampling ASIC in TSMC 65 nm process targetting one picosecond resolution at 20 mW power per channel. Each channel consists of four fast and one slow switched capacitor arrays (SCA), allowing for picosecond time resolution combined with a long effective buffer. Each fast SCA is 1.6 ns long and has a nominal sampling rate of 40 GS/s. The slow SCA is 204.8 ns long and samples at 5 GS/s. Recording of the analog data for each channel is triggered by a fast discriminator capable of multiple triggering during the window of the slow SCA. To achieve a large dynamic range, low leakage, and high bandwidth, the SCA sampling switches are implemented as 2.5 V nMOSFETs controlled by 1.2 V shift registers. Stored analog data are digitized by an external ADC at 10 bits or better. Specifications on operational parameters include a 4 GHz analog bandwidth and a dead time of 20 microseconds, corresponding to a 50 kHz readout rate, determined by the choice of the external ADC. PSEC5 has been submitted for fabrication.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Using quasar X-ray and UV flux measurements to constrain cosmological model parameters

ABSTRACT Risaliti and Lusso have compiled X-ray and UV flux measurements of 1598 quasars (QSOs) in the redshift range 0.036 ≤ z ≤ 5.1003, part of which, z ∼ 2.4 − 5.1, is largely cosmologically unprobed. In this paper we use these QSO measurements, alone and in conjunction with baryon acoustic oscillation (BAO) and Hubble parameter [H(z)] measurements, to constrain cosmological parameters in six different cosmological models, each with two different Hubble constant priors. In most of these models, given the larger uncertainties, the QSO cosmological parameter constraints are mostly consistent with those from the BAO + H(z) data. A somewhat significant exception is the non-relativistic matter density parameter Ωm0 where QSO data favour Ωm0 ∼ 0.5 − 0.6 in most models. As a result, in joint analyses of QSO data with H(z) + BAO data the 1D Ωm0 distributions shift slightly towards larger values. A joint analysis of the QSO + BAO + H(z) data is consistent with the current standard model, spatially-flat ΛCDM, but mildly favours closed spatial hypersurfaces and dynamical dark energy. Since the higher Ωm0 values favoured by QSO data appear to be associated with the z ∼ 2 − 5 part of these data, and conflict somewhat with strong indications for Ωm0 ∼ 0.3 from most z < 2.5 data as well as from the cosmic microwave background anisotropy data at z ∼ 1100, in most models, the larger QSO data Ωm0 is possibly more indicative of an issue with the z ∼ 2 − 5 QSO data than of an inadequacy of the standard flat ΛCDM model.

79 ASTRONOMY AND ASTROPHYSICS↗

Translytics: A Novel Approach for Runtime Selection of Database Layout Based on User’s Context

Currently, organizations have to maintain separate systems for transactions and analytics inside the company. Notably, different vendors provide the capabilities for either of these tasks that require specialized hardware or software. Data engineers are required to retrieve data from one source and transform it into another format to obtain the maximum benefit in a minimum time. Organizations strive for a competitive advantage that is achieved by fetching data from their customers and getting insights earliest for timely decision making. Present practices do not permit the view of the latest data for analytics since the first data have to be fetched from the source, transformed, and loaded to other systems to be utilized for analysis by relevant teams. This paper introduces a single system for both transactions and analytics. Our proposed solution would permit companies to seamlessly adapt our solution without the need to shift all of their data to newer systems and allow all the teams. It would grant all the teams to have a view of the latest available data without extra expertise and budget.

Tanvir, Muhammad Makhshif↗

A Data-Agnostic, Continuous Machine Learning Framework for Application in High Energy Physics and Beyond: Phase 1 Final Scientific/Technical Report

This Phase 1 effort has focused on the development of continual learning frameworks for use in machine learning, specifically in the applied context of High Energy Physics (HEP). Machine learning (ML) is a transformative technology by which computers, typically through the use of neural networks, are able to perform tasks with proficiency that rivals or surpasses that of human users. Model Degradation & Catastrophic Forgetting are two undesired phenomena which can occur in ML where the performance of a model degrades when either deployed on novel data streams, or trained on novel data which are sufficiently different than the data the models were initially trained on. A natural example where these sorts of effects can be observed is in the performance of detectors in harsh environments, where the detector signature may change over the lifetime of the detector as it ages and deteriorates — precisely what occurs in the experiments conducted in HEP. Real world HEP data is therefore an excellent test-ground and use-case for Continual Learning paradigms, which are techniques used in ML to counteract these problems. Ensemble learning is one such technique, where multiple smaller models are trained on subsets of the overall data and are ensembled together during inference. The intuition behind this technique is that, although there are shifts in the distributions which govern the incoming data streams, these shifts are not expected to be homogeneous or global. If a sufficient diversity in solutions within the various sub-models has been achieved, then at least one sub-model is expected to retain its performance within the overall ensemble. One further strength of this approach is that the architectures of the various models do not need to be identical, and in fact even different modalities of data can naturally be combined in this way. This work focused on applying ensemble learning techniques to derive results using two main datasets, anomaly detection in HEP data & time-series forecasting in semiconductor manufacturing data. Semiconductor manufacturing involves data with surprising similarity to that of HEP (e.g. wafer maps look very similar to digi-occupancy maps) and Cerium Lab’s prominence within the semiconductor industry makes semiconductor manufacturing a natural opportunity for commercialization of this work. Our efforts have led to two strong results. The first is that we evaluated the proposed ensembling techniques using previously proposed machine learning architectures for use in anomaly detection, namely AutoEncoder based models and their derivatives. We also developed new architectures which have not been evaluated in this context before. In fact, this work marks the first use of Vision Transformers for anomaly detection in HEP. Second, we demonstrated that ensemble learning significantly improves model performance in scenarios prone to degradation, validating its effectiveness across both HEP and semiconductor datasets. These results further support ensemble learning as a powerful strategy for mitigating catastrophic forgetting and maintaining robust performance in evolving data environments.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Preparing Earth Data Scientists for 'The Sexiest Job of the 21st Century'

What Exactly do Earth Data Scientists do, and What do They Need to Know, to do It? There is not one simple answer, but there are many complex answers. Data Science, and data analytics, are new and nebulas, and takes on different characteristics depending on: The subject matter being analyzed, the maturity of the research, and whether the employed subject specific analytics is descriptive, diagnostic, discoveritive, predictive, or prescriptive, in nature. In addition, in a, thus far, business driven paradigm shift, university curriculums teaching data analytics pertaining to Earth science have, as a whole, lagged behind, andor have varied in approach.This presentation attempts to breakdown and identify the many activities that Earth Data Scientists, as a profession, encounter, as well as provide case studies of specific Earth Data Scientist and data analytics efforts. I will also address the educational preparation, that best equips future Earth Data Scientists, needed to further Earth science heterogeneous data research and applications analysis. The goal of this presentation is to describe the actual need for Earth Data Scientists and the practical skills to perform Earth science data analytics, thus hoping to initiate discussion addressing a baseline set of needed expertise for educating future Earth Data Scientists.

data analytics↗

Effects of 1997-1998 El Nino on Tropospheric Ozone and Water Vapor

This paper analyzes the impact of the 1997-1998 El Nino on tropospheric column ozone and tropospheric water vapor derived respectively from the Total Ozone Mapping Spectrometer (TOMS) on Earth Probe and the Microwave Limb Scanning instrument on the Upper Atmosphere Research Satellite. The 1997-1998 El Nino, characterized by an anomalous increase in sea-surface temperature (SST) across the eastern and central tropical Pacific Ocean, is one of the strongest El Nino Southern Oscillation (ENSO) events of the century, comparable in magnitude to the 1982-1983 episode. The major impact of the SST change has been the shift in the convection pattern from the western to the eastern Pacific affecting the response of rain-producing cumulonimbus. As a result, there has been a significant increase in rainfall over the eastern Pacific and a decrease over the western Pacific and Indonesia. The dryness in the Indonesian region has contributed to large-scale burning by uncontrolled wildfires in the tropical rainforests of Sumatra and Borneo. Our study shows that tropospheric column ozone decreased by 4-8 Dobson units (DU) in the eastern Pacific and increased by about 10-20 DU in the western Pacific largely as a result of the eastward shift of the tropical convective activity as inferred from National Oceanic and Atmospheric Administration (NOAA) outgoing longwave radiation (OLR) data. The effect of this shift is also evident in the upper tropospheric water vapor mixing ratio which varies inversely as ozone (O3). These conclusions are qualitatively consistent with the changes in atmospheric circulation derived from zonal and vertical wind data obtained from the Goddard Earth Observing System data assimilation analyses. The changes in tropospheric column O3 during the course of the 1997-1998 El Nino appear to be caused by a combination of large-scale circulation processes associated with the shift in the tropical convection pattern and surface/boundary layer processes associated with forest fires in the Indonesian region.

Chandra, S.↗

Accessing numeric data via flags and tags: A final report on a real world experiment

An experiment is reported which: extended the concepts of data flagging and tagging to the aerospace scientific and technical literature; generated experience with the assignment of data summaries and data terms by documentation specialists; and obtained real world assessments of data summaries and data terms in information products and services. Inclusion of data summaries and data terms improved users' understanding of referenced documents from a subject perspective as well as from a data perspective; furthermore, a radical shift in document ordering behavior occurred during the experiment toward proportionately more requests for data-summarized items.

Kottenstette, J. P.↗

Creating User-Friendly Tools for Data Analysis and Visualization in K-12 Classrooms: A Fortran Dinosaur Meets Generation Y

During the summer of 2007, as part of the second year of a NASA-funded project in partnership with Christopher Newport University called SPHERE (Students as Professionals Helping Educators Research the Earth), a group of undergraduate students spent 8 weeks in a research internship at or near NASA Langley Research Center. Three students from this group formed the Clouds group along with a NASA mentor (Chambers), and the brief addition of a local high school student fulfilling a mentorship requirement. The Clouds group was given the task of exploring and analyzing ground-based cloud observations obtained by K-12 students as part of the Students' Cloud Observations On-Line (S'COOL) Project, and the corresponding satellite data. This project began in 1997. The primary analysis tools developed for it were in FORTRAN, a computer language none of the students were familiar with. While they persevered through computer challenges and picky syntax, it eventually became obvious that this was not the most fruitful approach for a project aimed at motivating K-12 students to do their own data analysis. Thus, about halfway through the summer the group shifted its focus to more modern data analysis and visualization tools, namely spreadsheets and Google(tm) Earth. The result of their efforts, so far, is two different Excel spreadsheets and a Google(tm) Earth file. The spreadsheets are set up to allow participating classrooms to paste in a particular dataset of interest, using the standard S'COOL format, and easily perform a variety of analyses and comparisons of the ground cloud observation reports and their correspondence with the satellite data. This includes summarizing cloud occurrence and cloud cover statistics, and comparing cloud cover measurements from the two points of view. A visual classification tool is also provided to compare the cloud levels reported from the two viewpoints. This provides a statistical counterpart to the existing S'COOL data visualization tool, which is used for individual ground-to-satellite correspondences. The Google(tm) Earth file contains a set of placemarks and ground overlays to show participating students the area around their school that the satellite is measuring. This approach will be automated and made interactive by the S'COOL database expert and will also be used to help refine the latitude/longitude location of the participating schools. Once complete, these new data analysis tools will be posted on the S'COOL website for use by the project participants in schools around the US and the world.

Chambers, L. H.↗

EMU Lessons Learned Database

As manned space exploration takes on the task of traveling beyond low Earth orbit, many problems arise that must be solved in order to make the journey possible. One major task is protecting humans from the harsh space environment. The current method of protecting astronauts during Extravehicular Activity (EVA) is through use of the specially designed Extravehicular Mobility Unit (EMU). As more rigorous EVA conditions need to be endured at new destinations, the suit will need to be tailored and improved in order to accommodate the astronaut. The Objective behind the EMU Lessons Learned Database(LLD) is to be able to create a tool which will assist in the development of next-generation EMUs, along with maintenance and improvement of the current EMU, by compiling data from Failure Investigation and Analysis Reports (FIARs) which have information on past suit failures. FIARs use a system of codes that give more information on the aspects of the failure, but if one is unfamiliar with the EMU they will be unable to decipher the information. A goal of the EMU LLD is to not only compile the information, but to present it in a user-friendly, organized, searchable database accessible to all familiarity levels with the EMU; both newcomers and veterans alike. The EMU LLD originally started as an Excel database, which allowed easy navigation and analysis of the data through pivot charts. Creating an entry requires access to the Problem Reporting And Corrective Action database (PRACA), which contains the original FIAR data for all hardware. FIAR data are then transferred to, defined, and formatted in the LLD. Work is being done to create a web-based version of the LLD in order to increase accessibility to all of Johnson Space Center (JSC), which includes converting entries from Excel to the HTML format. FIARs related to the EMU have been completed in the Excel version, and now focus has shifted to expanding FIAR data in the LLD to include EVA tools and support hardware such as the Pistol Grip Tool (PGT) and the Battery Charger Module (BCM), while adding any recently closed EMU-related FIARs.

Matthews, Kevin M., Jr.↗

High dimensional binary classification under label shift: phase transition and regularization

Label Shift has been widely believed to be harmful to the generalization performance of machine learning models. Researchers have proposed many approaches to mitigate the impact of the label shift, e.g., balancing the training data. However, these methods often consider the underparametrized regime, where the sample size is much larger than the data dimension. The research under the overparametrized regime is very limited. Here, to bridge this gap, we propose a new asymptotic analysis of the Fisher Linear Discriminant classifier for binary classification with label shift. Specifically, we prove that there exists a phase transition phenomenon: Under certain overparametrized regime, the classifier trained using imbalanced data outperforms the counterpart with reduced balanced data. Moreover, we investigate the impact of regularization to the label shift: The aforementioned phase transition vanishes as the regularization becomes strong.

binary classification↗

Assessing the ASME Section III, Division 5, Class A Primary Load Design Rules Against Creep Notch Effects

This report assesses the ASME Section III, Division 5, Subsection HB, Subpart B rules covering the design and construction of high temperature Class A nuclear reactor components for their robustness against creep notch effects. The creep notch effect combines the effect of multiaxial stresses on material creep deformation and damage. This report considers both effects, including the potential for creep mechanism shifts affecting the extrapolation of design data from high stress, short-time experimental data to low stress, long time operating component conditions. The report summarizes the history of and literature on multiaxial creep and surveys the current ASME design rules dealing with multiaxial effects. The report then describes the results of two dedicated numerical studies, one assessing the robustness of the ASME rules against uncertainty in extrapolating from uniaxial creep test data to multiaxial component conditions and a second study examining the potential effects of a mechanism shift from high stress, dislocation-mediated creep to diffusion-dominated creep at lower stresses. The final conclusion of the report is that the ASME rules adequately guard against multiaxial creep failure, though there are several aspects of the Code that could be optimized to provide less over-conservative design predictions or to provide a more consistent design margin as a function of temperature and stress. In a few areas, particularly the potential for mechanism shifts outside the currently-available experimental data, the Code rules should be revaluated as additional experimental data and new modeling and simulation results become available

22 GENERAL STUDIES OF NUCLEAR REACTORS↗