Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data shift”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Machine learning for domain transfer between simulated and experimental 2D X-ray diffraction patterns using generative adversarial networks

X-ray diffraction (XRD) is a well-established technique for analyzing materials at an atomic level. Dynamic compression experiments (DCE), in which materials are subject to extreme pressures, can provide fundamental understanding to pressure-induced phase transitions and compression of the crystal lattice. The analysis of XRD patterns from highly compressed samples is non-trivial given the sparsity of data, high experimental costs, and the fact that the data is often marred with X-ray background and other artifacts. While accurate computational frameworks exist, they solve the forward problem—from structures and orientations to XRD patterns. Solving the inverse problem for 2D experimental diffraction patterns is currently a complex manual process of matching and comparing experimentally observed patterns to computationally generated ones. Machine learning is a promising tool for automating the matching process but often requires data-intensive architectures. Here, in this study, we use a CycleGAN to translate the domain of limited experimental data to a domain in which there is readily available simulated data. This domain shift allows data-intensive machine learning models that have only been trained on simulated XRD patterns to be used in the analysis of experiments.

Brozak, Samantha Jean [Sandia National Laboratorie↗

Dynamic Testing of a Commercial FRAM Device Under Gamma Ray Dose and Neutron Beam

Here, this work presents dose rate and neutron testing on a ferroelectric random-access memory (FRAM) device in dynamic operation during radiation pulses. Radiation failure modes are shown to be unique. Dose rate failure is shown to be based primarily on integrated dose in the radiation pulse and terminates the FRAM operation and can alter up to 8 bytes of data in storage. Neutron failures are attributed to a shifting of data like a clocking error.

Harris, Nathan↗

Data and script associated with “Shifts in Rain-Snow Partitioning Drive Faster Water Transit Times in the US Pacific Northwest”

This data package contains the data and code to use and run the Water Tracer enabled version of the Weather Research and Forecasting Hydrologic model (WT-WRF-Hydro) with the Sequential Precipitation Input Tagging (SPIT) framework. It is associated with the publication “Shifts in Rain-Snow Partitioning Drive Faster Water Transit Times in the US Pacific Northwest” published in Scientific Reports (Butler et al., 2026; https://doi.org/10.1038/s41598-026-46539-1). We use the Continental U.S. (CONUSII; Rasmussen et al., 2021) dataset to force the model with an historical climate (2006–2013) and a future climate (2086–2093) with a representative carbon pathway (RCP) 8.5 scenario. We use the model to calculate water transit times in five headwater catchments within the U.S. Pacific Northwest. We also show key hydrologic and environmental variables that affect water transit times and changes in the future. Finally, we use observed data to validate the model such as stream water isotopes, snowpack characteristics, and stream discharge. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. In addition to a readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes all column/row headers and variable definitions. The data package consists of 11 folders: (1) "Figures" contains the exported figures used in the manuscript; (2) "Model_Isotope_Date" contains the WT-WRF-Hydro isotope date used in model validation; (3) “Model_Outputs_Future” contains the WT-WRF-Hydro future climate outputs; (4) “Model_Outputs_Historical” contains the WT-WRF-Hydro historical climate outputs; (5) “Model_Outputs_Weights_Areas” contains the WT-WRF-Hydro weights per catchment used to calculate water transit times and isotopes in stream water; (6) “MODIS_data_scripts” contains data used to validate snow conditions in the study area; (7) “Observed_Flow_Data” contains the observed streamflow data used in model validation; (8) “Observed_Isotope_Data” contains the observed stream water isotope data used in model validation; (9) “Scripts” contains the Python scripts used to general results and the figures; (10) “Statistic_Outputs” contains the water transit time statistical outputs reported in this manuscript; (11) “Validation_SNOTEL” contains the SNOTEL data used in model validation. The files in this data package have the following file extensions: .tif, .txt, .csv, .pdf, .py, .jpg, and .png.

American River↗

Spectral Properties of Globally Distributed ENA Fluxes across Diverse Regions of the Heliosphere

This study analyzes energetic neutral atom (ENA) spectral properties across distinct regions of globally distributed flux (GDF) sky maps, using Interstellar Boundary Explorer data from a full solar cycle, corrected for time dispersion. By time-shifting the data to the heliosheath using GDF source distances from D. B. Reisenfeld et al., we achieve a more accurate representation of heliosheath GDF energy spectra. We quantify ENA spectral characteristics, heliosheath line-of-sight-integrated proton pressure, and heliosheath proton temperature, comparing these to solar wind properties at 1 au and interplanetary scintillation-derived solar wind data. Our findings show that the spectral index is generally anticorrelated with heliosheath proton temperature and pressure, except in the central tail, where a partial positive correlation is observed. The lowest spectral index values occur when high-latitude heliosheath regions are dominated by fast solar wind from polar coronal holes. The south pole exhibits the flattest energy spectra due to plasma heating from both fast solar wind and a late-2014 pressure pulse. The central tail shows shorter variability (5–6 yr) for spectral index and heliosheath proton temperature, while proton pressure follows the 11 yr solar cycle. Most spectral shapes exhibit a “knee” distribution, peaking during solar maximum, with an “ankle” shape observed only at the south pole during solar cycle transitions. Asymmetry in proton pressure in the lobes is driven by the draping effect of the local interstellar magnetic field. This study provides insights into the energetic properties of GDF across the heliosphere, enhancing our understanding of the heliospheric environment.

79 ASTRONOMY AND ASTROPHYSICS↗

Out-of-Distribution Detection and Radiological Data Monitoring Using Statistical Process Control

Abstract Machine learning (ML) models often fail with data that deviates from their training distribution. This is a significant concern for ML-enabled devices as data drift may lead to unexpected performance. This work introduces a new framework for out of distribution (OOD) detection and data drift monitoring that combines ML and geometric methods with statistical process control (SPC). We investigated different design choices, including methods for extracting feature representations and drift quantification for OOD detection in individual images and as an approach for input data monitoring. We evaluated the framework for both identifying OOD images and demonstrating the ability to detect shifts in data streams over time. We demonstrated a proof-of-concept via the following tasks: 1) differentiating axial vs. non-axial CT images, 2) differentiating CXR vs. other radiographic imaging modalities, and 3) differentiating adult CXR vs. pediatric CXR. For the identification of individual OOD images, our framework achieved high sensitivity in detecting OOD inputs: 0.980 in CT, 0.984 in CXR, and 0.854 in pediatric CXR. Our framework is also adept at monitoring data streams and identifying the time a drift occurred. In our simulations tracking drift over time, it effectively detected a shift from CXR to non-CXR instantly, a transition from axial to non-axial CT within few days, and a drift from adult to pediatric CXRs within a day—all while maintaining a low false positive rate. Through additional experiments, we demonstrate the framework is modality-agnostic and independent from the underlying model structure, making it highly customizable for specific applications and broadly applicable across different imaging modalities and deployed ML models.

Zamzmi, Ghada↗

Phylogenomics and the first higher taxonomy of Placozoa, an ancient and enigmatic animal phylum

Placozoa is an ancient phylum of extraordinarily unusual animals: miniscule, ameboid creatures that lack most fundamental animal features. Despite high genetic diversity, only recently have the second and third species been named. While prior genomic studies suffer from incomplete placozoan taxon sampling, we more than double the count with protein sequences from seven key genomes and produce the first nuclear phylogenomic reconstruction of all major placozoan lineages. This leads us to the first complete Linnaean taxonomic classification of Placozoa, over a century after its discovery: This may be the only time in the 21st century when an entire higher taxonomy for a whole animal phylum is formalized. Our classification establishes 2 new classes, 4 new orders, 3 new families, 1 new genus, and 1 new species, namely classes Polyplacotomia and Uniplacotomia; orders Polyplacotomea, Trichoplacea, Cladhexea, and Hoilungea; families Polyplacotomidae, Cladtertiidae, and Hoilungidae; and genus Cladtertia with species Cladtertia collaboinventa, nov. Our likelihood and gene content tree topologies refine the relationships determined in previous studies. Adding morphological data into our phylogenomic matrices suggests sponges (Porifera) as the sister to other animals, indicating that modest data addition shifts this node away from comb jellies (Ctenophora). Furthermore, by adding the first genomic protein data of the exceptionally distinct and branching Polyplacotoma mediterranea , we solidify its position as sister to all other placozoans; a divergence we estimate to be over 400 million years old. Yet even this deep split sits on a long branch to other animals, suggesting a bottleneck event followed by diversification. Ancestral state reconstructions indicate large shifts in gene content within Placozoa, with Hoilungia hongkongensis and its closest relatives having the most unique genetics.

54 ENVIRONMENTAL SCIENCES↗

Emerging Technologies for Privacy Preservation in Energy Systems

This study explores the intersection of digitalization and privacy within the energy sector, focusing on the emerging challenges and opportunities presented by integrating Distributed Energy Resources (DERs) and advanced metering infrastructure. The need for robust digital privacy measures has become crucial as the energy industry evolves towards a more decentralized, digitalized, and decarbonized future. This study delves into four cutting-edge privacy-preserving technologies—Homomorphic Encryption (HE), Secure Multiparty Computation (SMPC), Differential Privacy (DP), and Federated Learning (FL)—each offering unique solutions to safeguard consumer data by increasing digital connectivity and data exchange. Through a detailed examination of these methods, the study explains how each technology operates, its applications within the energy sector, and the specific privacy challenges it addresses. Homomorphic Encryption allows for secure computations on encrypted data, enabling data analysis without compromising privacy. Secure Multiparty Computation enables collaborative data analysis across different entities while protecting the confidentiality of the inputs. Differential Privacy introduces randomness into the assembled data set, preventing the identification of individual records in statistical databases. Lastly, Federated Learning offers a paradigm shift in data analysis, where machine learning models are trained at the edge, minimizing the centralization of sensitive data. The research underscores the significance of implementing these privacy-enhancing technologies to comply with strict data protection regulations, foster consumer trust, and enhance the security of the energy infrastructure. By providing a comprehensive overview of these methodologies and their practical implications for the energy sector, this study aims to contribute to the ongoing discourse on digital privacy, offering insights into how the energy industry can navigate the complexities of data privacy in the digital age.

Cali, Umit↗

DOE COVID-19 Data Curation Effort: Overview of Initial Data Collection Coverage (March - June 2020)

During the COVID-19 pandemic of 2020, major case reporting outlets quickly coalesced around two or three primary vendors. Johns Hopkins University and The New York Times were among the more prominent, and all were of great value to the nation, particularly during the uncertain early stages of the pandemic. They primarily focused on three major attributes: number of new cases, deaths, and recovery, but only at the state level. Recognizing that many states were reporting very detailed data sets (e.g., hospital beds) at a count level or finer, the ORNL Pandemic Modeling team embarked on a major data curation effort from March to June 2020 for the purpose of capturing this wealth of detailed data. The challenge of curating this data was daunting. The number of attributes reported by the states grew on almost on a weekly basis. States were routinely shifting their web tool strategies away from easily parsable HTML-based formatting to new Tableau and ArcGIS content. This growth in the sheer number of attributes combined with the unpredictable shifts in data format meant an aggressive and agile combination of automated scripting and manual scraping was required to capture new daily streams. To keep up, the team had to scale up staff and widen its approach for capture and storage. The DOE COVID-19 data collection effort resulted in over 11 million data points being collected, covering over 13,000 unique geographies and over 2,000 unique attributes that spanned predominantly from early March through the end of June 2020.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

High-resolution laser system for the S 3 -Low Energy Branch

Here, in this paper we present the first high-resolution laser spectroscopy results obtained at the GISELE laser laboratory of the GANIL-SPIRAL2 facility, in preparation for the first experiments with the S 3 -Low Energy Branch. Studies of neutron-deficient radioactive isotopes of erbium and tin represent the first physics cases to be studied at S 3 . The measured isotope-shift and hyperfine structure data are presented for stable isotopes of these elements. The erbium isotopes were studied using the 4f 12 6s 2 3 H 6 → 4f 12 ( 3 H)6s6p J = 5 atomic transition (415 nm) and the tin isotopes were studied by the 5s 2 5p 2 ( 3 P 0 ) → 5s 2 5p6s( 3 P 1 ) atomic transition (286.4 nm), and are used as a benchmark of the laser setup. Additionally, the tin isotopes were studied by the 5s 2 5p6s( 3 P 1 ) → 5s 2 5p6p( 3 P 2 ) atomic transition (811.6 nm), for which new isotope-shift data was obtained and the corresponding field-shift F 812 and mass-shift M 812 factors are presented.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

A customizable data management framework for high-repetition-rate high-energy-density science

The high-energy-density (HED) physics community is moving toward a new paradigm of high-repetition-rate (HRR) operation. To fully leverage the scientific power of HRR HED facilities, all of the components of each subsystem (laser, targetry, and performance diagnostics) must be connected and synchronized in a reliable and robust manner while the data acquired are tagged and archived in real time. To this end, GA has begun developing a generalized NoSQL-database framework, the MongoDB repository for information and archiving. An organizational strategy has been developed that shifts HED data organization from a shot-based to a diagnostic-based approach in order to increase archival and retrieval efficiency that lends itself to optimization applications. This work is a first step in pushing HRR HED science toward data management solutions that emphasize machine actionability and aim to stimulate community engagement to define data standards in HED science.

Instruments & Instrumentation↗

Counter Data Paucity through Adversarial Invariance Encoding: A Case Study on Modeling Battery Thermal Runaway

Lithium-ion batteries, widely used for their durability and high energy storage, face the risk of internal short circuits leading to catastrophic thermal runaway events. These events, triggered by external stimuli like mechanical loads, pose safety concerns in applications such as electric vehicles. Detecting and understanding thermal runaway events is crucial, but physics-driven models struggle to explain the non-linear evolution of battery temperature during these events, considering factors like material composition and state-of-charge. Due to the rarity of these events and the cost of data collection, we propose a deep learning (DL) model to predict battery temperature responses during thermal runaway. The challenge lies in the scarcity of data, making traditional DL models prone to overfitting and learning low-quality representations of the complex process.Our approach introduces a novel few-shot architecture that incorporates an adversarially governed invariant encoding process. This architecture aims to distill "invariant" relationships by addressing distributional shifts in data across various battery properties, facilitating the detection of thermal runaway events. Specifically, our results demonstrate that deep learning models conditioned on these "invariant" representations outperform state-of-the-art baselines, achieving a remarkable 96.8% performance improvement in terms of the popular metric MAPE. This framework presents a promising direction for enhancing battery safety modeling, particularly in the context of rare and complex events like thermal runaway. Our code and code and dataset used for the paper are public1.

Tabassum, Anika [ORNL] (ORCID:0000000254600955)↗

Materials data science using CRADLE: A distributed, data-centric approach

Abstract There is a paradigm shift towards data-centric AI, where model efficacy relies on quality, unified data. The common research analytics and data lifecycle environment (CRADLE™) is an infrastructure and framework that supports a data-centric paradigm and materials data science at scale through heterogeneous data management, elastic scaling, and accessible interfaces. We demonstrate CRADLE’s capabilities through five materials science studies: phase identification in X-ray diffraction, defect segmentation in X-ray computed tomography, polymer crystallization analysis in atomic force microscopy, feature extraction from additive manufacturing, and geospatial data fusion. CRADLE catalyzes scalable, reproducible insights to transform how data is captured, stored, and analyzed. Graphical abstract

97 MATHEMATICS AND COMPUTING↗

In situ laser profilometry for material segmentation and digital reconstruction of a multicomponent additively manufactured part

In addition to its ability to produce geometrically complex parts, additive manufacturing offers a unique opportunity to collect data about a component while it is being fabricated. However, there has only been limited effort to characterize parts morphologically and compositionally in situ. In this article, we present a layer-by-layer, laser profilometry-based in situ characterization technique as a method to digitally reconstruct a multi-material part. Data collected by the laser profilometer yields height maps and grayscale images which are voxelized using purpose-built software to volumetrically reconstruct the part. Additionally, the same part was also analyzed using X-ray computed tomography (CT) which was not able to resolve the different compositional regions within the part, but captured the filament morphology. The part was then bisected to compare the digital reconstruction to the actual part morphology and composition. Overall, the digital reconstruction was in good agreement with both the CT and bisected images. Deviations between the digital reconstruction and the CT/bisected images are likely the result of image segmentation settings or material shifts after data was collected. The in situ characterization method demonstrated here sets the stage for real time process monitoring and paves the way for additively manufactured parts that are “born qualified.”

36 MATERIALS SCIENCE↗

Thermodynamic control on the decomposition of organic matter across different electron acceptors

The increasing availability of high-resolution characterization of natural organic matter (OM) data has shifted the paradigm of lumped descriptions of OM components and potential microbial activities. Our recent development of a substrate-explicit thermodynamic model uniquely enables incorporating complex OM pools to formulate biogeochemical reaction models based on their elemental compositions. While this previous work facilitates prediction of aerobic respiration of complex OM, it is equally imperative to consider the role of non-oxygenic electron acceptors in regulating OM turnover and the fate of carbon. In this study, we significantly expand our previous model by flexibly incorporating both detailed OM chemistry and electron acceptors other than oxygen. Here, our modeling analysis has revealed substantial variations in the energy status of OM molecules across different soils, which drive the co-occurrence of different electron-accepting processes. We demonstrated the effectiveness of the proposed model using a consistency check with experimental data. Through systematic evaluation of the impact of diverse chemical inputs (both electron donors and acceptors) on OM decomposition, the new model also revealed how key microbial growth parameters such as carbon use efficiency (CUE) and reaction rates vary across different electron-accepting processes. Our model provides a unified framework integrating thermodynamic and kinetic constraints on microbial metabolic activities. It complements traditional kinetic models, which are often designed solely to capture mass fluxes. We conclude that thermodynamic modeling emerges as a powerful tool for describing the mechanisms underlying the interplay between microbial growth and OM chemistry and cycling across different electron acceptors, enhancing our ability to project complex ecosystem behaviors in dynamic environments.

59 BASIC BIOLOGICAL SCIENCES↗

The future low-temperature geochemical data-scape as envisioned by the U.S. geochemical community

Data sharing benefits the researcher, the scientific community, and the public by allowing the impact of data to be generalized beyond one project and by making science more transparent. However, many scientific communities have not developed protocols or standards for publishing, citing, and versioning datasets. One community that lags in data management is that of low-temperature geochemistry (LTG). This paper resulted from an initiative from 2018 through 2020 to convene LTG and data scientists in the U.S. to strategize future management of LTG data. Through webinars, a workshop, a preprint, a townhall, and a community survey, the group of U.S. scientists discussed the landscape of data management for LTG – the data-scape. Currently this data-scape includes a “street bazaar” of data repositories. This was deemed appropriate in the same way that LTG scientists publish articles in many journals. The variety of data repositories and journals reflect that LTG scientists target many different scientific questions, produce data with extremely different structures and volumes, and utilize copious and complex metadata. Nonetheless, the group agreed that publication of LTG science must be accompanied by sharing of data in publicly accessible repositories, and, for sample-based data, registration of samples with globally unique persistent identifiers. LTG scientists should use certified data repositories that are either highly structured databases designed for specialized types of data, or unstructured generalized data systems. Recognizing the need for tools to enable search and cross-referencing across the proliferating data repositories, the group proposed that the overall data informatics paradigm in LTG should shift from “build data repository, data will come” to “publish data online, cybertools will find”. Funding agencies could also provide portals for LTG scientists to register funded projects and datasets, and forge approaches that cross national boundaries. Finally, the needed transformation of the LTG data culture requires emphasis in student education on science and management of data.

58 GEOSCIENCES↗

EmissionsFlexibility

Software to calculate the how changes in electricity consumption change the generator which provides the electricity. This is applied to find how load shifting (e.g. data centers) impact the carbon emissions of the power grid.

Rhodes, Noah↗

Analyzing Complex Energy Security Systems [Slides]

Motivation: DOE is investing in our technology for improving Energy Security; Many of these problems are grand challenges requiring moonshot type efforts; The geoscience paradigm is shifting from data sparse to data rich requiring us to take advantage of the latest computational and AI to tools optimize these systems.

54 ENVIRONMENTAL SCIENCES↗

Pooled Rideshare in the U.S.: An Exploratory Study of User Preferences

Pooled ridesharing offers on-demand, one-way, cost-effective transportation for passengers traveling in similar directions via a shared vehicle ride with others they do not know. Despite its potential benefits, the adoption of pooled rideshare remains low in the United States. This exploratory study aims to evaluate potential service improvements and features that may increase users’ willingness to adopt the service. The study analyzed transportation behaviors, rideshare preferences, and willingness to adopt pooled rideshare services among 8296 U.S. participants in 2025, building on findings from a 2021 nationwide survey of 5385 U.S. participants. The study incorporated 77 actionable items developed from the results of the 2021 survey to assess whether addressing specific user-generated topics such as safety, reliability, convenience, and privacy can improve pooled rideshare use. A side-by-side comparison of the 2021 and 2025 data revealed shifts in transportation behavior, with personal rideshare usage increasing from 22% to 28%, public transportation from 21% to 27%, and pooled rideshare from 6% to 8%, while personal vehicle (79%) use remained dominant. Participants rated features such as driver verification (94%), vehicle information (93%), peak time reliability (93%), and saving time and money (92–93%) as most important for improving rideshare services. A pre-to-post analysis of willingness to use pooled rideshare utilizing the actionable items as per respondents’ preferences showed improvement: “definitely will” increased from 15.9% to 20.1% and “probably will” rose from 35.6% to 47.7%. These results suggest that well-targeted service improvements may meaningfully enhance pooled rideshare acceptance. This study offers practical guidance for Transportation Network Companies (TNCs) and policymakers aiming to improve pooled rideshare as well as potential future research opportunities.

Transportation Network Companies (TNCs)↗