Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “heterogeneous data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Integration of Condition-Based, Diagnostic, Prognostic, And Anomaly Detection Data into Reliability Models to Support a Predictive Maintenance Context

Reliability data employed in plant reliability models are an approximated integral representation of the past industrywide operational experience, and they neglect the present asset health status (available, for example, from online monitoring data and diagnostic assessments) and forecasted health projection (when available from prognostic models). Ideally, in a predictive maintenance context, system reliability models should support decision making by propagating actual health information from the asset to the system level in order to provide a quantitative snapshot of system health and identify the most critical assets. Asset health should be informed solely by that specific asset’s current and historical performance data and should not be an approximated integral representation of the past industrywide operational experience (as currently performed by system reliability models through Bayesian updating processes). This paper proposes a reliability modeling approach that relies on asset diagnostic and prognostic assessments, along with monitoring data to measure asset health. We show how state-of-the art condition-based, diagnostic, prognostic, and anomaly detection models can be linked to system reliability models not in probability terms, but in terms of margin where margin is defined as the “distance” between the present status and an undesired event (e.g., failure or unacceptable performance). Then, we show how the propagation of margin data from the asset to the system level is performed through classical reliability models such as fault trees or reliability block diagrams. The described method is in fact able to propagate heterogenous health data from the asset to the system level in order to analytically assess system health.

97 MATHEMATICS AND COMPUTING↗

A unified large language model–based framework for heterogeneous PV image diagnosis

With advances in imaging technologies, modern photovoltaic (PV) systems generate large volumes of heterogeneous image data, including visible, electroluminescence (EL), and infrared (IR) images. Existing PV image analysis models, particularly deep learning approaches, are typically task-specific and lack cross-modality generalization. To address this limitation, this paper proposes an open-source large language model (LLM)–based unified framework for heterogeneous PV image diagnostics. Through task-aware diagnostic prompting, the framework enables analysis of visible, EL, and IR images within a single pipeline, supporting both zero-shot and few-shot inference and binary and multiclass classification. It is compatible with state-of-the-art multimodal LLMs, including ChatGPT, Gemini, Claude, Qwen, and CLIP. The framework is evaluated on PV module condition classification (clean, soiling, snow, hail, and bird droppings) using visible images, cell crack detection using EL images, and hotspot detection using IR images. GPT-5.1 in few-shot mode achieves the best performance, with classification accuracy exceeding 97.3%. Open-source models such as Qwen and CLIP also deliver competitive results on visible images (around 90% accuracy), though their performance is more limited on EL and IR modalities. On the full ELPV dataset, the framework achieves 83.5% zero-shot accuracy, within 2.8% of the supervised CNN baseline, confirming scalability to larger benchmarks. Practical aspects such as reproducibility, response latency, and confidence estimation are systematically analyzed. The framework operates across PV image modalities without modality- or task-specific training, making it well suited as a rapid pre-screening tool to support downstream detailed diagnostics. A benchmark dataset of diverse labeled PV images is also released.

Li, Baojie↗

Data transfers for full core heterogeneous reactor high- fidelity multiphysics studies

Multiphysics simulations for nuclear reactor analysis are usually performed by resorting to operator splitting and fixed point iterations between single-physics solvers. This enables the separate solution of each physics, such as neutronics, fuel performance, and thermal hydraulics, on meshes tailored to the requirements of the respective numerical discretizations of the equations. As the equations are coupled, several fields must be transferred between single-physics solves. Projecting fields between meshes while preserving order of accuracy, conservation properties, and mapping non-overlapping geometries is a complex endeavor. This conference paper will present the transfers as implemented in MOOSE, which can handle arbitrary meshes, arbitrary mappings, conservation of integral quantities, and are made to scale with distributed simulations on both ends of the transfers. Their adequacy for advanced nuclear reactor multiphysics coupling is shown through examples and numerical studies.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Developing machine learning for heterogeneous catalysis with experimental and computational data

Machine learning techniques have emerged as a useful tool for identifying complex patterns and correlations in large datasets, such as associating catalyst performance to its physicochemical properties. In the heterogeneous catalysis communities, machine learning models have mostly been developed using high-throughput quantum chemistry calculations, with only a few case studies resulting in experimentally validated catalyst improvements. This limited success may be due to the use of simplified catalyst structures in computational studies and the lack of comprehensive experimental datasets. In this Review, we bring together studies integrating high-throughput approaches and machine learning for the advancement of solid heterogeneous catalysis, leveraging both experimental and computational data. We systematically analyze trends in the field, based on the descriptors used as model input and output; the materials, devices, or reactions investigated; the dataset size; and the overall achievements. Furthermore, for models reporting unitless R 2 values, we compare the performances based on these mentioned trends.

Computational chemistry↗

Quantifying spatial and vertical variations in soil C:N relationships in permafrost-affected landscapes

Permafrost regions are experiencing rapid changes that affect carbon (C) and nitrogen (N) cycles, with implications for vegetation dynamics and gas exchanges with the atmosphere. Soil C:N ratio is a key indicator of organic matter quality, yet spatial estimates of N stocks and C:N ratios lag behind those for C. We used quantile regression forests to compare direct and indirect digital soil mapping approaches for predicting soil C:N ratios at 0–30, 30–60, and 60–100 cm depths across a latitudinal transect in Alaska. The indirect approach – deriving C:N from separately predicted C and N stocks – outperformed direct mapping for the surface layer (0–30 cm), while direct mapping was marginally better at greater depths. However, prediction accuracy decreased with depth for both methods. Temperature and topography were the most important predictors. Both approaches overestimated low and underestimated high C:N ratios, with direct mapping showing greater bias. Our results underscore the challenges of modeling C:N ratios in heterogeneous, data-sparse permafrost soils, but also suggest that indirect mapping holds promise if supported by more extensive datasets.

54 ENVIRONMENTAL SCIENCES↗

ZTF SN Ia DR2: Improved SN Ia colors through expanded dimensionality with SALT3+

Context. Type Ia supernovae (SNe Ia) are a key probe in modern cosmology, as they can be used to measure luminosity distances at gigaparsec scales. Models of their light curves are used to project heterogeneous observed data onto a common basis for analysis. Aims. The SALT model currently used for SN Ia cosmology describes SNe as having two sources of variability, accounted for by a color parameter c , and a “stretch” parameter x 1 . We extend the model to include an additional parameter we label x 2 , to investigate the cosmological impact of currently unaddressed light-curve variability. Methods. We constructed a new SALT model, that we dub “SALT3+”. This model was trained by an improved version of the SALTshaker code, using training data combining a selection of the second data release of cosmological SNe Ia from the Zwicky Transient Facility and the existing SALT3 training compilation. Results. We find additional, coherent variability in supernova light curves beyond SALT3. Most of this variation can be described as phase-dependent variation in g − r and r − i color curves, correlated with a boost in the height of the secondary maximum in i -band. These behaviors correlate with spectral differences, particularly in line velocity. We find that fits with the existing SALT3 model tend to address this excess variation with the color parameter, leading to less informative measurements of supernova color. We find that neglecting the new parameter in light-curve fits leads to a trend in Hubble residuals with x 2 of 0.039 ± 0.005 mag, representing a potential systematic uncertainty. However, we find no evidence of a bias in current cosmological measurements. Conclusions. We conclude that extended SN Ia light-curve models promise mild improvement in the accuracy of color measurements, and corresponding cosmological precision. However, models with more parameters are unlikely to substantially affect current cosmological results.

Kenworthy, W. D. (ORCID:0000000251535983)↗

L-VISP: LSTM Visualization for Interpretable Symptom Prediction in Patient Cohorts

Symptom modelling in head and neck cancer is challenged by the complexity of heterogeneous patient data, leading to an interest in deep learning approaches. Although Long Short-Term Memory Networks (LSTMs) have shown great results in patient risk prediction, their low interpretability requires data modellers to collaborate with clinical experts to validate the results. We present L-VISP, a human–machine solution that uses visual analytics for LSTM modelling in clinical research. L-VISP uses custom visual encodings to make multiple LSTM variants interpretable, supporting a full range of analysis, from understanding model operations and evaluating performance to interpreting results in a clinical context. We evaluate L-VISP with data modellers and a clinical oncologist and present the takeaways from this multidisciplinary collaboration.

LSTM modeling↗

MLCommons Science Benchmarks

Benchmarks are a cornerstone of modern machine learning practice, providing standardized eval- uations that enable reproducibility, comparison, and scientific progress. Yet, as AI systems particularly deep learning models become increasingly dynamic, traditional static benchmarking approaches are losing their relevance. Models rapidly evolve in architecture, scale, and capability; datasets shift; and deployment contexts continuously change, creating a moving target for evaluation. Without adaptive benchmarking frame- works, both scientific assessment and real-world de- ployment risk becoming misaligned with actual system behavior. Drawing on our experience from MLCommons, educa- tional initiatives, and government programs such as the DOE s Million Parameter Consortium, we identify key barriers that hinder the broader adoption and utility of benchmarking in AI. These include substantial resource demands, limited access to specialized hardware, lack of expertise in benchmark design, and uncertainty among practitioners about how to relate benchmark results to their own application domains. Moreover, current benchmarks often emphasize peak performance on leadership-class hardware, offering limited guidance for more diverse, real-world deployment scenarios. We argue that benchmarking itself must become dy- namic in order to incorporate evolving models, updated data, and heterogeneous computational platforms while maintaining transparency, reproducibility, and inter- pretability. Democratizing this process requires not only technical innovation, but also systematic educational efforts spanning undergraduate to professional levels to develop sustained expertise in benchmark design and use. Finally, benchmarks should be framed and com- municated to support application-relevant comparisons, enabling both developers and users to make informed, context-sensitive decisions. Advancing dynamic and inclusive benchmarking practices will be essential to ensure that evaluation keeps pace with the evolving AI landscape and supports responsible, reproducible, and accessible AI deployment.

Hawks, Benjamin G. [Fermilab]↗

Opportunities in AI/ML for the Rubin LSST Dark Energy Science Collaboration

The Vera C. Rubin Observatory's Legacy Survey of Space and Time (LSST) will produce unprecedented volumes of heterogeneous astronomical data (images, catalogs, and alerts) that challenge traditional analysis pipelines. The LSST Dark Energy Science Collaboration (DESC) aims to derive robust constraints on dark energy and dark matter from these data, requiring methods that are statistically powerful, scalable, and operationally reliable. Artificial intelligence and machine learning (AI/ML) are already embedded across DESC science workflows, from photometric redshifts and transient classification to weak lensing inference and cosmological simulations. Yet their utility for precision cosmology hinges on trustworthy uncertainty quantification, robustness to covariate shift and model misspecification, and reproducible integration within scientific pipelines. This white paper surveys the current landscape of AI/ML across DESC's primary cosmological probes and cross-cutting analyses, revealing that the same core methodologies and fundamental challenges recur across disparate science cases. Since progress on these cross-cutting challenges would benefit multiple probes simultaneously, we identify key methodological research priorities, including Bayesian inference at scale, physics-informed methods, validation frameworks, and active learning for discovery. With an eye on emerging techniques, we also explore the potential of the latest foundation model methodologies and LLM-driven agentic AI systems to reshape DESC workflows, provided their deployment is coupled with rigorous evaluation and governance. Finally, we discuss critical software, computing, data infrastructure, and human capital requirements for the successful deployment of these new methodologies, and consider associated risks and opportunities for broader coordination with external actors.

Aubourg, Eric [APC, Paris] (ORCID:000000025592023X↗

AI Benchmark Democratization and Carpentry

Benchmarks are a cornerstone of modern machine learning, enabling reproducibility, comparison, and scientific progress. However, AI benchmarks are increasingly complex, requiring dynamic, AI-focused workflows. Rapid evolution in model architectures, scale, datasets, and deployment contexts makes evaluation a moving target. Large language models often memorize static benchmarks, causing a gap between benchmark results and real-world performance. Beyond traditional static benchmarks, continuous adaptive benchmarking frameworks are needed to align scientific assessment with deployment risks. This calls for skills and education in AI Benchmark Carpentry. From our experience with MLCommons, educational initiatives, and programs like the DOE's Trillion Parameter Consortium, key barriers include high resource demands, limited access to specialized hardware, lack of benchmark design expertise, and uncertainty in relating results to application domains. Current benchmarks often emphasize peak performance on top-tier hardware, offering limited guidance for diverse, real-world scenarios. Benchmarking must become dynamic, incorporating evolving models, updated data, and heterogeneous platforms while maintaining transparency, reproducibility, and interpretability. Democratization requires both technical innovation and systematic education across levels, building sustained expertise in benchmark design and use. Benchmarks should support application-relevant comparisons, enabling informed, context-sensitive decisions. Dynamic, inclusive benchmarking will ensure evaluation keeps pace with AI evolution and supports responsible, reproducible, and accessible AI deployment. Community efforts can provide a foundation for AI Benchmark Carpentry.

von Laszewski, Gregor [Virginia U.]↗

IPC-Fusion (Infrastructure Perception and Control (IPC): Multisensor Data Fusion Software) [SWR-25-153]

As part of the National Laboratory of the Rockies' (NLR’s) Infrastructure Perception and Control Laboratory, the IPC-Fusion toolkit provides a probabilistic, scalable, multi-sensor fusion framework that integrates (late-stage fusion) heterogeneous object detection data from traffic sensors to enable robust, real-time tracking of roadway occupants. The algorithmic design of the toolkit is motivated by the need for creating a digital twin of traffic at the edge in a scalable and affordable manner. The software operates by combining object-level measurements (such as position and velocity) from a suite of sensors (such as radar, lidar, camera) using Kalman filtering and probabilistic data association techniques to overcome individual sensor limitations and achieve superior tracking performance in complex traffic zones. The framework addresses key challenges including heterogeneous measurement uncertainties, asynchronous data streams, varying spatiotemporal data resolutions, robust data association, and adaptive object lifecycle management. Validated on real-world traffic intersection data including vehicles and pedestrians, IPC-Fusion demonstrates enhanced tracking reliability across scenarios involving occlusions, sensor failures, and varying traffic densities, supporting the broader IPC initiative's goal of transforming transportation infrastructure through advanced perception capabilities for intelligent transportation systems, traffic safety applications, and autonomous vehicle support.

Sandhu, Rimple [National Laboratory of the Rockies↗

Heterogeneous Multi-Domain Dataset Synthesis to Facilitate Privacy and Risk Assessments in Smart City IoT

The emergence of the Smart Cities paradigm and the rapid expansion and integration of Internet of Things (IoT) technologies within this context have created unprecedented opportunities for high-resolution behavioral analytics, urban optimization, and context-aware services. However, this same proliferation intensifies privacy risks, particularly those arising from cross-modal data linkage across heterogeneous sensing platforms. To address these challenges, this paper introduces a comprehensive, statistically grounded framework for generating synthetic, multimodal IoT datasets tailored to Smart City research. The framework produces behaviorally plausible synthetic data suitable for preliminary privacy risk assessment and as a benchmark for future re-identification studies, as well as for evaluating algorithms in mobility modeling, urban informatics, and privacy-enhancing technologies. As part of our approach, we formalize probabilistic methods for synthesizing three heterogeneous and operationally relevant data streams—cellular mobility traces, payment terminal transaction logs, and Smart Retail nutrition records—capturing the behaviors of a large number of synthetically generated urban residents over a 12-week period. The framework integrates spatially explicit merchant selection using K-Dimensional (KD)-tree nearest-neighbor algorithms, temporally correlated anchor-based mobility simulation reflective of daily urban rhythms, and dietary-constraint filtering to preserve ecological validity in consumption patterns. In total, the system generates approximately 116 million mobility pings, 5.4 million transactions, and 1.9 million itemized purchases, yielding a reproducible benchmark for evaluating multimodal analytics, privacy-preserving computation, and secure IoT data-sharing protocols. To show the validity of this dataset, the underlying distributions of these residents were successfully validated against reported distributions in published research. We present preliminary uniqueness and cross-modal linkage indicators; comprehensive re-identification benchmarking against specific attack algorithms is planned as future work. This framework can be easily adapted to various scenarios of interest in Smart Cities and other IoT applications. By aligning methodological rigor with the operational needs of Smart City ecosystems, this work fills critical gaps in synthetic data generation for privacy-sensitive domains, including intelligent transportation systems, urban health informatics, and next-generation digital commerce infrastructures.

IoT↗

Capturing Historic Reliability Performance Through Graph Databases: A Model Based System Engineering Approach

With the goal of improving the performance and reliability of high dependable technological systems such as nuclear power plants, advanced monitoring and health management systems are employed to inform system engineers on observed degradation processes and anomalous behaviors of assets and components. This information is captured in the form of large amount of data which can be heterogenous in nature (e.g., numeric, textual). Such large data availability poses challenges when system engineers are required to parse and analyze them in order to track historic reliability performance of assets and components. This paper tackles directly this challenge by providing means to organize data in the form of a graph: a knowledge graph. The presented approach distinguish itself from current knowledge graph-based methods by the fact that model-based system engineering (MBSE) models are used to “put data into context”. In particular, MBSE models are used as skeleton of a knowledge graph; numeric and textual data elements, once processed, are associated to MBSE model elements. Thus, a knowledge graph captures both system architecture (though MBSE models) and health/performance data. Such feature opens the door to new data analytics methods designed to identify causal relations between observed phenomena.

97 - MATHEMATICS AND COMPUTING↗

Data and Code for Understanding Generative AI Content with Embedding Models

This repository contains code for the experiments in the paper "Understanding Generative AI Content with Embedding Models". Constructing high-quality features is critical to any quantitative data analysis. While feature engineering was historically addressed by carefully hand-crafting data representations based on domain expertise, deep neural networks (DNNs) now offer a radically different approach. DNNs implicitly engineer features by transforming their input data into hidden feature vectors called embeddings. For embedding vectors produced by foundation models -- which are trained to be useful across many contexts -- we demonstrate that simple and well-studied dimensionality-reduction techniques such as Principal Component Analysis uncover inherent heterogeneity in input data concordant with human-understandable explanations. Of the many applications for this framework, we find empirical evidence that there is intrinsic separability between real samples and those generated by artificial intelligence (AI).

Vargas, Max [Pacific Northwest National Laboratory↗

Evolution of DUNE’s Production System

The DUNE experiment will start running in 2029 and record 30 PB/year of raw waveforms from Liquid Argon TPCs and photon detectors. The size of individual readouts can range from 100 MB to a typical 8 GB full readout of the detector, and even 100 TB for extended readouts from supernova candidates. These data then need to be cataloged, stored and distributed for processing worldwide. This massive amount of data and a heterogeneous computing environment necessitates a powerful and robust distributed computing infrastructure. In the process of building up that infrastructure, DUNE’s production system has recently undergone an overhaul, in which it has integrated 1) a new workflow management system (justIN) 2) a new data catalog (MetaCat) and 3) a state-of-the-art data management system (Rucio). Simulations of DUNE’s Far Detector and its prototypes ProtoDUNE Horizontal Drift (ProtoDUNE-HD) and ProtoDUNE Vertical Drift (ProtoDUNE-VD), as well as data from ProtoDUNE-HD serve as the first tests of this infrastructure.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Multi-omics data resource: Data package 23 (Pck023)

The data package consists of isolated pancreatic islets from adult male C57BL6/J mice treated with IL-1β + IFNγ, IL-1β + IFNγ + NMMA, or NMMA alone for 18 h and submitted for scRNA-seq. This study focused on the cell-type-specific effects of nitric oxide signaling in islets and characterized the heterogeneity of responses. Data contributors: Jennifer S Stancill & John A Corbett: Department of Biochemistry, Medical College of Wisconsin, Milwaukee, WI, USA Data repository: GSE183010 Publication: 10.1093/function/zqab063

Sarkar, Soumyadeep [Pacific Northwest National Lab↗

Quantitative kinetic rules for plastic strain-induced α - ω phase transformation in Zr under high pressure

Plastic strain-induced phase transformations (PTs) and chemical reactions under high pressure are broadly spread in modern technologies, friction and wear, geophysics, and astrogeology. However, because of very heterogeneous fields of plastic strain $E$ p and stress σ tensors and volume fraction c of phases in a sample compressed in a diamond anvil cell (DAC) and impossibility of measurements of σ and $E$ p , there are no strict kinetic equations for them. Here, we develop a kinetic model, finite element method (FEM) approach, and combined FEM-experimental approaches to determine all fields in strongly plastically predeformed Zr compressed in DAC, and specific kinetic equation for α-ω PT consistent with experimental data for the entire sample. Since all fields in the sample are very heterogeneous, data are obtained for numerous complex 7D paths in the space of 3 components of the plastic strain tensor and 4 components of the stress tensor. Kinetic equation depends on accumulated plastic strain (instead of time) and pressure and is independent of plastic strain and deviatoric stress tensors, i.e., it can be applied for various above processes. Our results initiate kinetic studies of strain-induced PTs and provide efforts toward more comprehensive understanding of material behavior in extreme conditions.

36 MATERIALS SCIENCE↗