Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “sparse data analytics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Regional-scale estimates of surface moisture availability and thermal inertia using remote thermal measurements

A review is presented of numerical models which were developed to interpret thermal IR data and to identify the governing parameters and surface energy fluxes recorded in the images. Analytic, predictive, diagnostic and empirical models are described. The limitations of each type of modeling approach are explored in terms of the error sources and inherent constraints due to theoretical or measurement limitations. Sample results of regional-scale soil moisture or evaporation patterns derived from the Heat Capacity Mapping Mission and GOES satellite data through application of the predictive model devised by Carlson (1981) are discussed. The analysis indicates that pattern recognition will probably be highest when data are collected over flat, arid, sparsely vegetated terrain. The soil moisture data then obtained may be accurate to within 10-20 percent.

Carlson, T. N.↗

Nonnegative canonical tensor decomposition with linear constraints: nnCANDELINC

Abstract There is an emerging interest for tensor factorization applications in big‐data analytics and machine learning. To speed up the factorization of extra‐large datasets, organized in multidimensional arrays (also known as tensors), easy to compute compression‐based tensor representations, such as, Tucker and tensor train formats, are used to approximate the initial large‐tensor. Further, tensor factorization is used to extract latent features that can facilitate discoveries of new mechanisms and signatures hidden in the data, where the explainability of the latent features is of principal importance. Nonnegative tensor factorization extracts latent features that are naturally sparse and parts of the data, which makes them easily interpretable. However, to take into account available domain knowledge and subject matter expertise, often additional constraints need to be imposed, which lead us to canonical decomposition with linear constraints (CANDELINC), a canonical polyadic decomposition with rank deficient factors. In CANDELINC, Tucker compression is used as a preprocessing step, which lead to a larger residual error but to more explainable latent features. Here, we propose a nonnegative CANDELINC (nnCANDELINC) accomplished via a specific nonnegative Tucker decomposition; we refer to as minimal or canonical nonnegative Tucker. We derive several results required to understand the specificity of nnCANDELINC, focusing on the difficulties of preserving the nonnegative rank of a tensor to its Tucker core and comparing the real valued to nonnegative case. Finally, we demonstrate nnCANDELINC performance on synthetic and real‐world examples.

97 MATHEMATICS AND COMPUTING↗

Disruption of the Globular Cluster Pal 5

Orbit calculations suggest that the sparse globular cluster, Pal 5, will pass within 7 kpc of the Galactic center the next time it crosses the plane, where it might be destroyed by tidal stresses. We study this problem, treating Pal 5 as a self-consistent dynamical system orbiting through an external potential that represents the Galaxy. The first part of the problem is to find suitable analytic approximations to the Galactic potential. They must be valid in all regions the cluster is likely to explore. Observed velocity and positional data for Pal 5 are used as initial conditions to determine the orbit. Methods we used for a different problem some 12 years ago have been adapted to this problem. Three experiments have been run, with M/L= 1, 3, and 10, for the cluster model. The cluster blew up shortly after passing through the Galactic plane (about 130 Myrs after the beginning of the run) with M/L=1. At M/L = 3 and 10 the cluster survived, although it got quite a kick in the fundamental mode on passing through the plane. But the fundamental mode oscillation died out in a couple of oscillation cycles at M/L=10. Pal 5 will probably be destroyed on its next crossing of the Galactic plane if M/L=1, but it can survive (albeit with fairly heavy damage) if NI/L=3. We haven't tried to trap the mass limits more closely than that. Pal 5 comes through pretty well unscathed at M/L=10. An interesting follow-up experiment would be to back the cluster up along its orbit to look at its previous passage through the Galactic plane, to see what kind of object it might have been at earlier times.

Miller, R. H.↗

Biological Research and Space Health Enabled by Machine Learning to Support Deep Space Missions

A key science goal of the NASA “Moon to Mars” campaign is to understand how biology responds to the Lunar, Martian, and deep space environments in order to advance fundamental knowledge and support human space missions. Through artificial intelligence (AI) and machine learning (ML), a paradigm shift has begun in space biosciences and engineered astronaut health systems, to enable Earth-independence and mission operations autonomy. We describe current AI/ML methods to support 1) fundamental biology, 2) in situ analytics, 3) high performance computing, 4) automated science, 5) self-driving labs, 6) remote data management, 7) integrated mission biomonitoring, and 8) a Precision Space Health system. AI/ML approaches that can be integrated to support these domains include active learning, explainable AI, adaptive learning, causal inference, knowledge graphs, federated learning, transfer learning, and large language models. Finally, we present results from several current ML projects that are underway in the space biology field to address key challenges of small sample n, high feature count, heterogeneity, and sparse data. These include 1) connecting omics to phenotypic data using an ensemble model to infer causality of rodent liver health disruption, 2) usage of explainable ML to interrogate muscular underpinnings of muscle atrophy, 3) ML models analyzing and determining directed acyclic graphs of human health risk leveraging rodent bone datasets, 4) usage of large pre-trained models connecting biomedical knowledgebases with small spaceflight datasets to understand gene-to-gene interactions, and 5) a suite of benchmarked open science datasets enabling programmers to identify best algorithms to answer space biology questions.

space biology↗

Biological Research and Space Health Enabled by Machine Learning to Support Deep Space Missions

A key science goal of the NASA “Moon to Mars” campaign is to understand how biology responds to the Lunar, Martian, and deep space environments in order to advance fundamental knowledge, reduce risk, and support safe, productive human space missions. Through the powerful emerging computer science approaches of artificial intelligence (AI) and machine learning (ML), a paradigm shift has begun in biomedical science and engineered astronaut health systems, to enable Earth-independence and autonomy of mission operations. We present a decadal view of AI/ML architecture to support deep space mission goals, developed in concert with leaders in the field. We describe current AI/ML methods to support 1) fundamental biology, 2) in situ analytics, 3) high performance computing hardware, 4) automated science, 5) self-driving labs, 6) remote data management, 7) integrated real-time mission biomonitoring, and 8) a Precision Space Health system. Cutting-edge AI/ML approaches that can be integrated to support these domains include active learning, explainable AI, adaptive learning, causal inference, knowledge graphs, federated learning, transfer learning, and large language models. Finally, we present results from several current ML projects that are underway in the field to address key challenges of small sample n, high feature count, heterogeneity, and sparse data. These include 1) connecting omics data to phenotypic data using an ensemble model to infer causality of spaceflight rodent liver health disruption, 2) usage of explainable ML to interrogate the muscular underpinnings of spaceflight muscle atrophy, 3) ML models analyzing and determining directed acyclic graphs of human space health risk leveraging rodent bone datasets, 4) usage of large pre-trained models connecting biomedical knowledgebases with small spaceflight datasets to understand gene-to-gene interaction networks, and 5) a suite of benchmarked open science datasets (spaceflight mouse liver; radiation DNA damage) enabling programmers to identify the best ML algorithms to answer space biological science questions.

space biology↗

Biological Research and Space Health Enabled by Machine Learning to Support Deep Space Missions

A key science goal of the NASA “Moon to Mars” campaign is to understand how biology responds to the Lunar, Martian, and deep space environments in order to advance fundamental knowledge, reduce risk, and support safe, productive human space missions. Through the powerful emerging computer science approaches of artificial intelligence (AI) and machine learning (ML), a paradigm shift has begun in biomedical science and engineered astronaut health systems, to enable Earth-independence and autonomy of mission operations. We present a decadal view of AI/ML architecture to support deep space mission goals, developed in concert with leaders in the field. We describe current AI/ML methods to support 1) fundamental biology, 2) in situ analytics, 3) high performance computing hardware, 4) automated science, 5) self-driving labs, 6) remote data management, 7) integrated real-time mission biomonitoring, and 8) a Precision Space Health system. Cutting-edge AI/ML approaches that can be integrated to support these domains include active learning, explainable AI, adaptive learning, causal inference, knowledge graphs, federated learning, transfer learning, and large language models. Finally, we present results from several current ML projects that are underway in the field to address key challenges of small sample n, high feature count, heterogeneity, and sparse data. These include 1) connecting omics data to phenotypic data using an ensemble model to infer causality of spaceflight rodent liver health disruption, 2) usage of explainable ML to interrogate the muscular underpinnings of spaceflight muscle atrophy, 3) ML models analyzing and determining directed acyclic graphs of human space health risk leveraging rodent bone datasets, 4) usage of large pre-trained models connecting biomedical knowledgebases with small spaceflight datasets to understand gene-to-gene interaction networks, and 5) a suite of benchmarked open science datasets (spaceflight mouse liver; radiation DNA damage) enabling programmers to identify the best ML algorithms to answer space biological science questions.

space biology↗

Expediting field-effect transistor chemical sensor design with neuromorphic spiking graph neural networks

Improving the sensitive and selective detection of analytes in a variety of applications requires accelerating the rational design of field-effect transistor (FET) chemical sensors. Achieving high-performance detection relies on identifying optimal probe materials that can effectively interact with target analytes, a process traditionally driven by chemical intuition and time-consuming trial-and-error methods. To address the difficulties in probe screening for FET sensor development, this work presents a methodology that combines neuromorphic machine learning (ML) architectures, specifically a hybrid spiking graph neural network (SGNN), with an enriched dataset of physicochemical properties through semi-automated data extraction using large language models. Achieving a classification accuracy of 0.89 in predicting sensor sensitivity categories, the SGNN model outperformed traditional ML techniques by leveraging its ability to capture both global physicochemical properties and sparse topological features through a hybrid modeling framework. Next-generation sensor design was informed by the actionable insights into the connections between material properties and sensing performance offered by the SGNN framework. Through virtual screening for the detection of per- and polyfluoroalkyl substances (PFAS) as a use case, the effectiveness of the SGNN model was further validated. Density functional theory simulations confirmed graphene as a promising active material for PFAS detection as suggested by the SGNN framework. By bridging gaps in predictive modeling and data availability, this integrated approach provides a strong foundation for accelerating advancements in FET sensor design and innovation.

Ferreira, Rodrigo Pires [Univ. of Chicago, IL (Uni↗

Frictionless knowledge injection for few-shot learning

Cutting-edge machine learning methods often require large volumes of curated training data, precluding their use in national security problems with rare events in massive datasets. We present a method for incorporating abstract knowledge into models tailored for sparse data. A subject matter expert defines salient concepts using data examples, which are encoded in the model’s embedding space. Models are then trained to respect these concepts. This method enables knowledge injection, yielding effective models with limited labeled data and the ability to assess model sensitivity for subject matter expertise across the nonproliferation mission space, as demonstrated with Raman spectra analysis.

Stomps, Jordan [ORNL] (ORCID:0000000178114479)↗

Rare earth element enrichment in coal and coal-adjacent strata of the Uinta Region, Utah and Colorado

This study aims to quantify rare earth element enrichment within coal and coal-adjacent strata in the Uinta Region of central Utah and western Colorado. Rare earth elements are a subset of critical minerals as defined by the U.S. Geological Survey. These elements are used for a wide variety of applications, including renewable energy technology in the transition toward carbon-neutral energy. While rare earth element enrichment has been associated with Appalachian coals, there has been a more limited evaluation of western U.S. coals. Here, samples from six active mines, four idle/historical mines, four mine waste piles, and seven stratigraphically complete cores within the Uinta Region were geochemically evaluated using portable X-ray fluorescence ( n = 3,113) and inductively coupled plasma-mass spectrometry ( n = 145) elemental analytical methods. Results suggest that 24%–45% of stratigraphically coal-adjacent carbonaceous shale and siltstone units show rare earth element enrichment (>200 ppm), as do 100% of sampled igneous material. A small subset (5%–8%) of coal samples display rare earth element enrichment, specifically in cases containing volcanic ash. This study proposes two multi-step depositional and diagenetic models to explain the enrichment process, requiring the emplacement and mobilization of rare earth element source material due to hydrothermal and other external influences. Historical geochemical evaluations of Uinta Region coal and coal-adjacent data are sparse, emphasizing the statistical significance of this research. These results support the utilization of active mines and coal processing waste piles for the future of domestic rare earth element extraction, offering economic and environmental solutions to pressing global demands.

Coe, Haley H.↗

GPU-Accelerated Analytic Simulation of Sparse Ionization Signal Formation in Pixelated Projection Detector

This paper presents a GPU-accelerated simulation package, TRED, for next-generation neutrino detectors with pixelated charge readout, leveraging community-driven software ecosystems to ensure adaptability and extensibility. We introduce two generic contributions: (i) an effective-charge representation based on Gaussian quadrature rules, in which the linear- interpolation factors for the field response inside each voxel are absorbed into the effective charge, and (ii) a sparse, block- binned tensor representation that enables efficient FFT-based computation of induced signals on readout electrodes for sparsely activated detector volumes. The former captures structure inside a voxel without dense sampling, while the latter achieves low memory usage and scalable runtime, as demonstrated in bench- mark studies. The underlying data representation is applicable to large-scale detectors and to other computational problems involving sparse activity.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Variance Preserving Spectral Subsampling

Generating statistically faithful short-duration gamma-ray spectra from a single long measurement is essential in nuclear safeguards, supporting tasks such as algorithm development and machine-learning applications, especially when list-mode data are unavailable. Existing subsampling methods often distort the statistical characteristics of genuine short-duration measurements, leading to biased or unreliable analytical outcomes and thereby undermining downstream tasks. In this work, we compare five subsampling approaches using a benchmark set of 156 genuine replicate spectra collected with a high-purity germanium detector. We evaluate each method with respect to run-to-run variance, channel-to-channel variance, and preservation of total counts (losslessness). Across a wide range of subsampling ratios, only binomial subsampling without replacement consistently reproduces the statistical properties of genuine short-duration spectra, maintaining proper dispersion even in sparse spectral regions and perfectly preserving total counts. These results provide a mathematically principled and practically validated framework for generating synthetically shortened spectra when true short-duration measurements are unavailable.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Exact Gaussian processes for massive datasets via non-stationary sparsity-discovering kernels

Abstract A Gaussian Process (GP) is a prominent mathematical framework for stochastic function approximation in science and engineering applications. Its success is largely attributed to the GP’s analytical tractability, robustness, and natural inclusion of uncertainty quantification. Unfortunately, the use of exact GPs is prohibitively expensive for large datasets due to their unfavorable numerical complexity of $$O(N^3)$$ O ( N 3 ) in computation and $$O(N^2)$$ O ( N 2 ) in storage. All existing methods addressing this issue utilize some form of approximation—usually considering subsets of the full dataset or finding representative pseudo-points that render the covariance matrix well-structured and sparse. These approximate methods can lead to inaccuracies in function approximations and often limit the user’s flexibility in designing expressive kernels. Instead of inducing sparsity via data-point geometry and structure, we propose to take advantage of naturally-occurring sparsity by allowing the kernel to discover—instead of induce—sparse structure. The premise of this paper is that the data sets and physical processes modeled by GPs often exhibit natural or implicit sparsities, but commonly-used kernels do not allow us to exploit such sparsity. The core concept of exact, and at the same time sparse GPs relies on kernel definitions that provide enough flexibility to learn and encode not only non-zero but also zero covariances. This principle of ultra-flexible, compactly-supported, and non-stationary kernels, combined with HPC and constrained optimization, lets us scale exact GPs well beyond 5 million data points.

97 MATHEMATICS AND COMPUTING↗

Bayesian Physics Informed Spatio-Temporal Network for Streamflow Data Imputation

Reliable reconstruction of incomplete streamflow records is critical for improving hydrological forecasting, flood preparedness, and water resource management. However, large observational gaps and uncertainties in governing physical parameters limit the accuracy of traditional statistical and machinelearning imputation frameworks. To address these challenges, we develop a Bayesian Physics-Informed Spatio-Temporal Network (BPI-STNet) that jointly captures spatial and temporal dependencies while enforcing hydrologic consistency through embedded physical constraints. The framework integrates a GraphSAGE-LSTM architecture to model spatial connectivity across gauges and temporal flow dynamics, coupled with a Bayesian update mechanism to estimate uncertain parameters in a simplified water-balance framework. Unlike conventional physics-informed networks that rely on sampling-based posterior estimation, BPI-STNet derives an analytic solution to the inverse problem, allowing closed-form Bayesian updates of uncertain parameters Λ={α,β,k} using Gaussian priors and likelihoods. Applied to daily observations from the Susquehanna River Basin (1980-2022), BPI-STNet achieves substantial improvements over a purely data-driven RGNN baseline, which reduced RMSE by 23 % and MAE by 9 %, and achieving an average NSE values up to 0.96. The results demonstrate that coupling Bayesian inference with physics-informed learning yields physically consistent, uncertainty-aware reconstructions that preserve the temporal persistence and statistical distribution of observed flows. The proposed framework establishes a generalizable paradigm for data-sparse hydrologic systems where both data fidelity and physical interpretability are essential.

Krishnan Kutty Ambika, Anukesh [ORNL] (ORCID:00000↗

A Hybrid Biophysical‐Machine Learning Framework for Diurnal Surface Energy Flux Estimation Using Proximal Sensing

Thermal infrared-based remote sensing of surface energy fluxes has traditionally relied on high spatial resolution satellite data with revisit frequencies on the order of weeks. In this study, we evaluate a biophysics-based analytical surface energy balance model for predicting latent energy ( LE ) and sensible heat ( H ) fluxes using proximal sensing observations. The Surface Temperature Initiated Closure (STIC1.2) model has been extensively validated across a wide range of spatial and temporal scales using various satellite-derived thermal infrared data sets. Here we extend this validation by applying STIC at sub-hourly temporal resolution over multiple growing seasons for four distinct agricultural systems. We further develop and evaluate novel STIC variants that incorporate machine learning (ML) techniques to eliminate the need for surface energy balance observations, specifically net radiation and soil heat flux, thereby enhancing model applicability in data-sparse settings. The integration of a ML component to estimate surface available energy is shown to have strong predictive performance for both LE (R 2 = 0.81–0.94) and H (R 2 = 0.46–0.72) across all agricultural systems examined here, demonstrating the potential of hybrid biophysical-machine learning approaches for surface energy balance modeling with minimal data requirements. This study concludes with a novel application of explainable machine learning (exML) to diagnose sources of model error. This exML framework attributes residual prediction errors to both model input variables and environmental drivers not explicitly included in the simulation experiments. This approach provides a new pathway for improving model design and integrating previously overlooked yet influential variables into future model iterations.

evapotranspiration↗

Stochastic modeling and statistical calibration with model error and scarce data

This paper introduces a procedure to assess the predictive accuracy of stochastic models subject to model error and sparse data. Model error is introduced as uncertainty on the coefficients of appropriate polynomial chaos expansions (PCE). The error associated with finite sample size allows us to conceive of these coefficients as statistics of the data that we describe as random variables whose influence on output quantities of interest is evaluated through the extended polynomial chaos expansion (EPCE). A Bayesian data assimilation scheme is introduced to update these expansions by considering the resulting nested chaos expansion as a hierarchical probabilistic model. Stochastic models of quantities of interest (QoI) are thus constructed and efficiently evaluated. Here, the Metropolis–Hastings Markov chain Monte Carlo procedure is used to sample the posterior. Two illustrative analytical and numerical problems are used to demonstrate the proposed approach.

Bayesian inference↗

Coarse-Grain Bandwidth Estimation Scheme for Large-Scale Network

A large-scale network that supports a large number of users can have an aggregate data rate of hundreds of Mbps at any time. High-fidelity simulation of a large-scale network might be too complicated and memory-intensive for typical commercial-off-the-shelf (COTS) tools. Unlike a large commercial wide-area-network (WAN) that shares diverse network resources among diverse users and has a complex topology that requires routing mechanism and flow control, the ground communication links of a space network operate under the assumption of a guaranteed dedicated bandwidth allocation between specific sparse endpoints in a star-like topology. This work solved the network design problem of estimating the bandwidths of a ground network architecture option that offer different service classes to meet the latency requirements of different user data types. In this work, a top-down analysis and simulation approach was created to size the bandwidths of a store-and-forward network for a given network topology, a mission traffic scenario, and a set of data types with different latency requirements. These techniques were used to estimate the WAN bandwidths of the ground links for different architecture options of the proposed Integrated Space Communication and Navigation (SCaN) Network. A new analytical approach, called the "leveling scheme," was developed to model the store-and-forward mechanism of the network data flow. The term "leveling" refers to the spreading of data across a longer time horizon without violating the corresponding latency requirement of the data type. Two versions of the leveling scheme were developed: 1. A straightforward version that simply spreads the data of each data type across the time horizon and doesn't take into account the interactions among data types within a pass, or between data types across overlapping passes at a network node, and is inherently sub-optimal. 2. Two-state Markov leveling scheme that takes into account the second order behavior of the store-and-forward mechanism, and the interactions among data types within a pass. The novelty of this approach lies in the modeling of the store-and-forward mechanism of each network node. The term store-and-forward refers to the data traffic regulation technique in which data is sent to an intermediate network node where they are temporarily stored and sent at a later time to the destination node or to another intermediate node. Store-and-forward can be applied to both space-based networks that have intermittent connectivity, and ground-based networks with deterministic connectivity. For groundbased networks, the store-and-forward mechanism is used to regulate the network data flow and link resource utilization such that the user data types can be delivered to their destination nodes without violating their respective latency requirements.

Cheung, Kar-Ming↗

The importance of cycle-by-cycle data in performing rapid battery technology development and validation

Lithium-ion battery (LiB) technology is playing a crucial role in transforming the predominantly fossil fuel-based transportation and stationary storage sectors to achieve a low-carbon economy. Rapid innovation in the LiB materials to electrode to cell design is happening to satisfy the performance, life, and safety metrics required by those myriads of applications. Lately, advanced analytics, such as machine-learning or artificial intelligence (ML/AI) techniques, are being used more frequently to aid in expedited LiB technology development, performance validation, and life prediction. The success of these techniques often relies on a large volume of well-defined and high-quality battery test data. On the other hand, most battery developers and research and development (R&D) communities are still following a classical approach to develop batteries, which is running calendar- and/or cycle-aging tests, performing reference performance tests (RPTs), and conducting post-mortem analyses periodically without paying attention to the wealth of data often not collected during the calendar or cycle life aging tests. This sparse data collection approach is time- and resource-intensive, requiring data capture and evaluation of months to years of RPT data to diagnose accurate battery state of performance, health, and safety. Even so, the underlying aging modes and mechanisms can be missed. If collected properly, battery test data during cycling or calendaring can be efficiently combined with ML/AI techniques to create powerful tools in the rapid diagnosis of battery state of performance, health, and safety along with insights into underlying aging modes and mechanisms. In this report, we discuss the importance of effective cycle-by-cycle (CBC) data collection with example case studies. Within a reasonable timeframe, RPT data are often inadequate in capturing many of the crucial battery aging dynamics, which often predominantly show up in CBC test data. Finally, we also show examples of ML/AI techniques that use CBC data in rapid diagnosis and projection of LiB state of health (SOH) to motivate the scientific community in collecting and using CBC data to facilitate expeditious technology development and validation.

25 ENERGY STORAGE↗

Transferring a Molecular Foundation Model for Polymer Property Predictions

Transformer-based large language models have remarkable potential to accelerate design optimization for applications such as drug development and material discovery. Self-supervised pretraining of transformer models requires large-scale data sets, which are often sparsely populated in topical areas such as polymer science. Further, state-of-the-art approaches for polymers conduct data augmentation to generate additional samples but unavoidably incur extra computational costs. In contrast, large-scale open-source data sets are available for small molecules and provide a potential solution to data scarcity through transfer learning. In this work, we show that using transformers pretrained on small molecules and fine-tuned on polymer properties achieves comparable accuracy to those trained on augmented polymer data sets for a series of benchmark prediction tasks.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗