Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Sparse Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Advanced Material Studies for High Intensity Proton Production Targets and Windows

The titanium alloy Ti6Al4V is widely used in accelerator facilities as beam windows, which are exposed to high cycle compressive stress waves from intense pulsed proton beams. Such materials interacting with the beam are subject to various forms of radiation damage, which can adversely affect their endurance limit. However, no fatigue data is currently available for high energy proton irradiated titanium alloy. Due to limitations in proton irradiation facilities, only miniature flat samples can be used for irradiation. To address this issue, we have developed a custom-made bend fatigue tester at Fermilab specifically for testing proton irradiated titanium alloy. In this presentation, we report on the fatigue test results obtained from this custom-fatigue tester using a non-ASTM standard specimen design. We plan to validate these sparse fatigue data with ASTM standard samples using standard fatigue samples. Furthermore, we have modified another commercial bend fatigue tester to accommodate miniature samples, and discuss some inherent deficiencies of the commercial fatigue tester to test miniature samples. To overcome this issue, a new fixture design is presented, which enables satisfactory fatigue testing on miniature samples over long periods. Finally, we present an upgrade to the custom-fatigue tester, featuring this new fixture design.

43 PARTICLE ACCELERATORS↗

Pan-sharpening for microscopy

Techniques for generating full-spatial resolution, full spectral resolution image(s) from a 3D spectral-data cube for any spectral value within a given spectral range are provided without requiring the acquisition of all full-spatial resolution, full spectral resolution data by an instrument. The 3D spectral-data cube is generated from a limited number of full-spatial resolution, sparse spectral resolution data and a sparse-spatial resolution, full-spectral resolution data of the same area of the sample. The use of the 3D spectral-data cube reduces the data acquisition time.

Ovchinnikova, Olga S.↗

Physics constrained learning for data-driven inverse modeling from sparse observations

Deep neural networks (DNN) have been used to model nonlinear relations between physical quantities. Those DNNs are embedded in physical systems described by partial differential equations (PDE) and trained by minimizing a loss function that measures the discrepancy between predictions and observations in some chosen norm. This loss function often includes the PDE constraints as a penalty term when only sparse observations are available. As a result, the PDE is only satisfied approximately by the solution. However, the penalty term typically slows down the convergence of the optimizer for stiff problems. We present a new approach that trains the embedded DNNs while numerically satisfying the PDE constraints. We develop an algorithm that enables differentiating both explicit and implicit numerical solvers in reverse-mode automatic differentiation. This allows the gradients of the DNNs and the PDE solvers to be computed in a unified framework. We demonstrate that our approach enjoys faster convergence and better stability in relatively stiff problems compared to the penalty method. Furthermore, our approach allows for the potential to solve and accelerate a wide range of data-driven inverse modeling, where the physical constraints are described by PDEs and need to be satisfied accurately.

97 MATHEMATICS AND COMPUTING↗

Increasing the Scale of the Mass Spectrometry Query Language Compendium with Explainable AI

A significant bottleneck in metabolomics data interpretation is the effective use of domain knowledge to assign structural information based on fragmentation patterns. The mass spectrometry query language (MassQL) aims to make this process accessible and applicable across multiple analysis platforms. While advanced computational methods are capable of predicting compound structures from fragmentation data, AI/ML approaches often rely on complex, opaque criteria that are difficult to interpret or modify. As a result, their predictive patterns cannot be readily translated into human-readable rules, such as those used in MassQL. Here, in this study, we introduce ChemEcho, a machine learning embedding method that converts tandem mass spectrometry data into sparse feature vectors containing peak and neutral mass subformulae to enhance explainable AI/ML-based methods. An advantage of this approach is that decision trees trained using these feature vectors can be directly translated to MassQL. Using a battery of decision trees trained using ChemEcho embeddings to predict molecular attributes, we generated over 1500 MassQL queries for 765 molecular features and evaluated their precision and recall. From these queries, the 50 highest-performing queries were integrated into the MassQL compendium. This set of generated MassQL queries included environmentally and biologically relevant classes such as PFAS and molecules containing phosphate or sulfate substructures. To illustrate the impact these queries would have on a typical metabolomics experiment, these MassQL queries were applied to a public metabolomics data set─resulting in a marked increase in the structural information derived from tandem mass spectra. Access and reuse of these queries is expected to enhance structural annotation in untargeted experiments, leading to more specific claims and advancing many applications in metabolomics.

Harwood, Thomas V. [USDOE Joint Genome Institute (↗

Data-Driven Discovery of Active Nematic Hydrodynamics

Active nematics can be modeled using phenomenological continuum theories that account for the dynamics of the nematic director and fluid velocity through partial differential equations (PDEs). While these models provide a statistical description of the experiments, the relevant terms in the PDEs and their parameters are usually identified indirectly. We adapt a recently developed method to automatically identify optimal continuum models for active nematics directly from spatio-temporal data, via sparse regression of the coarse-grained fields onto generic low order PDEs. After extensive benchmarking, we apply the method to experiments with microtubule-based active nematics, finding a surprisingly minimal description of the system. Furthermore, our approach can be generalized to gain insights into active gels, microswimmers, and diverse other experimental active matter systems.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Forecasting high-dimensional spatio-temporal systems from sparse measurements

This paper introduces a new neural network architecture designed to forecast high-dimensional spatio-temporal data using only sparse measurements. The architecture uses a two-stage end-to-end framework that combines neural ordinary differential equations (NODEs) with vision transformers. Initially, our approach models the underlying dynamics of complex systems within a low-dimensional space; and then it reconstructs the corresponding high-dimensional spatial fields. Many traditional methods involve decoding high-dimensional spatial fields before modeling the dynamics, while some other methods use an encoder to transition from high-dimensional observations to a latent space for dynamic modeling. In contrast, our approach directly uses sparse measurements to model the dynamics, bypassing the need for an encoder. This direct approach simplifies the modeling process, reduces computational complexity, and enhances the efficiency and scalability of the method for large datasets. We demonstrate the effectiveness of our framework through applications to various spatio-temporal systems, including fluid flows and global weather patterns. Although sparse measurements have limitations, our experiments reveal that they are sufficient to forecast system dynamics accurately over long time horizons. Our results also indicate that the performance of our proposed method remains robust across different sensor placement strategies, with further improvements as the number of sensors increases. This robustness underscores the flexibility of our architecture, particularly in real-world scenarios where sensor data is often sparse and unevenly distributed.

97 MATHEMATICS AND COMPUTING↗

Aboveground biomass density models for NASA’s Global Ecosystem Dynamics Investigation (GEDI) lidar mission

NASA's Global Ecosystem Dynamics Investigation (GEDI) is collecting spaceborne full waveform lidar data with a primary science goal of producing accurate estimates of forest aboveground biomass density (AGBD). This paper presents the development of the models used to create GEDI's footprint-level (~25 m) AGBD (GEDI04_A) product, including a description of the datasets used and the procedure for final model selection. The data used to fit our models are from a compilation of globally distributed spatially and temporally coincident field and airborne lidar datasets, whereby we simulated GEDI-like waveforms from airborne lidar to build a calibration database. We used this database to expand the geographic extent of past waveform lidar studies, and divided the globe into four broad strata by Plant Functional Type (PFT) and six geographic regions. GEDI's waveform-to-biomass models take the form of parametric Ordinary Least Squares (OLS) models with simulated Relative Height (RH) metrics as predictor variables. From an exhaustive set of candidate models, we selected the best input predictor variables, and data transformations for each geographic stratum in the GEDI domain to produce a set of comprehensive predictive footprint-level models. We found that model selection frequently favored combinations of RH metrics at the 98th, 90th, 50th, and 10th height above ground-level percentiles (RH98, RH90, RH50, and RH10, respectively), but that inclusion of lower RH metrics (e.g. RH10) did not markedly improve model performance. Second, forced inclusion of RH98 in all models was important and did not degrade model performance, and the best performing models were parsimonious, typically having only 1-3 predictors. Third, stratification by geographic domain (PFT, geographic region) improved model performance in comparison to global models without stratification. Fourth, for the vast majority of strata, the best performing models were fit using square root transformation of field AGBD and/or height metrics. There was considerable variability in model performance across geographic strata, and areas with sparse training data and/or high AGBD values had the poorest performance. These models are used to produce global predictions of AGBD, but will be improved in the future as more and better training data become available.

54 ENVIRONMENTAL SCIENCES↗

Data-driven analysis of dipole strength functions using artificial neural networks

Here, we present a data-driven analysis of dipole strength functions across the nuclear chart, employing an artificial neural network to model nuclear dipole responses. We train the network on a dataset of experimentally measured dipole strength functions for 216 different nuclei. To assess its predictive capability, we test the trained model on an additional set of 10 new nuclei, where experimental data exist. We demonstrate that the artificial neural network not only accurately reproduces known data but also identifies potential inconsistencies in experimental datasets, indicating which results may warrant further review or possible rejection. For nuclei where experimental data are sparse or unavailable, the network confirms theoretical calculations, reinforcing its utility as a predictive tool in nuclear physics. Finally, utilizing the predicted electric dipole polarizability, we extract the value of the symmetry energy at saturation density and find it consistent with results from the literature.

artificial neural networks↗

Design of a graphical user interface for few-shot machine learning classification of electron microscopy data

The recent growth in data generation by modern electron microscopes requires rapid, scalable, and flexible approaches to image segmentation and analysis. Few-shot machine learning, which can richly classify images from a handful of user-provided examples, is a promising route to high-throughput analysis. However, current command-line implementations of such approaches can be slow and unintuitive to use, lacking the real-time feedback necessary to perform effective classification. Here we report on the development of a Python-based graphical user interface that enables end users to easily conduct and visualize the output of few-shot learning models. This interface is portable and can be hosted locally or on the web, providing the opportunity to reproducibly conduct, share, and crowd-source few-shot analyses.

97 MATHEMATICS AND COMPUTING↗

On a Simplified Approach to Achieve Parallel Performance and Portability Across CPU and GPU Architectures

This paper presents software advances to easily exploit computer architectures consisting of a multi-core CPU and CPU+GPU to accelerate diverse types of high-performance computing (HPC) applications using a single code implementation. The paper describes and demonstrates the performance of the open-source C++ matrix and array (MATAR) library that uniquely offers: (1) a straightforward syntax for programming productivity, (2) usable data structures for data-oriented programming (DOP) for performance, and (3) a simple interface to the open-source C++ Kokkos library for portability and memory management across CPUs and GPUs. The portability across architectures with a single code implementation is achieved by automatically switching between diverse fine-grained parallelism backends (e.g., CUDA, HIP, OpenMP, pthreads, etc.) at compile time. The MATAR library solves many longstanding challenges associated with easily writing software that can run in parallel on any computer architecture. This work benefits projects seeking to write new C++ codes while also addressing the challenges of quickly making existing Fortran codes performant and portable over modern computer architectures with minimal syntactical changes from Fortran to C++. We demonstrate the feasibility of readily writing new C++ codes and modernizing existing codes with MATAR to be performant, parallel, and portable across diverse computer architectures.

97 MATHEMATICS AND COMPUTING↗

Evaluation of an event-driven 3FI ASIC for spectroscopic X-ray detection with synchrotron radiation

The novel design and evaluation on the NSLS-II beamline of the 3FI application-specific integrated circuit (ASIC) bump-bonded to a simple, planar, 2D segmented silicon sensor are presented. The ASIC was developed for full-field fluorescence spectral X-ray imaging (3FI). It is a small-scale prototype that features a square array of 32 × 32 pixels, and the size of the pixels is 100 µm × 100 µm. The ASIC was implemented in a 65 nm CMOS integrated circuit fabrication process. Each pixel incorporates a charge-sensitive amplifier, a shaping filter, a discriminator, a peak detector and a sample-and-hold circuit, allowing detection of events and storage of signal amplitudes. The system operates in a frameless event-driven readout mode, outputting analog values for threshold-triggered events, allowing high-speed multi-element X-ray fluorescence data acquisition. The 3FI ASIC achieves per-channel spectrometric performance at a power consumption of only 200 µW per pixel, with nearly all dissipation confined to the analog front-end. An energy resolution is measured at the level of 308 eV full width at half-maximum (FWHM) at 8.04 keV (Cu Kα), and 138 eV FWHM at 3.69 keV (Ca Kα). This per-pixel capability makes the prototype suitable for in situ trace element microanalysis in biological and environmental studies. Moreover, the frameless architecture of the detector is designed to address limitations of conventional X-ray fluorescence microscopy, which typically requires mechanical scanning, by enabling continuous high-throughput data acquisition in future full-field implementations.

47 OTHER INSTRUMENTATION↗

Accurate data-driven surrogates of dynamical systems for forward propagation of uncertainty

Stochastic collocation (SC) is a well-known non-intrusive method of constructing surrogate models for uncertainty quantification. In dynamical systems, SC is especially suited for full-field uncertainty propagation that characterizes the distributions of the high-dimensional solution fields of a model with stochastic input parameters. However, due to the highly nonlinear nature of the parameter-to-solution map in even the simplest dynamical systems, the constructed SC surrogates are often inaccurate. Here, this work presents an alternative approach, where we apply the SC approximation over the dynamics of the model, rather than the solution. By combining the data-driven sparse identification of nonlinear dynamics framework with SC, we construct dynamics surrogates and integrate them through time to construct the surrogate solutions. We demonstrate that the SC-over-dynamics framework leads to smaller errors, both in terms of the approximated system trajectories as well as the model state distributions, when compared against full-field SC applied to the solutions directly. We present numerical evidence of this improvement using three test problems: a chaotic ordinary differential equation, and two partial differential equations from solid mechanics.

42 ENGINEERING↗

Integrating Analytical Solutions and U-Net Model for Predicting Groundwater Contaminant Plumes in Pump-and-Treat Systems

Pump-and-treat (P&T) is a common technique for groundwater remediation involving the extraction and treatment of contaminated water above ground. Optimizing the design and operation of the P&T well network is essential for maximizing the system’s effectiveness and efficiency. However, this optimization often necessitates many model evaluations, leading to computationally demanding tasks. This study introduces a novel approach that integrates analytical solutions for groundwater dynamics with the U-Net (Ronneberger et al., 2015) deep learning framework to predict groundwater contaminant plume migration under dynamic pumping conditions. By incorporating the Thiem equation (Thiem, 1906) into the input preprocessing, the U-Net model transforms sparse well data into a continuous spatial field that captures the hydraulic impacts of pumping activities. This integration enables the model to leverage both deep learning capabilities and classical physics-based groundwater theories, enhancing prediction accuracy and computational efficiency. These advancements can facilitate rapid, large-scale evaluations of P&T optimization simulations, allowing for timely and effective decision-making in well placement and system management. We demonstrate the model's robust performance across both simplified transient 2D models and a more complex 3D heterogeneous site model at the 200 West P&T facility at the Hanford Site. The U-Net-based model offers substantial computational advantages, reducing simulation times significantly compared to full physics-based models and providing a powerful tool for rapid site evaluation and P&T system optimization, such as evaluating alternative P&T well network designs. Our findings highlight the potential of advanced machine learning models to significantly enhance the efficiency and sustainability of groundwater remediation efforts, offering a novel application of U-Net architecture in environmental science.

Pump-and-treat↗

Hierarchical semi-Markov models with duration-aware dynamics for activity sequences

Residential electricity demand at granular scales is driven by what people do and for how long. Accurately forecasting this demand for applications like microgrid management and demand response therefore requires generative models for activities that can produce realistic daily activity sequences, capturing both the timing and duration of human behavior. This paper develops a generative model of human activity sequences using nationally representative time-use diaries at a 10-min resolution. We use this model to quantify which demographic factors are most critical for improving predictive performance. We propose a hierarchical semi-Markov framework that addresses two key modeling challenges. First, a time-inhomogeneous Markov router learns the patterns of “which activity comes next.” Second, a semi-Markov hazard component explicitly models activity durations, capturing “how long” activities realistically last. To ensure statistical stability when data are sparse, the model pools information across related demographic groups and time blocks. The entire framework is trained and evaluated using survey design weights to ensure our findings are representative of the U.S. population. On a held-out test set, we demonstrate that explicitly modeling durations with the hazard component provides a substantial and statistically significant improvement over purely Markovian models. Furthermore, our analysis reveals a clear hierarchy of demographic factors: Sex, Day-Type, and Household Size provide the largest predictive gains, while Region and Season, though important for energy calculations, contribute little to predicting the activity sequence itself. The result is an interpretable and robust generator of synthetic activity traces, providing a high-fidelity foundation for downstream energy systems modeling.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Field testing and validation of a low-cost MPC for demand flexibility for grid-interactive K-12 schools

K-12 school buildings account for the highest energy consumption within the public sector. Implementing advanced HVAC controls in grid-interactive K-12 schools could bring substantial economic advantages and grid flexibility. Our previous study demonstrated that a low-cost model predictive control (MPC) solution, which coordinates multiple packaged units, can enable demand flexibility without major hardware upgrades. However, a significant gap remains between academic pilots and market-ready scalable solutions. This paper extends the previous single-site pilot to a multi-site demonstration involving three school campuses (95 total units) through a commercial technology transfer process. Addressing the challenge of verifying performance with sparse field data, we present a new statistical approach using Bayesian methods to estimate the MPC’s effect on peak demand. Unlike traditional methods, this approach robustly quantifies uncertainty in non-normal, limited datasets. The results confirm the solution’s replicability, achieving a 21.6–38.9% reduction in HVAC peak demand (10.8–22.1% at the site-level) with > 98% probability across diverse locations. Finally, we document critical barriers to scaling software-as-a-service (SaaS) solutions–such as API instability and diverse legacy systems–and offer practical strategies to accelerate the commercial adoption of grid-interactive efficient buildings.

Ham, Sang Woo↗

Stratigraphic Identification with Airborne Electromagnetic Methods at the Hanford Site, Washington

Stratigraphic units can influence the fate and transport of subsurface contaminants within groundwater. Units having coarse-grained sediments act as preferential flow pathways, and therefore can accelerate the transport of contaminants to reach human and ecological receptors. At legacy waste sites, detailed knowledge of subsurface stratigraphy can be used for effective monitoring and remediation planning to help minimize risk to human health and the environment. Airborne electromagnetic (AEM) methods can non-invasively provide information on kilometer-scale or larger subsurface stratigraphic features and fill informational gaps in directly sampled data from sparsely located boreholes. In this paper, we present inversion results of a 412 line-km frequency-domain AEM survey to delineate subsurface stratigraphic features at the Hanford Site, located in southeastern Washington State. The inversion was performed using a massively parallel 3D electromagnetic modeling and inversion code, where the modeling is based on solving frequency-domain Maxwell’s equations using an unstructured-mesh finite-element method and the inversion employs a Gauss-Newton optimization scheme. The results are compared to an underlying geologic framework model (GFM), built by interpolating contact depths of stratigraphic units interpreted from site borehole datasets. In areas with good borehole coverage, the inversion results show a good match with the GFM to a depth of about 60 m. Outside of these areas, the inversion results exhibit inconsistencies from the assumptions made to create the GFM, demonstrating that the AEM survey results can be used to improve the understanding of the geological conceptual model.

47 OTHER INSTRUMENTATION↗

Modeling magnetic fields with helical solutions to Laplace’s equation

The series solution to Laplace's equation in a helical coordinate system is derived and refined using symmetry and chirality arguments. These functions and their more commonplace counterparts are used to model solenoidal magnetic fields via linear, multidimensional curve-fitting. A judicious choice of functional forms, a small number of free parameters and sparse input data can lead to highly accurate, fine-grained modeling of solenoidal magnetic fields, including helical features arising from the winding of the solenoid, with overall field accuracy at better than one part per million.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗