Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “hierarchical data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Distributed Transient Safety Verification via Robust Control Invariant Sets: A Microgrid Application

Modern safety-critical energy infrastructures are increasingly operated in a hierarchical and modular control framework which allows for limited data exchange between the modules. In this context, it is important for each module to synthesize and communicate constraints on the values of exchanged information in order to assure system-wide safety. To ensure transient safety in inverter-based microgrids, we develop a set invariance-based distributed safety verification algorithm for each inverter module. Applying Nagumo's invariance condition, we construct a robust polynomial optimization problem to jointly search for safety-admissible set of control set-points and design parameters, under allowable disturbances from neighbors. We use sum-of-squares (SOS) programming to solve the verification problem and we perform numerical simulations using grid-forming inverters to illustrate the algorithm.

Bouvier, Jean-Baptiste H.↗

Conformal Hierarchical Simulation-Based Inference with Local Validity

Trustworthy and interpretable uncertainty quantification is a long-standing challenge in artificial intelligence. Simulation-based inference (SBI) comprises a broad swath of approaches for estimating latent parameters with uncertainties. Although flexible neural density estimators in SBI can be remark- ably expressive capturing highly structured, high-dimensional posteriors their credible regions can be badly mis-calibrated and are often only accompanied by heuristic coverage checks. We present the first SBI framework that delivers finite-sample local valid coverage guarantees that hold in the neighborhood of each observation. Our framework can couple any off-the-shelf hierarchical SBI engine with a confor- mal Bayesian post-processing step that operates on the posterior predictive density. A kernel-weighted conformity score adapts the conformal quantile to the local geometry of the data, yielding prediction sets that are simultaneously (i) marginally calibrated, (ii) locally valid, and (iii) hierarchical, handling global and observation-specific parameters in a single pass. Through experiments on synthetic data and benchmarks from neuroscience and physics, we show that our approach attains 1 − α coverage, where prior SBI methods under- or over-cover. Our approach also maintains a competitive, credible set size with minimal computational overhead. Finally, our approach can be used to make predictions on real data and give valid credible regions modulo weight-initialization-based model mis-specification.

Trivedi, Shubhendu [Fermilab]↗

Identifying Vehicle Signals in Continuous Seismic Data Using Unsupervised Machine-Learning Techniques

Seismic sensors deployed near roadways effectively capture ground vibrations generated by passing vehicles. Although both traditional and machine‐learning algorithms have been utilized for analyzing such signals, independent validation of detected vehicle events remains limited. We applied two unsupervised machine‐learning algorithms, uniform manifold approximation and projection for dimension reduction, and hierarchical density‐based spatial clustering of applications with noise, to continuous seismic data collected along a road on the main campus of Oak Ridge National Laboratory. The algorithms identified seven distinct cluster labels across the entire dataset. By comparing these cluster labels with precipitation records from a nearby weather station and image‐derived labels from a local camera system, we identified one cluster associated with rainfall and another with vehicle activity. Our algorithms identified a greater number of vehicle‐related labels compared to the camera‐derived labels because seismic data are unaffected by poor lighting conditions. The arrival times of the newly detected vehicle signals corresponded well with the road’s speed limit, supporting our findings. Our algorithm outperformed the short‐term average/long‐term average method and k‐means clustering. Our results suggest that seismic data, when analyzed with machine‐learning algorithms, can complement existing vehicle monitoring systems, particularly under challenging environmental conditions.

Chai, Chengping [Oak Ridge National Laboratory (OR↗

Scaling pair count to next galaxy surveys

ABSTRACT Counting pairs of galaxies or stars according to their distance is at the core of real-space correlation analyses performed in astrophysics and cosmology. Upcoming galaxy surveys (LSST, Euclid) will measure properties of billions of galaxies challenging our ability to perform such counting in a minute-scale time relevant for the usage of simulations. The problem is only limited by efficient access to the data, hence belongs to the big data category. We use the popular Apache Spark framework to address it and design an efficient high-throughput algorithm to deal with hundreds of millions to billions of input data. To optimize it, we revisit the question of non-hierarchical sphere pixelization based on cube symmetries and develop a new one dubbed the ‘Similar Radius Sphere Pixelization’ (SARSPix) with very close to square pixels. It provides the most adapted indexing over the sphere for all distance-related computations. Using LSST-like fast simulations, we compute autocorrelation functions on tomographic bins containing between a hundred million to one billion data points. In each case, we achieve the construction of a standard pair-distance histogram in about 2 min, using a simple algorithm that is shown to scale, over a moderate number of nodes (16–64). This illustrates the potential of this new techniques in the field of astronomy where data access is becoming the main bottleneck. They can be easily adapted to other use-cases as nearest-neighbours search, catalogue cross-match or cluster finding. The software is publicly available from https://github.com/astrolabsoftware/SparkCorr.

79 ASTRONOMY AND ASTROPHYSICS↗

Piecewise linear approximation with minimum number of linear segments and minimum error: A fast approach to tighten and warm start the hierarchical mixed integer formulation

In several areas of economics and engineering, it is often necessary to fit discrete data points or approximate nonlinear functions with continuous functions. Piecewise linear (PWL) functions are a convenient way to achieve this. PWL functions can be modeled in mathematical problems using only linear and integer variables. Moreover, there is a computational benefit in using PWL functions that have the least possible number of segments. This work proposes a novel hierarchical mixed integer linear programming (MILP) formulation that identifies a continuous PWL approximation with minimum number of linear segments for a given target maximum error. The proposed MILP formulation also identifies the solution with the least maximum error among the solutions with minimum number of segments. Then, this work proposes a fast iterative algorithm that identifies non necessarily continuous PWL approximations by solving O(S log N) linear programming (LP) problems, where N is the number of data points and S is the minimum number of segments in the non necessarily continuous case. This work demonstrates that tight bounds for the MILP problem can be derived from these approximations. Next, a fast algorithm is introduced to transform a non necessarily continuous PWL approximation into a continuous one. Finally, the tight bounds and the continuous PWL approximations are used to tighten and warm start the MILP problem. The tightened formulation is shown in experimental results to be more efficient, especially for large data sets, with a solution time that is up to two orders of magnitude less than the existing literature.

97 MATHEMATICS AND COMPUTING↗

Galaxy Cruise: Deep Insights into Interacting Galaxies in the Local Universe

Abstract We present the first results from GALAXY CRUISE, a community (or citizen) science project based on data from the Hyper Suprime-Cam Subaru Strategic Program (HSC-SSP). The current paradigm of galaxy evolution suggests that galaxies grow hierarchically via mergers, but our observational understanding of the role of mergers is still limited. The data from HSC-SSP are ideally suited to improve our understanding with improved identifications of interacting galaxies thanks to the superb depth and image quality of HSC-SSP. We launched a community science project, GALAXY CRUISE, in 2019 and have collected over two million independent classifications of 20686 galaxies at z < 0.2. We first characterize the accuracy of the participants’ classifications and demonstrate that it surpasses previous studies based on shallower imaging data. We then investigate various aspects of interacting galaxies in detail. We show that there is a clear sign of enhanced activities of super-massive black holes and star formation in interacting galaxies compared to those in isolated galaxies. The enhancement seems particularly strong for galaxies undergoing violent mergers. We also show that the mass growth rate inferred from our results is roughly consistent with the observed evolution of the stellar mass function. The second season of GALAXY CRUISE is currently underway and we conclude with future prospects. We make the morphological classification catalog used in this paper publicly available at the GALAXY CRUISE website, which will be particularly useful for machine-learning applications.

Tanaka, Masayuki↗

Challenges and Vision for Standardization of Biopolymer Data Sets for Machine Learning

Machine learning (ML) is transforming materials research, yet potential for biopolymer discovery remains constrained by fragmented data and nonstandardized reporting. Biopolymers differ significantly from synthetic polymers, requiring specialized approaches to represent their biosynthetic origins, hierarchical structures, and application-specific metrics. In this Perspective, we identify three core challenges limiting biopolymer representation: information encoding, data quality, and data sharing. We describe the most pressing issues and propose commensurate approaches to address each key challenge. Recommendations include the design and adoption of biopolymer-specific fingerprinting and representation frameworks, development of hybrid human-large language model (LLM) data extraction strategies, and expanding Findable, Accessible, Interoperable, Reusable (FAIR)-compliant repositories. We propose a robust foundation to define interoperable, high-quality data sets that capture the full context of biopolymer materials. Standardized metadata, shared ontologies, and community-driven infrastructure would enable scalable, reproducible workflows and accelerate the ML-driven development of biopolymers.

36 MATERIALS SCIENCE↗

Coupling Noah-Multiparameterization land-surface Model with Energy Research and Forecasting Model

The Energy Research and Forecasting (ERF) model is a high-performance atmospheric model built on the AMReX adaptive mesh refinement (AMR) framework, enabling efficient simulations on heterogeneous computing platforms that combine multicore processors with hardware accelerators. To support land–atmosphere interactions within ERF’s AMR-based environment, a land-surface model must be capable of operating directly on hierarchically refined meshes. In this work, we present a methodology for coupling the Fortran-based Noah-Multiparameterization (Noah-MP) land-surface model with ERF’s C++ codebase. Rather than rewriting Noah-MP, we construct a Fortran–C interoperability layer using CodeScribe, a tool that leverages large language models (LLMs) to automate the generation of interface code. CodeScribe applies structured prompting techniques to generate bindings that support efficient data exchange and function calls between ERF and Noah-MP. The coupling framework also incorporates AMR-aware data handling strategies, allowing NoahMP to operate seamlessly within ERF’s hierarchical mesh structure. This work provides a structured approach for integrating legacy Fortran models into modern C++-based modeling systems using LLM-assisted code generation.

54 ENVIRONMENTAL SCIENCES↗

Utah FORGE Project 3-2417: DAS Microseismic Event Catalog from the 16A/16B Circulation Test, 2023

This preliminary data archive includes the relocated microseismic event catalog, 1D velocity model, and methods report from DAS acquisition conducted during the Well 16A and 16B circulation test (July 19th and 20th, 2023) at Utah FORGE. The methods report describes all processing steps, including real-time event detection, hierarchical clustering, joint velocity/hypocenter inversion, and relocation. The resulting work is accepted and will be presented at IMAGE 2024. This dataset was acquired by the FOGMORE R&D project (Fiber Optic MOnitoring for Reservoir Evolution), Utah FORGE R&D Project 3-2417.

15 GEOTHERMAL ENERGY↗

AICCA: AI-Driven Cloud Classification Atlas

Clouds play an important role in the Earth’s energy budget, and their behavior is one of the largest uncertainties in future climate projections. Satellite observations should help in understanding cloud responses, but decades and petabytes of multispectral cloud imagery have to date received only limited use. This study describes a new analysis approach that reduces the dimensionality of satellite cloud observations by grouping them via a novel automated, unsupervised cloud classification technique based on a convolutional autoencoder, an artificial intelligence (AI) method good at identifying patterns in spatial data. Our technique combines a rotation-invariant autoencoder and hierarchical agglomerative clustering to generate cloud clusters that capture meaningful distinctions among cloud textures, using only raw multispectral imagery as input. Cloud classes are therefore defined based on spectral properties and spatial textures without reliance on location, time/season, derived physical properties, or pre-designated class definitions. We use this approach to generate a unique new cloud dataset, the AI-driven cloud classification atlas (AICCA), which clusters 22 years of ocean images from the Moderate Resolution Imaging Spectroradiometer (MODIS) on NASA’s Aqua and Terra instruments—198 million patches, each roughly 100 km × 100 km (128 × 128 pixels)—into 42 AI-generated cloud classes, a number determined via a newly-developed stability protocol that we use to maximize richness of information while ensuring stable groupings of patches. AICCA thereby translates 801 TB of satellite images into 54.2 GB of class labels and cloud top and optical properties, a reduction by a factor of 15,000. The 42 AICCA classes produce meaningful spatio-temporal and physical distinctions and capture a greater variety of cloud types than do the nine International Satellite Cloud Climatology Project (ISCCP) categories—for example, multiple textures in the stratocumulus decks along the West coasts of North and South America. We conclude that our methodology has explanatory power, capturing regionally unique cloud classes and providing rich but tractable information for global analysis. AICCA delivers the information from multi-spectral images in a compact form, enables data-driven diagnosis of patterns of cloud organization, provides insight into cloud evolution on timescales of hours to decades, and helps democratize climate research by facilitating access to core data.

97 MATHEMATICS AND COMPUTING↗

Secure hierarchical processing using a secure ledger

Disclosed is a system and method for processing data using blockchain technology. The system includes a memory having programmable instructions stored thereon that, when executed by a processor, cause the system to: authenticate one or more sensors in anticipation of receiving component data; receive component data, upon successful authentication; store the component data locally or to a cloud-based server and/or calculate a root value for the component data; store or embed the root value with the stored component data; condense the component data and link the condensed component data to the stored component data via the root value. The system further includes instructions to log the condensed data, including the root value, to a ledger, and to identify a tag or transaction id corresponding to the logging event for subsequent retrieval of the condensed data using the tag or transaction id.

Zhao, Wenbing↗

Accurate and uncertainty-aware multi-task prediction of HEA properties using prior-guided deep Gaussian processes

Surrogate modeling techniques have become indispensable in accelerating the discovery and optimization of high-entropy alloys (HEAs), especially when integrating computational predictions with sparse experimental observations. This study systematically evaluates the training and testing performance of four prominent surrogate models—conventional Gaussian processes (cGP), Deep Gaussian processes (DGP), encoder-decoder neural networks for multi-output regression and eXtreme Gradient Boosting (XGBoost)—applied to a hybrid dataset of experimental and computational properties of the 8-component HEA system Al-Co-Cr-Cu-Fe-Mn-Ni-V. We specifically assess their capabilities in predicting correlated material properties, including yield strength, hardness, modulus, ultimate tensile strength, elongation, and average hardness under dynamic/quasi-static conditions, alongside auxiliary computational properties. The comparison highlights the strengths of hierarchical deep modeling approaches in handling heteroscedastic, heterotopic, and incomplete data commonly encountered in materials science. Our findings illustrate that combined surrogate models such as DGPs infused with machine-learned priors outperform other surrogates by effectively capturing inter-property correlations and by assimilating prior knowledge. This enhanced predictive accuracy positions the combined surrogate models as powerful tools for robust and data-efficient materials design.

36 MATERIALS SCIENCE↗

TRACE Input Modernization

This work presents a Tom’s Obvious Minimal Language (TOML)-based representation of input for the US Nuclear Regulatory Commission’s TRAC/RELAP Advanced Computational Engine (TRACE) thermal hydraulics code. Implemented using the Workbench Analysis Sequence Processor (WASP), the approach maps traditional TRACE input structures to a hierarchical format composed of named parameters, typed values, and native data collections. The resulting representation preserves TRACE’s existing modeling capabilities while providing a modern, structured interface for model development and management. WASP further extends TOML through a file import directive that supports modular model composition and reusable input organization. In addition, WASP provides extended array data entry convenience with various data repeat and interpolation capabilities. Examples of the new TOML syntax are provided for major TRACE input categories, including hydraulic components, heat structures, control systems, and trip logic. The TOML representation establishes a foundation for improved validation, tooling, automation, and model maintainability while remaining compatible with existing TRACE workflows. To facilitate migration to the TOML-based input format, the TRACE executable now supports conversion of native TRACE input into an intermediate JSON representation. A Python utility subsequently transforms the JSON data into an equivalent TOML model. Lastly, the TRACE executable now supports execution using TOML-formatted input.

Lefebvre, Robert A. [Oak Ridge National Laboratory↗

Iterative self-organizing SCEne-LEvel sampling (ISOSCELES) for large-scale building extraction

Convolutional neural networks (CNN) provide state-of-the-art performance in many computer vision tasks, including those related to remote-sensing image analysis. Successfully training a CNN to generalize well to unseen data, however, requires training on samples that represent the full distribution of variation of both the target classes and their surrounding contexts. With remote sensing data, acquiring a sufficiently representative training set is a challenge due to both the inherent multi-modal variability of satellite or aerial imagery and the general high cost of labeling data. To address this challenge, we have developed ISOSCELES, an Iterative Self-Organizing SCEne LEvel Sampling method for hierarchical sampling of large image sets. Using affinity propagation, ISOSCELES automates the selection of highly representative training images. Compared to random sampling or using available reference data, the distribution of the training is principally data driven, reducing the chance of oversampling uninformative areas or undersampling informative ones. In comparison to manual sample selection by an analyst, ISOSCELES exploits descriptive features, spectral and/or textural, and eliminates human bias in sample selection. Using a hierarchical sampling approach, ISOSCELES can obtain a training set that reflects both between-scene variability, such as in viewing angle and time of day, and within-scene variability at the level of individual training samples. We verify the method by demonstrating its superiority to stratified random sampling in the challenging task of adapting a pre-trained model to a new image and spatial domain for country-scale building extraction. Using a pair of hand-labeled training sets comprising 1,987 sample image chips, a total of 496,000,000 individually labeled pixels, we show, across three distinct model architectures, an increase in accuracy, as measured by F1-score, of 2.2–4.2%.

42 ENGINEERING↗

MAGNET Scaling and Methodology

The purpose of this study was to analyze the heat transfer of the Microreactor Agile Non-nuclear Experimental Testbed (MAGNET) within the Dynamic Energy Transport and Integration Laboratory (DETAIL) and develop scaling equations and models to couple with other systems. Hierarchical Two-Tiered Scaling (H2TS) methodologies were applied to DETAIL’s MAGNET facility to scale and project data sets while conserving the observed behavior based on first principles. The MAGNET system was successfully scaled using H2TS and multiple system parameters were determined or calculated from experimental data including steady state and transient data.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Data-assimilated time-lapse visco-acoustic full-waveform inversion: Theory and application for injected CO 2 plume monitoring

Continuous seismic monitoring for quantifying CO 2 plume migration and detection of any potential leakages in the subsurface is essential for the security of long-term anthropogenic carbon dioxide geologic storage. Traditional time-lapse full-waveform inversion (TLFWI) methods aim to map the CO 2 distribution by estimating seismic velocity changes, but recent studies find that CO 2 -induced attenuation is an important complement to seismic velocity for tracking the CO 2 plumes and even quantifying the CO 2 saturation. We have developed a novel data-assimilated TLFWI method to construct high-resolution time-lapse velocity and attenuation changes from dense time-lapse monitoring data. This method consists of two theoretical developments: visco-acoustic full-waveform inversion (QFWI) and multiparameter hierarchical matrix-powered extended Kalman filter (mHiEKF). The method is capable of (1) posing temporal constraints to retrieve time-lapse information from dense monitoring data by using mHiEKF, (2) accurately recovering high-spatial-resolution velocity and attenuation perturbations using first-order equation system-based QFWI, and (3) providing the model uncertainty by estimating their model standard deviation. With numerical examples, we first find the effectiveness of the new QFWI on estimating accurate velocity and attenuation models simultaneously. Then, a CO 2 leakage case and a realistic Frio-II CO 2 monitoring case are presented to find the advantages and applicability of our data-assimilated QFWI method for estimating time-lapse changes using dense time-lapse monitoring surveys. Here, by assimilating time-lapse seismic monitoring data over time, our data-assimilated QFWI method can improve the resolution of velocity and attenuation changes and decrease their model uncertainties.

58 GEOSCIENCES↗

Development of an open-source regional data assimilation system in PEcAn v. 1.7.2: application to carbon cycle reanalysis across the contiguous US using SIPNET

Abstract. The ability to monitor, understand, and predict the dynamics of the terrestrial carbon cycle requires the capacity to robustly and coherently synthesize multiple streams of information that each provide partial information about different pools and fluxes. In this study, we introduce a new terrestrial carbon cycle data assimilation system, built on the PEcAn model–data eco-informatics system, and its application for the development of a proof-of-concept carbon “reanalysis” product that harmonizes carbon pools (leaf, wood, soil) and fluxes (GPP, Ra, Rh, NEE) across the contiguous United States from 1986–2019. We first calibrated this system against plant trait and flux tower net ecosystem exchange (NEE) using a novel emulated hierarchical Bayesian approach. Next, we extended the Tobit–Wishart ensemble filter (TWEnF) state data assimilation (SDA) framework, a generalization of the common ensemble Kalman filter which accounts for censored data and provides a fully Bayesian estimate of model process error, to a regional-scale system with a calibrated localization. Combined with additional workflows for propagating parameter, initial condition, and driver uncertainty, this represents the most complete and robust uncertainty accounting available for terrestrial carbon models. Our initial reanalysis was run on an irregular grid of ∼ 500 points selected using a stratified sampling method to efficiently capture environmental heterogeneity. Remotely sensed observations of aboveground biomass (Landsat LandTrendr) and leaf area index (LAI) (MODIS MOD15) were sequentially assimilated into the SIPNET model. Reanalysis soil carbon, which was indirectly constrained based on modeled covariances, showed general agreement with SoilGrids, an independent soil carbon data product. Reanalysis NEE, which was constrained based on posterior ensemble weights, also showed good agreement with eddy flux tower NEE and reduced root mean square error (RMSE) compared to the calibrated forecast. Ultimately, PEcAn's new open-source regional data assimilation framework provides a scalable workflow for harmonizing multiple data constraints and providing a uniform synthetic platform for carbon monitoring, reporting, and verification (MRV) as well as accelerating terrestrial carbon cycle research.

54 ENVIRONMENTAL SCIENCES↗

MFNets: data efficient all-at-once learning of multifidelity surrogates as directed networks of information sources

We present an approach for constructing a surrogate from ensembles of information sources of varying cost and accuracy. The multifidelity surrogate encodes connections between information sources as a directed acyclic graph, and is trained via gradient-based minimization of a nonlinear least squares objective. While the vast majority of state-of-the-art assumes hierarchical connections between information sources, our approach works with flexibly structured information sources that may not admit a strict hierarchy. The formulation has two advantages: (1) increased data efficiency due to parsimonious multifidelity networks that can be tailored to the application; and (2) no constraints on the training data—we can combine noisy, non-nested evaluations of the information sources. Finally, numerical examples ranging from synthetic to physics-based computational mechanics simulations indicate the error in our approach can be orders-of-magnitude smaller, particularly in the low-data regime, than single-fidelity and hierarchical multifidelity approaches.

97 MATHEMATICS AND COMPUTING↗