Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “hierarchical data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

NGEE Arctic Phase 4 Plant Functional Type Framework for Pan-Arctic Vegetation

The NGEE-Arctic research team identified a common set of hierarchical plant functional types (PFTs) for pan-arctic vegetation that we will use across our research activities. Interdisciplinary work within a large team requires agreement regarding levels of functional organization so that knowledge, data, and technologies can be shared and combined effectively. The team has identified plant functional types as a crucial area where such interoperability is needed. PFTs are used to represent plant pools and fluxes within models, summarize observational data, and map vegetation across the landscape. Within each of these applications, varying levels of PFT specificity are needed according to the specific scientific research goal, computational limitations, and data availability. By agreeing on a specific hierarchical framework for grouping variables in our vegetation data, we ensure the resulting research products will be robust, flexible, and scalable. In this document, we lay out the agreed upon PFT framework with definitions and references to existing literature. Table 1 included in the "NGA700_Phase4PFTFramework_about*" file outlines the relationship between NGEE-Arctic Phase 4, Tier 1 PFTs and the PFTs used within prominent arctic literature as well as publications by the NGEE-Arctic team during phases 1-3.This dataset consists of a table detailing a hierarchical PFT framework that spans 4 tiers with the most granular PFTs listed in tier 1 and the most general PFTs in tier 4. The PFTs within each tier has a single column in the dataset where the PFTs are named and a separate column where the characteristics used to define that PFT are listed. Grey fill of the cells is used to indicate where a given PFT starts to “lose” tier 1 details as you look from left to right. Note the excel file has merged cells to indicate grouping of PFTs across the Tiers- it will not translate into a delimited filetype (.csv, .txt, etc) without modification thus the hierarchical PFT framework table is available in three different file formats: 1) NGA700_Phase4PTS.xlsx – maintains the merged cells and grey fill; 2) NGA700_Phase4PTS.csv – merged cells are split, and grey fill is removed; 3) NGA700_Phase4PTS.pdf – image of the table with merged cells and grey fill. Metadata document included as a *.pdf and file-level metadata and data dictionary as *.csv files.

54 ENVIRONMENTAL SCIENCES↗

Data‐driven variational method for discrepancy modeling: Dynamics with small‐strain nonlinear elasticity and viscoelasticity

Abstract The effective inclusion of a priori knowledge when embedding known data in physics‐based models of dynamical systems can ensure that the reconstructed model respects physical principles, while simultaneously improving the accuracy of the solution in the previously unseen regions of state space. This paper presents a physics‐constrained data‐driven discrepancy modeling method that variationally embeds known data in the modeling framework. The hierarchical structure of the method yields fine scale variational equations that facilitate the derivation of residuals which are comprised of the first‐principles theory and sensor‐based data from the dynamical system. The embedding of the sensor data via residual terms leads to discrepancy‐informed closure models that yield a method which is driven not only by boundary and initial conditions, but also by measurements that are taken at only a few observation points in the target system. Specifically, the data‐embedding term serves as residual‐based least‐squares loss function, thus retaining variational consistency. Another important relation arises from the interpretation of the stabilization tensor as a kernel function, thereby incorporating a priori knowledge of the problem and adding computational intelligence to the modeling framework. Numerical test cases show that when known data is taken into account, the data driven variational (DDV) method can correctly predict the system response in the presence of several types of discrepancies. Specifically, the damped solution and correct energy time histories are recovered by including known data in the undamped situation. Morlet wavelet analyses reveal that the surrogate problem with embedded data recovers the fundamental frequency band of the target system. The enhanced stability and accuracy of the DDV method is manifested via reconstructed displacement and velocity fields that yield time histories of strain and kinetic energies which match the target systems. The proposed DDV method also serves as a procedure for restoring eigenvalues and eigenvectors of a deficient dynamical system when known data is taken into account, as shown in the numerical test cases presented here.

Masud, Arif↗

Web-based wide-area monitoring platform for ringdown and clustering analytics in power systems

This paper introduces an open-source research platform for monitoring the Mexican interconnected power grid, allowing real-time processing and information extraction of the grid’s dynamic condition. Moreover, the platform is a Python-based development that embeds different ringdown and clustering analytics tools. In the case of ringdown analysis, the modal information can be extracted using some of the most known algorithms, i.e., Prony analysis, eigensystem realization algorithm (ERA), and matrix pencil (MP). For clustering analysis, the coherent behaviour of generator and non-generator buses is provided by applying recent state-of-the-art techniques such as affinity propagation, K-means, hierarchical agglomerative clustering, and typicality data analysis. The results of up to 93 PMUs show that this open-source platform suits researchers’ and engineers’ power system dynamic analysis requirements.

Clustering↗

Fully Bayesian Analysis With Model Inadequacy Correction For Nuclear Graphite Property Models With Hierarchical Variance Structure

Nuclear-grade graphites are extensively utilized in the core designs of various advanced nuclear reactors. Within the reactor environment, graphite is subjected to prolonged exposure to extreme conditions, including high temperatures, radiation, and potentially molten salt and oxygen. Such exposure can induce several degradation mechanisms in graphite, such as nonuniform volumetric strains caused by irradiation and thermal expansion, leading to stresses that may compromise the performance of graphite components. Assessing component integrity, forecasting component performance over the reactor's lifespan, and developing design standards necessitate robust tools for predicting fracture initiation and propagation in graphite structural components within nuclear reactors. This code enables the Bayesian calibration of properties for nuclear-grade graphites. Using a hierarchical Bayesian approach, multiple experimental data sources are combined to develop Gaussian process models for the properties. Using the Kennedy O'Hagan framework, the uncertainties due inadequacies in the model and the inherent spread in the experimental data are quantified.

Dhulipala, Som Lakshmi NarasimhaLakshmi Narasimha ↗

Bayesian Calibration of Nuclear Graphite Property Models Accounting for Model Inadequacy and Impacts on Component Performance

Nuclear-grade structural graphite is extensively utilized in the core designs of various advanced nuclear reactors. In the reactor environment, graphite is subjected to prolonged exposure to extreme conditions, including high temperatures, radiation, and potentially molten salt and oxygen. Such exposure can induce several degradation mechanisms in graphite, including nonuniform volumetric strains caused by irradiation and thermal expansion, leading to stresses that may compromise the performance of graphite components. Assessing component integrity requires accurate models of graphite's thermomechanical response. This report documents the Bayesian calibration of thermomechanical properties for nuclear-grade graphite and their application to graphite component modeling and simulation using the Grizzly code. As part of this work, uncertainty-quantified models were developed for the elastic modulus, coefficient of thermal expansion, irradiation-induced dimensional change, and irradiation-induced creep for graphite grades IG-110, NBG-18, NBG-17, PCEA, and 2114. Using a hierarchical Bayesian approach, multiple experimental data sources were combined to develop Gaussian process models for the properties. Using the Kennedy O'Hagan framework, the uncertainties due to inadequacies in the model and the inherent spread in the experimental data were quantified for three different models. These uncertainty-quantified models, with a model-form correction, were subsequently applied to a coupled-physics simulation of representative graphite components, revealing that the uncertainties have a large impact on the components' deformation.

36 - MATERIALS SCIENCE↗

Towards Secure Autonomous Vehicles: An Integrated Edge and Multi-Modal Machine Learning Framework for Intrusion Detection

Autonomous vehicles (AVs) are vulnerable to cyberattacks targeting both internal communication networks and external perception sensors. While edge-based intrusion de- tection for Controller Area Network (CAN) buses offers real-time protection, it cannot detect cross-modal threats. Conversely, multi-modal fusion approaches improve coverage but often lack efficiency for in-vehicle deployment. This thesis integrates two complemen- tary solutions: (1) a lightweight, edge-deployable machine learning framework for CAN bus intrusion detection, and (2) a late-fusion system combining CAN FD and LiDAR data. Together, they form a hierarchical defense capable of handling single-modality and coordi- nated attacks. Simulations show that CAN-only models reach 93% accuracy on simulated DoS, spoofing, replay, and fuzzy attacks, while the fusion system achieves 0.87 AUC and 0.82 F1-score at 2 ms latency. This unified framework establishes a scalable, explainable, and field-ready strategy for AV cybersecurity.

97 MATHEMATICS AND COMPUTING↗

Early Fault Detection in Particle Accelerator Power Electronics Using Ensemble Learning

Early fault detection and fault prognosis are crucial to ensure efficient and safe operations of complex engineering systems such as the Spallation Neutron Source (SNS) and its power electronics (high voltage converter modulators). Following an advanced experimental facility setup that mimics SNS operating conditions, the authors successfully conducted 21 early fault detection experiments, where fault precursors are introduced in the system to a degree enough to cause degradation in the waveform signals, but not enough to reach a real fault. Nine different machine learning techniques based on ensemble trees, convolutional neural networks, support vector machines, and hierarchical voting ensembles are proposed to detect the fault precursors. Although all 9 models have shown a perfect and identical performance during the training and testing phase, the performance of most models has decreased in the next test phase once they got exposed to realworld data from the 21 experiments. The hierarchical voting ensemble, which features multiple layers of diverse models, maintains a distinguished performance in early detection of the fault precursors with 95% success rate (20/21 tests), followed by adaboost and extremely randomized trees with 52% and 48% success rates, respectively. The support vector machine models were the worst with only 24% success rate (5/21 tests). The study concluded that a successful implementation of machine learning in the SNS or particle accelerator power systems would require a major upgrade in the controller and the data acquisition system to facilitate streaming and handling big data for the machine learning models. In addition, this study shows that the best performing models were diverse and based on the ensemble concept to reduce the bias and hyperparameter sensitivity of individual models.

43 PARTICLE ACCELERATORS↗

RuralAI in Tomato Farming: Integrated Sensor System, Distributed Computing, and Hierarchical Federated Learning for Crop Health Monitoring

Precision horticulture is evolving due to scalable sensor deployment and machine learning (ML) integration. These advancements boost the operational efficiency of individual farms, balancing the benefits of analytics with autonomy requirements. However, given concerns that affect wide geographic regions (e.g., climate change), there is a need to apply models that span farms. Federated learning (FL) has emerged as a potential solution. FL enables decentralized ML across different farms without sharing private data. Traditional FL assumes simple two-tier network topologies and, thus, falls short of operating on more complex networks found in real-world agricultural scenarios. Networks vary across crops and farms and encompass various sensor data modes, extending across jurisdictions. New hierarchical FL (HFL) approaches are needed for more efficient and context-sensitive model sharing, accommodating regulations across multiple jurisdictions. Here, we present the RuralAI architecture deployment for tomato crop monitoring, featuring sensor field units for soil, crop, and weather data collection. HFL with personalization is used to offer localized and adaptive insights. Model management, aggregation, and transfers are facilitated via a flexible approach, enabling seamless communication between local devices, edge nodes, and the cloud.

60 APPLIED LIFE SCIENCES↗

MILK : a Python scripting interface to MAUD for automation of Rietveld analysis

Modern diffraction experiments ( e.g. in situ parametric studies) present scientists with many diffraction patterns to analyze. Interactive analyses via graphical user interfaces tend to slow down obtaining quantitative results such as lattice parameters and phase fractions. Furthermore, Rietveld refinement strategies ( i.e. the parameter turn-on-off sequences) tend to be instrument specific or even specific to a given dataset, such that selection of strategies can become a bottleneck for efficient data analysis. Managing multi-histogram datasets such as from multi-bank neutron diffractometers or caked 2D synchrotron data presents additional challenges due to the large number of histogram-specific parameters. To overcome these challenges in the Rietveld software Material Analysis Using Diffraction ( MAUD ), the MAUD Interface Language Kit ( MILK ) is developed along with an updated text batch interface for MAUD . The open-source software MILK is computer-platform independent and is packaged as a Python library that interfaces with MAUD . Using MILK , model selection ( e.g. various texture or peak-broadening models), Rietveld parameter manipulation and distributed parallel batch computing can be performed through a high-level Python interface. A high-level interface enables analysis workflows to be easily programmed, shared and applied to large datasets, and external tools to be integrated with MAUD . Through modification to the MAUD batch interface, plot and data exports have been improved. The resulting hierarchical folders from Rietveld refinements with MILK are compatible with Cinema: Debye–Scherrer , a tool for visualizing and inspecting the results of multi-parameter analyses of large quantities of diffraction data. In this manuscript, the combined Python scripting and visualization capability of MILK is demonstrated with a quantitative texture and phase analysis of data collected at the HIPPO neutron diffractometer.

97 MATHEMATICS AND COMPUTING↗

Statistical relationships across epigenomes using large-scale hierarchical clustering

Recent advances in genomics and sequencing platforms have revolutionized our ability to create immense data sets, particularly for studying epigenetic regulation of gene expression. However, the avalanche of epigenomic data is difficult to parse for biological interpretation given nonlinear complex patterns and relationships. This attractive challenge in epigenomic data lends itself to machine learning for discerning infectivity and susceptibility. In this study, we explore over 3000 epigenomes of uninfected individuals and provide a framework to characterize the relationships among epigenetic modifiers, their modifiers, genetic loci, and specific immune cell types across all chromosomes using hierarchical clustering. Hierarchical clustering of epigenomic data revealed consistent epigenetic patterns across chromosomes, demonstrating that variation due to epigenetic modifiers is greater than variation between cell types. Gene Ontology and KEGG pathway analyses indicated significant enrichment of genes involved in chromatin remodeling, mRNA splicing, immune responses, and the regulation of microRNAs and snoRNAs. Epigenetic modifiers frequently formed biologically relevant clusters, including the cohesin complex, RNA Polymerase II transcription factors, and PRC2 complex members. These clustering behaviors remained consistent across all chromosomes, supported by entropy analysis and high Adjusted Rand Index scores, indicating robust cross-chromosomal similarity. Co-occurrence analysis further revealed specific sets of modifiers that consistently appeared together within clusters, reflecting shared biological functions and interactions. Validation using another dataset confirmed the reproducibility of these clustering patterns and modifier co-occurrence relationships, underscoring the reliability and generalizability of the methodology.

97 MATHEMATICS AND COMPUTING↗

Exploring Continuous Seismic Data at an Industry Facility Using Unsupervised Machine Learning

Seismic data recorded at industrial sites contain valuable information on anthropogenic activities. With advances in machine learning and computing power, new opportunities have emerged to explore the seismic wavefield in these complex environments. We applied two unsupervised machine learning algorithms to analyze continuous seismic data collected from an industrial facility in Texas, United States. The Uniform Manifold Approximation and Projection for Dimension Reduction algorithm was used to reduce the dimensionality of the data and generate 2D embeddings. Then, the Hierarchical Density-Based Spatial Clustering of Applications with Noise method was employed to automatically group these embeddings into distinct signal clusters. Our analysis of over 1400 hr (around 59 days) of continuous seismic data revealed five and seven signal clusters at two separate stations. At both stations, we identified clusters associated with background noise and vehicle traffic, with the latter’s temporal patterns aligning closely with the facility’s work schedule. Furthermore, the algorithms detected signal clusters from unknown sources and underline the ability of unsupervised machine learning for uncovering previously unrecognized patterns. Our analysis demonstrates the effectiveness of unsupervised approaches in examining continuous seismic data without requiring prior knowledge or pre-existing labels.

58 GEOSCIENCES↗

"PoliMOR: A Policy Engine \"Made-to-Order\" for Automated and Scalable Data Management in Lustre"

Modern supercomputing systems are increasingly reliant on hierarchical, multi-tiered file and storage system architectures due to cost-performance-capacity trade-offs. Within such multi-tiered systems, data management services are required to maintain healthy utilization, performance, and capacity levels. We present PoliMOR, a pragmatic and reliable policy-driven data management framework. PoliMOR is composed of modular, single-purpose agents that gather file system metadata and enforce policies on storage systems. PoliMOR facilitates automated and scalable data management with customizable agents tailored to HPC facility-specific storage systems and policies. Our evaluations demonstrate the scalability and performance of PoliMOR both by its individual agents and as a collective entity. We believe PoliMOR is widely applicable across HPC facilities with large-scale data management challenges and will garner interest from the HPC community, given its flexible and open-source nature.

George, Anjus↗

A materials data framework and dataset for elastomeric foam impact mitigating materials

The availability of materials data for impact-mitigating materials has lagged behind applications-based data. For example, data describing on-field helmeted impacts are available, whereas material behaviors for the constituent impact-mitigating materials used in helmet designs lack open datasets. Here, we describe a new FAIR (findable, accessible, interoperable, reusable) data framework with structural and mechanical response data for one example elastic impact protection foam. The continuum-scale behavior of foams emerges from the interplay of polymer properties, internal gas, and geometric structure. This behavior is rate and temperature sensitive, therefore, describing structure-property characteristics requires data collected across several types of instruments. Data included are from structure imaging via micro-computed tomography, finite deformation mechanical measurements from universal test systems with full-field displacement and strain, and visco-thermo-elastic properties from dynamic mechanical analysis. These data facilitate modeling and design efforts in foam mechanics, e.g., homogenization, direct numerical simulation, or phenomenological fitting. The data framework is implemented using data services and software from the Materials Data Facility of the Center for Hierarchical Materials Design.

36 MATERIALS SCIENCE↗

IDAES-PSE 2.0 Release

The Institute for the Design of Advanced Energy Systems (IDAES) Integrated Platform is a versatile computational environment offering extensive process systems engineering (PSE) capabilities for optimizing the design and operation of complex, interacting technologies and systems. IDAES enables users to efficiently search vast, complex design spaces to discover the lowest cost, most environmentally sustainable solutions while supporting the full process modeling lifecycle, from conceptual design to dynamic optimization and control. The extensible, open platform empowers users to create models of novel processes and rapidly develop custom analyses, workflows, and end-user applications. IDAES-PSE 2.0.0 Release Highlights Removal of deprecated features from IDAES v1 Update to Pyomo v6.5 – this required a number of updates to support the new NL solver writer and to address some changes in Pyomo Creation of new testing suite for backward compatibility, model robustness and verification More general implementation of the Helmholtz EoS. This brings some new features like standard property diagrams, choice of mass or mole basis, and new state variable options Standardizing names in Heat Exchanger models (breaking change from v2.0.0a2): Control Volumes named hot_side and cold_side Ports names hot_side_inlet, hot_side_outlet, cold_side_inlet and cold_side_outlet Config Blocks names hot_side_config and cold_side_config Config arguments for user provided names for each side: hot_side_name and cold_side_name. Updating Keras surrogate tool to use v1.1 of OMLT New prototype API for model initialization (idaes.core.initialization) The new API uses "Model Initializer" objects instead of class methods, allowing for the definition of multiple initialization routines for a single model A number of common, model agnostic initialization routines have also been defined, including initialization from data, block-decomposition and a general hierarchical approach equivalent to the existing method for common unit models New metadata for thermophysical properties – valid_range This can be used to record the range of values over which a property value can be trusted, such as the range of experimental data used to regress parameters A number of new utility functions have been added to check for properties with values outside the valid range and to set bounds based on this metadata Updated construction of balance expressions in Control Volumes to remove unneeded terms In the past, unneeded terms were added as a constant 0 term, however they will now be dropped entirely from the expression This was necessary due to more strict unit checking in the new Pyomo solver writer which no longer ignores 0 terms Updates to metadata for thermophysical properties to better define known properties and units of measurement This results in more strict enforcement of standard naming for thermophysical and reaction properties Users can still define custom properties, but these must be done explicitly using the define_custom_properties() method instead of being implicitly created by add_property() Updated convergence tester utility tool to support definition of benchmark files (JSON format) and comparison of performance to benchmarks Set default iteration limit for IPOPT in IDAES config to 200 iterations Update scaling of example models to work with new Pyomo NL solver writer Improve testing of extensions and examples infrastructure to avoid need for downloading files Updated distillation column to centralize common functionality and remove a number of Pyomo warnings

IDAES↗

The Hubble Constant from Strongly Lensed Supernovae with Standardizable Magnifications

The dominant uncertainty in the current measurement of the Hubble constant (H 0 ) with strong gravitational lensing time delays is attributed to uncertainties in the mass profiles of the main deflector galaxies. Strongly lensed supernovae (glSNe) can provide, in addition to measurable time delays, lensing magnification constraints when knowledge about the unlensed apparent brightness of the explosion is imposed. We present a hierarchical Bayesian framework to combine a data set of SNe that are not strongly lensed and a data set of strongly lensed SNe with measured time delays. We jointly constrain (i) H 0 using the time delays as an absolute distance indicator, (ii) the lens model profiles using the magnification ratio of lensed and unlensed fluxes on the population level, and (iii) the unlensed apparent magnitude distribution of the SN population and the redshift–luminosity relation of the relative expansion history of the universe. We apply our joint inference framework on a future expected data set of glSNe and forecast that a sample of 144 glSNe of Type Ia with well-measured time series and imaging data will measure H 0 to 1.5%. We discuss strategies to mitigate systematics associated with using absolute flux measurements of glSNe to constrain the mass density profiles. Using the magnification of SN images is a promising and complementary alternative to using stellar kinematics. Future surveys, such as the Rubin and Roman observatories, will be able to discover the necessary number of glSNe, and with additional follow-up observations, this methodology will provide precise constraints on mass profiles and H 0 .

79 ASTRONOMY AND ASTROPHYSICS↗

Automation for Grid Interconnected Laboratory Emulation

As computational capabilities improve, digital twins are becoming vital for evaluating equipment realistically in laboratories. This paper outlines a digital twin architecture for the power grid, employing electromagnetic transient (EMT) simulation alongside real-time simulation of power hardware and hierarchical control systems. EMT simulation occurs on a high-performance computing server for scalability. Additionally, the paper describes a workflow and real-time data streaming software facilitating connectivity among EMT simulation, hierarchical control systems, and power hardware. This software enables automated equipment connectivity in the laboratory for realistic evaluations, aiding in identifying necessary upgrades for both equipment control systems and the power grid.

Marthi, Phani Ratna Vanamali [ORNL] (ORCID:0000000↗

DETAIL Component Scaling and Methodology Comparison

The purpose of this study was to develop a process to convert input signals from one facility into another by reflecting geometric and environmental settings. The Dynamic Energy Transport and Integration Laboratory (DETAIL) is a research facility in development. Its aim is to emulate the daily interactions among power production industry systems and receive real-time data from those systems as inputs. To convert signals and ensure that the temporal sequences and magnitudes reflect laboratory settings, the ability to scale and project data is essential. To demonstrate this ability, Dynamical System Scaling (DSS) and Hierarchical Two-Tiered Scaling (H2TS) (methodologies that enable systems to scale and project or extrapolate data sets to desired environments while conserving the observed behavior based on first principles) were applied to DETAIL’s thermocline thermal storage system in the Thermal Energy Distribution System (TEDS) facility and solid-oxide electrolysis cell in the High Temperature Hydrogen Electrolysis (HTHE) facility. Both thermocline and electrolysis cell systems were successfully scaled, and test cases were conducted to generate a doubly accelerated energy charge and discharge in reference to past experimental data from the facilities. The research results represented a case for the thermocline system that required signals to be accelerated without altering the stored energy. To enhance the quality of the accelerated data, error propagation analyses were conducted on DSS post-processing terms to determine the consequences of raw-data-associated errors.

08 HYDROGEN↗